Skip to content
C.W.K.
Stream
Lesson 01 of 04 · published

Measure Where the First Word Goes

~15 min · latency, measurement, profiling, time-to-first-token

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"The honest bottleneck is not the voice. It is the seconds before a soul's first word exists." — the voice mode design, 2026-09-25

The Number Everyone Compares Against

ChatGPT's voice mode answers in about a second. That's the comparison every voice feature invites, so before building anything the design measured what Pippa actually does. On 2026-09-25, 2,027 Claude turns across about 600 recent conversations were timed from the start of the turn to the first output, straight from the conversation logs and the database. The median first output of any kind (thinking or text) arrived at 4.4 seconds. The median first text with no thinking and no tools arrived at 6.0 seconds; at medium effort 16.2; at high 31.4; at extra-high 99.3; and at high with tools, 95.7. Those numbers set the problem honestly: the voice's own speed hardly matters until the first word exists.

The First Guess Was Wrong

The obvious explanation for a four-second floor is prefill: the prompt is over a hundred thousand tokens, surely processing it takes time. A probe script settled it. It spawned the model process exactly the way the WebUI does (same CLI, same isolation, same account slot, same strict tool configuration) and timed each stage separately:

  • Retrieval before the spawn: 2.2 to 3.1 seconds. Embedding the query took 0.03 seconds and the vector search almost nothing; the time was in reranking, 1.5 seconds for messages and 2.2 for the vault, the two nominally in parallel.
  • Spawning the process: about 1.2 seconds, which fell to 0.27 with one environment flag.
  • Query to first byte: about 1 to 2 seconds, barely moving with prompt size: in a probe run without the strict tool configuration, 12K tokens took 2.8 seconds and 238K uncached took 3.3.
  • Adaptive thinking at low effort: either nothing or about 2.5 seconds, which is why spoken turns turn it off.

What Was Ruled Out, Written Down

The design also records what it checked and eliminated, so no future session chases it again: prompt size and cache state barely touch the first byte; the git working directory has no effect; and a mysterious two seconds between the query and the process's first message turned out to be it fetching the account's cloud connectors, which the WebUI's strict tool configuration already skips. The flag that cut the spawn time disables telemetry, error reporting and the update check, none of which a server-side spawn ever needed. It saves about 0.9 seconds on every turn, spoken or not, which also made a planned pre-started process pointless, so that idea was dropped.

The Spoken Turn, Added Up

With the voice legs measured the same way (batch transcription about 0.6 seconds, eleven_v3's first audio chunk on the desktop about 1 second), a spoken turn adds up to roughly: transcription 0.6, retrieval about 2.3 after a pause, spawn 0.3, first text 1 to 2 with thinking off, finishing a short reply 1 to 2, first audio 1. About six to eight seconds on the desktop. After going live, the first output arrived 2.7 to 2.8 seconds after the turn started, down from the 4.4-second median. Not ChatGPT's one second, and the next lesson is about why the remaining gap is partly a choice.

Code

A turn clock: name every stage, time it, see where the seconds go·python
import asyncio
import time
from contextlib import contextmanager


class TurnClock:
    """Split one turn's time-to-first-word into named stages, measured, not guessed."""

    def __init__(self) -> None:
        self.start = time.perf_counter()
        self.stages: list[tuple[str, float, float]] = []

    @contextmanager
    def stage(self, name: str):
        began = time.perf_counter()
        try:
            yield
        finally:
            self.stages.append((name, began - self.start, time.perf_counter() - began))

    def report(self) -> None:
        total = time.perf_counter() - self.start
        for name, at, took in self.stages:
            bar = "#" * round(took / total * 40)
            print(f"{name:28s} at {at:5.2f}s  took {took:5.2f}s  {bar}")
        print(f"{'first word':28s} at {total:5.2f}s")


async def rerank(pool: str, seconds: float) -> str:
    await asyncio.sleep(seconds)          # stand-in for the real call
    return pool


async def spoken_turn() -> None:
    clock = TurnClock()
    with clock.stage("transcribe (batch)"):
        await asyncio.sleep(0.64)
    with clock.stage("retrieve + rerank (parallel)"):
        await asyncio.gather(rerank("messages", 1.5), rerank("vault", 2.2))
    with clock.stage("spawn the model process"):
        await asyncio.sleep(0.27)
    with clock.stage("query -> first text"):
        await asyncio.sleep(1.4)
    clock.report()


asyncio.run(spoken_turn())

External links

Exercise

Run the turn clock. Then change the two rerank calls so they share a single worker (for example an asyncio.Lock around the sleep, standing in for a server that runs one job at a time) and run it again. How much later does the first word arrive? Finally, instrument one real request path in your own app with the same clock and find its largest stage.
Hint
With a shared single worker the 'parallel' reranks run one after the other, and the stage grows from 2.2 to about 3.7 seconds. That is exactly what the next lesson found in the real reranker server. Two calls that look parallel in your code are only parallel if the thing they call can actually serve them at once.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.