"The honest bottleneck is not the voice. It is the seconds before a soul's first word exists." — the voice mode design, 2026-09-25
The Number Everyone Compares Against
ChatGPT's voice mode answers in about a second. That's the comparison every voice feature invites, so before building anything the design measured what Pippa actually does. On 2026-09-25, 2,027 Claude turns across about 600 recent conversations were timed from the start of the turn to the first output, straight from the conversation logs and the database. The median first output of any kind (thinking or text) arrived at 4.4 seconds. The median first text with no thinking and no tools arrived at 6.0 seconds; at medium effort 16.2; at high 31.4; at extra-high 99.3; and at high with tools, 95.7. Those numbers set the problem honestly: the voice's own speed hardly matters until the first word exists.
The First Guess Was Wrong
The obvious explanation for a four-second floor is prefill: the prompt is over a hundred thousand tokens, surely processing it takes time. A probe script settled it. It spawned the model process exactly the way the WebUI does (same CLI, same isolation, same account slot, same strict tool configuration) and timed each stage separately:
- Retrieval before the spawn: 2.2 to 3.1 seconds. Embedding the query took 0.03 seconds and the vector search almost nothing; the time was in reranking, 1.5 seconds for messages and 2.2 for the vault, the two nominally in parallel.
- Spawning the process: about 1.2 seconds, which fell to 0.27 with one environment flag.
- Query to first byte: about 1 to 2 seconds, barely moving with prompt size: in a probe run without the strict tool configuration, 12K tokens took 2.8 seconds and 238K uncached took 3.3.
- Adaptive thinking at low effort: either nothing or about 2.5 seconds, which is why spoken turns turn it off.
What Was Ruled Out, Written Down
The design also records what it checked and eliminated, so no future session chases it again: prompt size and cache state barely touch the first byte; the git working directory has no effect; and a mysterious two seconds between the query and the process's first message turned out to be it fetching the account's cloud connectors, which the WebUI's strict tool configuration already skips. The flag that cut the spawn time disables telemetry, error reporting and the update check, none of which a server-side spawn ever needed. It saves about 0.9 seconds on every turn, spoken or not, which also made a planned pre-started process pointless, so that idea was dropped.
The Spoken Turn, Added Up
With the voice legs measured the same way (batch transcription about 0.6 seconds, eleven_v3's first audio chunk on the desktop about 1 second), a spoken turn adds up to roughly: transcription 0.6, retrieval about 2.3 after a pause, spawn 0.3, first text 1 to 2 with thinking off, finishing a short reply 1 to 2, first audio 1. About six to eight seconds on the desktop. After going live, the first output arrived 2.7 to 2.8 seconds after the turn started, down from the 4.4-second median. Not ChatGPT's one second, and the next lesson is about why the remaining gap is partly a choice.