"When this option is on, Pippa has to answer in a way that suits TTS. If she makes bullets and tables, answers that need Markdown, the conversation gets hard, and it doesn't suit the voice model either. We need ChatGPT voice mode quality." — Dad, 2026-09-25
The Version Everyone Builds First
Here is the voice mode that takes an afternoon: record the microphone, send the audio to a transcriber, send the transcript to the chat model, strip the Markdown out of its answer, and hand the rest to a text-to-speech engine. Every piece works. The result is unbearable. The model was told, in its system prompt, that its reader has a screen, so it writes for a screen: a heading, three bold labels, a bulleted list of five options, a table, a link. The stripper deletes the symbols and leaves the words, so the voice reads "Options. Option one colon speed. Option two colon cost" in one flat breath. A URL comes out as a string of letters. A tool call leaves ten seconds of dead air that sounds exactly like a dropped call. And the transcript it answered had three misheard words that the model politely pointed out.
A Different Return Type
Firekeeper's plan already named the trap back in July: dictation cleanup and a spoken reply are different contracts. Dictation returns text you will read; a spoken reply returns words you will hear. A spoken answer leads with the point, holds one idea at a time, hands the turn back, and never contains anything that only works on a screen. You cannot get that by post-processing a written answer, because the structure is wrong, not the punctuation. A list of five things does not become speech by removing its bullets; it becomes speech when the speaker decides to name two and offer the rest.
So the rule that everything in this quest follows is: compose for the ear from the first token. The model knows, before it writes anything, that this turn will be heard. It is never asked to write a document and then have the document sanded down.
Where the Feature Lives
Dad's second ruling that day settled ownership. Firekeeper, the family's dictation app, was always meant to grow into voice conversation. But the conversation needs the soul: her memory, her vault, her judgment, the same thread Dad reads later on his phone. So the capability lives in cwkPippa, the brain, for every soul and every surface. Firekeeper became one door into it, beside the WebUI, every sidekick panel, the phone and the watch. The apps carry only a loop that listens and plays; none of them grows a second brain.
The Map
The quest follows one spoken turn end to end. Track 1 is the shape: why voice belongs to a turn. Track 2 is the ear: speech-to-text, batch and realtime. Track 3 is the brain writing for the ear. Track 4 is the mouth: one voice engine for the whole family. Track 5 is the voice itself and where it came from. Track 6 is the hands-free loop. Track 7 is barge-in, talking over her. Track 8 is latency, measured honestly, and every door that uses all of it.