"The system prompt never changes. That is the whole trick behind free switching." — the voice-mode design doc, which is right about the prompt and too generous about 'free'
How Big a Turn Really Is
A soul's system prompt is not a paragraph. It carries her identity files, the vault notes that load every session, the family roster, formatting rules, the emotion-tag instruction, and more. Measured on 2026-09-25, a Claude turn's prompt was about 140K tokens at the median and about 217K at the 90th percentile, and about 67K of the median was read back from the prompt cache instead of being processed fresh. That cache is not a nicety. It is most of what makes a long conversation affordable and quick to start.
Prompt caching is a prefix match. The provider stores the processed prefix of your request, and the next request reuses it only if every byte up to the breakpoint is identical. Change one character near the top and everything after it is processed again at full price. So the house has a test that fails the build if a conversation's system prompt stops being byte-identical from turn to turn.
Where a Voice Flag Would Break It
Now look at what the system prompt already says. It contains a readability instruction that describes a visual reading surface and asks for Markdown headings, bold labels and bullets where they help, plus an instruction for rich response parts. Those are exactly the rules a spoken turn must not follow. The naive fix is to swap them out in the system prompt when voice is on. That fix rewrites the top of the prompt on every switch, and every switch throws away the cache. Dad switches freely, by his own rule, so the naive fix would bill him for a cold 140K-token prompt over and over.
The Per-Turn Prefix
cwkPippa already had a place for things that change every turn. The context engine builds a small block of per-turn material in front of the current user message: the clock, the conversation timeline, the retrieved memory, and the questions a memory service answered. That block is sent with the turn but never written into the conversation's record, and all seven chat routes share it. Voice instructions ride there. When reply_modality is spoken, a spoken-reply instruction is added that says, in its first lines, that for this turn only it replaces the readability and rich-parts guidance in the system prompt. When input_modality is dictated, a dictated-input instruction is added. The system prompt itself never learns that voice exists.
The same discipline extends to settings. A spoken turn runs at a faster reasoning posture than the conversation's own, but the system prompt is always built from the conversation's own level, so its vessel-meta line doesn't change on a voice switch either. That choice has a side effect you will meet in Track 3: the soul needs to be told, per turn, which posture this turn actually runs at.
A still system prompt does not make the switch free, though, and the design doc's own word overstates it. Reasoning settings are part of what the cache is keyed on. Changing the top-level effort, or turning thinking off, between two requests always invalidates the cached conversation history, and on some models the cached system prompt as well. A spoken turn runs at a lower effort than the written turn before it, so the first turn after each switch re-reads the history at full price, even though no byte of the system prompt moved; the next turn in the same mode reads it from the cache again. The newest Claude models offer a way around even that: a system-role message inside the conversation that carries its own effort changes the setting from the next turn on without invalidating the cache. The example below sidesteps the question by pinning one effort for every turn.
The General Rule
Put what is stable at the top and what varies at the bottom. A per-turn fact never belongs in the system prompt, however natural it feels to write it there, because the system prompt is the most expensive place in the request to change.