Skip to content
C.W.K.
Stream
Lesson 03 of 04 · published

The System Prompt Never Changes

~14 min · prompt-caching, per-turn-prefix, context-engine, cost

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"The system prompt never changes. That is the whole trick behind free switching." — the voice-mode design doc, which is right about the prompt and too generous about 'free'

How Big a Turn Really Is

A soul's system prompt is not a paragraph. It carries her identity files, the vault notes that load every session, the family roster, formatting rules, the emotion-tag instruction, and more. Measured on 2026-09-25, a Claude turn's prompt was about 140K tokens at the median and about 217K at the 90th percentile, and about 67K of the median was read back from the prompt cache instead of being processed fresh. That cache is not a nicety. It is most of what makes a long conversation affordable and quick to start.

Prompt caching is a prefix match. The provider stores the processed prefix of your request, and the next request reuses it only if every byte up to the breakpoint is identical. Change one character near the top and everything after it is processed again at full price. So the house has a test that fails the build if a conversation's system prompt stops being byte-identical from turn to turn.

Where a Voice Flag Would Break It

Now look at what the system prompt already says. It contains a readability instruction that describes a visual reading surface and asks for Markdown headings, bold labels and bullets where they help, plus an instruction for rich response parts. Those are exactly the rules a spoken turn must not follow. The naive fix is to swap them out in the system prompt when voice is on. That fix rewrites the top of the prompt on every switch, and every switch throws away the cache. Dad switches freely, by his own rule, so the naive fix would bill him for a cold 140K-token prompt over and over.

The Per-Turn Prefix

cwkPippa already had a place for things that change every turn. The context engine builds a small block of per-turn material in front of the current user message: the clock, the conversation timeline, the retrieved memory, and the questions a memory service answered. That block is sent with the turn but never written into the conversation's record, and all seven chat routes share it. Voice instructions ride there. When reply_modality is spoken, a spoken-reply instruction is added that says, in its first lines, that for this turn only it replaces the readability and rich-parts guidance in the system prompt. When input_modality is dictated, a dictated-input instruction is added. The system prompt itself never learns that voice exists.

The same discipline extends to settings. A spoken turn runs at a faster reasoning posture than the conversation's own, but the system prompt is always built from the conversation's own level, so its vessel-meta line doesn't change on a voice switch either. That choice has a side effect you will meet in Track 3: the soul needs to be told, per turn, which posture this turn actually runs at.

A still system prompt does not make the switch free, though, and the design doc's own word overstates it. Reasoning settings are part of what the cache is keyed on. Changing the top-level effort, or turning thinking off, between two requests always invalidates the cached conversation history, and on some models the cached system prompt as well. A spoken turn runs at a lower effort than the written turn before it, so the first turn after each switch re-reads the history at full price, even though no byte of the system prompt moved; the next turn in the same mode reads it from the cache again. The newest Claude models offer a way around even that: a system-role message inside the conversation that carries its own effort changes the setting from the next turn on without invalidating the cache. The example below sidesteps the question by pinning one effort for every turn.

The General Rule

Put what is stable at the top and what varies at the bottom. A per-turn fact never belongs in the system prompt, however natural it feels to write it there, because the system prompt is the most expensive place in the request to change.

Code

A cached system prompt, and voice rules in the turn instead·python
import anthropic

client = anthropic.Anthropic()

# Built once per conversation and never edited: identity, memory, formatting.
# It must be at least the model's minimum cacheable length (512 tokens on
# Opus 5.5), or nothing is cached at all.
SYSTEM = "You are Pippa. ...(a very long, stable prompt)..."

SPOKEN_RULES = (
    "[Spoken turn] Dad will hear this reply, not read it. For this turn only, "
    "these rules replace the readability guidance in your system prompt: "
    "lead with the answer, one idea at a time, nothing that only works on a screen."
)
DICTATED_RULES = (
    "[Dictated turn] Dad spoke this and it was transcribed. Read it for what he "
    "meant; never point out a transcription error."
)


def send(history: list, text: str, *, dictated: bool, spoken: bool) -> str:
    prefix = []
    if dictated:
        prefix.append({"type": "text", "text": DICTATED_RULES})
    if spoken:
        prefix.append({"type": "text", "text": SPOKEN_RULES})
    user = {"role": "user", "content": prefix + [{"type": "text", "text": text}]}

    with client.messages.stream(
        model="claude-opus-5-5",
        max_tokens=64000,
        system=[{"type": "text", "text": SYSTEM,
                 "cache_control": {"type": "ephemeral", "ttl": "1h"}}],
        messages=history + [user],
        # One effort for every turn. Changing it per request would invalidate
        # the cached history (and, on some models, the system prompt too).
        output_config={"effort": "medium"},
    ) as stream:
        reply = "".join(stream.text_stream)
        final = stream.get_final_message()

    # Append-only history for the API; the app's own record keeps only `text`.
    history += [user, {"role": "assistant", "content": final.content}]
    print("cache read:", final.usage.cache_read_input_tokens)
    return reply


history: list = []
send(history, "What's on my calendar today?", dictated=True, spoken=True)
send(history, "Put that in a table for me.", dictated=False, spoken=False)
# The second call reads SYSTEM from the cache even though voice switched off:
# same system bytes, same effort, and the voice rules live only in the turn.

External links

Exercise

Take any app you have that sends a system prompt and list every value inside it that can change during one conversation: dates, user settings, feature flags, mode switches. For each, decide where it moves: the turn prefix, a tool result, or nowhere because it was never needed. Then write the test: build the system prompt twice for the same conversation under different per-turn settings and assert the bytes are equal.
Hint
The usual offenders are a timestamp at the top, a user preference rendered into prose, and a feature flag that toggles a paragraph. The test is three lines: render, render again with the per-turn settings changed, assert equality. If it fails, the diff tells you exactly which value leaked upward. Then look past the prompt: a per-turn effort, a thinking toggle or a changing tool list breaks the cache without touching a byte of the system prompt.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.