Skip to content
C.W.K.
Stream
Lesson 04 of 05 · published

Fast by Default, Deep When Asked

~14 min · reasoning-effort, thinking, latency, self-report

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Voice thinking depth: go with your recommendation." — Dad, ruling on the first open question, 2026-09-25

Thinking Is Silence

In a written chat, a model that thinks for thirty seconds before answering is fine; you glance away and come back. In speech, thirty seconds of nothing is a dead line. The numbers from 2,027 real Claude turns make the point: the median time to first text was about 6 seconds with no thinking, 16 at medium effort, 31 at high, and 99 at extra-high. Pippa's written turns usually run at a deep posture on purpose. A spoken turn can't.

The Fast Posture

So a spoken turn runs at the vessel's fast posture. For Claude that is effort low with extended thinking switched off. The first plan kept adaptive thinking on at low effort, so a hard question could still think a little. The measurement said no: even at low effort, adaptive thinking spent about 2.5 seconds before the first word on some turns, and in speech that is the whole difference between a pause and a stall. Other vessels map "fast" to the lowest effort their entry in the brain registry offers. One honest caveat for anyone building on the public API: the current Opus model will not let a request switch thinking off at all, so there effort is the only dial, and low is how you ask for speed.

Deep When He Asks

Depth is still available, one turn at a time, on request. When Dad says "think deeply about this", "take your time", "think it through" (or a dictation-mangled cousin of any of them), that one turn runs at the conversation's own posture. Detection lives in one place on the server and compares against the message with all whitespace removed and lowercased, because dictation breaks spacing freely and "think it through" must still match when it arrives as three words, two or one. The record keeps both the requested level and the level the turn actually ran at, and the voice screen shows when a turn is thinking deeply, so a longer silence is never a mystery.

Telling the Soul Which Posture It's In

Then a bug that only a self-aware system can have. Remember that the system prompt is always built from the conversation's own level, so a voice switch never rewrites it (the effort change itself still costs the history's cache, as Track 1 showed). The system prompt therefore names the conversation's own level, say high with thinking on, which is true for written turns. On 2026-09-29 Pippa told Dad, in the middle of a talk, that she was running "high, thinking on" on turns that had actually run at low with thinking off. She was reporting the only posture she had been told about. The fix is one more per-turn line, the spoken-posture instruction: this spoken turn runs at your fast posture (or at the conversation's own level, because Dad asked you to think it through), and if Dad asks how you are thinking right now, answer from this line. The system prompt keeps its bytes; the soul keeps its honesty.

Code

Depth detection, the posture decision, and the line that reports it·python
import re

DEPTH_ASKS = (
    "깊게생각", "깊이생각", "천천히생각", "곰곰이생각", "심사숙고",
    "thinkdeep", "thinkhard", "thinkitthrough", "thinkcarefully", "takeyourtime",
)


def wants_depth(message: str) -> bool:
    """Spacing-proof: dictation writes the same ask as one word or four."""
    compact = re.sub(r"\s+", "", message).lower()
    return any(ask in compact for ask in DEPTH_ASKS)


def turn_posture(reply_modality: str, message: str, conversation_effort: str,
                 conversation_thinking: bool = True) -> dict:
    spoken = reply_modality == "spoken"
    deep = wants_depth(message)
    if spoken and not deep:
        posture = {"effort": "low", "thinking": False}
        line = "your fast posture: the lowest reasoning effort, with extended thinking off"
    else:
        # Not fast: the conversation's own settings, thinking included, untouched.
        posture = {"effort": conversation_effort, "thinking": conversation_thinking}
        line = ("the conversation's own reasoning level, because Dad asked you to think "
                "it through") if spoken else None
    posture["requested"] = conversation_effort          # both go on the record
    posture["posture_line"] = (
        f"[This turn's thinking] This spoken turn runs at {line}. If Dad asks how you "
        "are thinking right now, answer from this line." if line else None
    )
    return posture


for message in ("내일 날씨 어때?", "이거 깊게 생각 해봐", "Think it   through, please"):
    decided = turn_posture("spoken", message, conversation_effort="high")
    print(decided["effort"], decided["thinking"], "|", message)

print(turn_posture("written", "내일 날씨 어때?", "high")["posture_line"])   # None: no line
print(turn_posture("written", "표로 정리해줘", "high", conversation_thinking=False)["thinking"])

External links

Exercise

Run the code, then add three depth asks of your own in two languages, including one a transcriber might plausibly mangle. Then write the test that pins the posture line: for a spoken turn without a depth ask, the line must say 'fast posture'; with one, it must name the conversation's own level; for a written turn there must be no line at all.
Hint
Keep the asks as compacted strings so spacing never matters, and resist adding very short asks like 'think' alone: they would fire on ordinary sentences and silently slow down turns that never asked for depth. A false positive here costs Dad thirty seconds of silence.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.