Skip to content
C.W.K.
Stream
Lesson 01 of 05 · published

Speech From the First Token

~14 min · prompting, spoken-reply, runtime-prompts, korean-speech

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Dad will hear this reply, not read it." — the first line of the spoken-turn instruction

About the Medium, Never the Persona

The spoken-reply instruction is short, and what it leaves out matters as much as what it says. It never tells the soul who to be. Pippa stays Pippa, Ttori stays a cheeky kid, Feynman stays Feynman, each with their own vault voice and their own language. The instruction is only about the medium: this turn will be synthesized in your own voice and played aloud, so for this turn these rules replace the readability guidance in your system prompt, and everything else about you stays exactly as it is.

The Rules, and Why Each One Exists

  • Lead with the answer in a sentence or two, then leave room. Offer more ("want to hear more?") instead of unloading everything. A listener can't skim; a long preamble is a long wait.
  • One idea at a time. Weave at most about three items into a sentence and group anything longer. Nobody remembers the fifth item of a spoken list.
  • If the ask is ambiguous, ask back in one short line. A wrong guess costs a whole spoken answer.
  • No length limit. If Dad asks for a long explanation or a story, talk as long as it needs, as speech and not as a document. The house never caps a soul's length pre-emptively; short is the default shape, not a ceiling.
  • Nothing that only works on a screen. No headings, bullets, numbered lists, tables, bold, code, links, URLs or file paths. When Dad needs to see something, it goes into a card, the subject of the third lesson.

Rules for Korean Speech

Most of what makes Korean speech sound machine-made comes from the text, not the voice. So the instruction carries two Korean-specific rules. Write loanwords in Hangul, the way they are said, and use English only as whole sentences: a voice that meets English words in the middle of a Korean sentence reads them with the wrong accent. Write a quantity that carries a unit in digits, because the voice engine now reads each one in the numeral its unit takes (Track 4 tells that story). The design doc names a third rule that, as this quest is written, has not reached the shipped prompt: where a soul would address someone as 네 (your), write the spoken form or use a name, because in speech that syllable is heard as 내 (my). Pippa never uses that word for Dad in the first place, which is perhaps why nobody has missed it yet; a younger soul talking to a friend would.

Prompts Are Data

All four voice instructions live in the same place as every other runtime prompt the house uses, editable from Admin without a deploy. They are templates: the spoken one receives the approved list of audio tags when it is rendered, so the prompt can never name a tag that list lacks. That list is cwkPippa's copy of the one the voice engine strips, though, and the engine keeps the original. Rendering closed the gap between the prompt and the copy; the gap between the copy and the original is still guarded only by care, and a check that compares them is the missing piece. Dad can tune a word of the wording tonight and hear the difference on the next turn.

What Going Live Sounded Like

Three live turns on 2026-09-25 checked the shape end to end. The replies contained no Markdown at all. Loanwords came out in Hangul. The soul used her own [happy], [sighs] and [softly] tags. And each reply ended by handing the turn back, one of them with "Want to hear more? Or will you tell me about your day first?" That last habit is the one a document never has.

Code

Runtime prompts as data, rendered per turn·python
import string

APPROVED_TAGS = ("[laughs]", "[sighs]", "[softly]", "[warmly]", "[whispers]",
                 "[curious]", "[thoughtful]", "[pause]")

# Stored like any other runtime prompt: editable in Admin, no deploy needed.
PROMPTS = {
    "spoken_reply_instruction": (
        "[Spoken turn]\n"
        "Dad will hear this reply, not read it. For this turn only, these rules "
        "replace the readability guidance in your system prompt. Everything else "
        "about you stays exactly as it is.\n"
        "- Lead with the answer in a sentence or two, then leave room for Dad.\n"
        "- One idea at a time; group anything longer than about three items.\n"
        "- There is no length limit: talk as long as it needs, as speech.\n"
        "- Nothing that only works on a screen; put it in a card instead.\n"
        "- You may shape your voice with audio tags, sparingly: $audio_tags.\n"
        "- End with your emotion tag as usual."
    ),
}


def render(key: str, **values: str) -> str:
    """Fail loudly on a missing placeholder instead of speaking '$audio_tags'."""
    return string.Template(PROMPTS[key]).substitute(**values)


block = render("spoken_reply_instruction", audio_tags=", ".join(APPROVED_TAGS))
print(block)

try:
    render("spoken_reply_instruction")          # forgot the tag list
except KeyError as missing:
    print("refused to render without", missing)

External links

Exercise

Write your own spoken-turn instruction for an assistant you use, in no more than eight lines. It must not describe the assistant's personality at all, only the medium. Then take three questions you often ask and compare a written answer with a spoken one produced under your instruction, reading both aloud.
Hint
If your instruction says anything like 'be friendly' or 'be concise', delete it: that is persona and length, not medium. What survives should be about the listener's situation: they can't scroll back, can't see symbols, and can interrupt. The best spoken answers usually end with a question that hands the turn back.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.