Skip to content
C.W.K.
Stream
Lesson 01 of 04 · published

Not a Document Read Aloud

~12 min · voice-mode, product-contract, tts, spoken-reply

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"When this option is on, Pippa has to answer in a way that suits TTS. If she makes bullets and tables, answers that need Markdown, the conversation gets hard, and it doesn't suit the voice model either. We need ChatGPT voice mode quality." — Dad, 2026-09-25

The Version Everyone Builds First

Here is the voice mode that takes an afternoon: record the microphone, send the audio to a transcriber, send the transcript to the chat model, strip the Markdown out of its answer, and hand the rest to a text-to-speech engine. Every piece works. The result is unbearable. The model was told, in its system prompt, that its reader has a screen, so it writes for a screen: a heading, three bold labels, a bulleted list of five options, a table, a link. The stripper deletes the symbols and leaves the words, so the voice reads "Options. Option one colon speed. Option two colon cost" in one flat breath. A URL comes out as a string of letters. A tool call leaves ten seconds of dead air that sounds exactly like a dropped call. And the transcript it answered had three misheard words that the model politely pointed out.

A Different Return Type

Firekeeper's plan already named the trap back in July: dictation cleanup and a spoken reply are different contracts. Dictation returns text you will read; a spoken reply returns words you will hear. A spoken answer leads with the point, holds one idea at a time, hands the turn back, and never contains anything that only works on a screen. You cannot get that by post-processing a written answer, because the structure is wrong, not the punctuation. A list of five things does not become speech by removing its bullets; it becomes speech when the speaker decides to name two and offer the rest.

So the rule that everything in this quest follows is: compose for the ear from the first token. The model knows, before it writes anything, that this turn will be heard. It is never asked to write a document and then have the document sanded down.

Where the Feature Lives

Dad's second ruling that day settled ownership. Firekeeper, the family's dictation app, was always meant to grow into voice conversation. But the conversation needs the soul: her memory, her vault, her judgment, the same thread Dad reads later on his phone. So the capability lives in cwkPippa, the brain, for every soul and every surface. Firekeeper became one door into it, beside the WebUI, every sidekick panel, the phone and the watch. The apps carry only a loop that listens and plays; none of them grows a second brain.

The Map

The quest follows one spoken turn end to end. Track 1 is the shape: why voice belongs to a turn. Track 2 is the ear: speech-to-text, batch and realtime. Track 3 is the brain writing for the ear. Track 4 is the mouth: one voice engine for the whole family. Track 5 is the voice itself and where it came from. Track 6 is the hands-free loop. Track 7 is barge-in, talking over her. Track 8 is latency, measured honestly, and every door that uses all of it.

Code

Why stripping Markdown is not composing for the ear·python
import re

WRITTEN_REPLY = """## Your options

| Option | Speed | Cost |
|---|---|---|
| Batch | slow | low |
| Realtime | fast | per second |

See https://example.com/pricing for details.
"""


def strip_markdown(text: str) -> str:
    """The afternoon version: delete the symbols, keep the words."""
    text = re.sub(r"^#+\s*", "", text, flags=re.MULTILINE)   # headings
    text = re.sub(r"^\|?[-| ]+\|?$", "", text, flags=re.MULTILINE)  # table rules
    text = text.replace("|", " ")                               # table cells
    return re.sub(r"\s+", " ", text).strip()


print(strip_markdown(WRITTEN_REPLY))
# Your options Option Speed Cost Batch slow low Realtime fast per second See
# https://example.com/pricing for details.   <- one breath, all of it
#
# Every word survived and none of it is speech. The fix is not a better
# regex; it is a reply that was never a table in the first place:

SPOKEN_REPLY = (
    "Two ways, really. Batch is cheaper but you wait for the whole take; "
    "realtime hears you as you go and bills by the second. "
    "Want me to put the prices on your screen?"
)
print(SPOKEN_REPLY)

External links

Exercise

Take three real answers from any chat assistant: one with a list, one with a table, one with a link. Read each one aloud exactly as written, then rewrite it as you would actually say it to a friend across the table. Write down every structural change you made, not just the symbols you dropped.
Hint
You will find you did not remove bullets so much as choose: you named two items and offered the rest, you said where the link goes instead of reading it, you led with the answer the table was hiding. Those choices are the spoken contract, and only the author of the reply can make them.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.