Skip to content
C.W.K.
Stream
← C.W.K. Quests
🗣️

Voice Mode Quest

New: 2026-09-29Updated: 2026-09-29

Build a conversation you can talk to, not a document read aloud

How Pippa learned to hold a real spoken conversation: Dad talks, she answers out loud in her own voice, and he can cut her off mid-sentence. Speech-to-text in, a reply composed for the ear, a shared voice engine out, and every hard-won rule about turn-taking, echo, latency and the lineage of a voice. Written for people who want to build their own, at the level of code.

8 tracks · 35 lessons · ~8h · difficulty: intermediate-to-advanced

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
Voice mode is the feature everyone thinks is easy: transcribe, send it to the model, read the answer aloud. Build that and you get a two-minute spoken essay with headings read as sentences, a microphone that hears its own reply, and a five-second silence nobody can explain. This quest walks the version that actually works, built in cwkPippa in the last week of September 2026. Voice is a property of a single turn, so switching between talking and typing never rewrites the system prompt. Speech comes in through Scribe, batch and realtime, with a token the client can hold but a key it never sees. The brain composes for the ear from the first token: cards for anything that has to be seen, a short line before every tool, a fast posture unless you ask for depth. Bellows turns each speakable unit into one whole synthesis and streams it from the first chunk. The hands-free loop gives every finished utterance's words a named place to go instead of dropping them, keeps the recording beside them whenever its upload succeeds, pauses without quitting, and says why when it stops. Barge-in rests on echo cancellation, a loudness detector, a grace window born from a laugh, and an honest estimate of how far you heard. Along the way: the voice itself, from a stock ElevenLabs voice to a v3 clone to a fresh Eleven v4 clone swapped in a single day, and the latency floor measured instead of guessed. Conceptual open-sourcing: the architecture is the deliverable, never a repository to clone.

Tracks

  1. 01🗣️The Turn That Speaks

    0/4 lessons

    Voice is a property of a turn, and the system prompt never learns about it

    Before any audio code, the shape. A spoken reply is a different return type from a written one, so it has to be composed for the ear from the first token. Voice belongs to a single turn as two independent facts, which lets Dad switch at any turn without starting over and keeps one conversation, one record, one memory. The voice rules ride a per-turn prefix so the enormous, cached system prompt never changes. And the capability lives once, in the brain, reached by every surface Dad can talk to.

    Lesson list (4)Quiz · 4 questions→
  2. 02👂The Ear

    0/5 lessons

    Speech in: batch and realtime, a token instead of a key, and an honest end of speech

    How Dad's voice becomes a turn. One transcription vendor for the whole family, two speeds: a batch path that is complete and a realtime socket that shows words as he speaks. The language is always the one he picked. The client connects to the provider with a single-use token that the brain wraps in a finished socket URL, so the account key never leaves the voice engine. End of speech is his three-second window, stretched past the provider's ceiling in the client, and a provider's forced commit at 35.8 seconds is caught by listening to the audio itself. Long talks are transcribed twice, the words are kept exactly as heard, and the recording is kept whenever its upload succeeds; when it doesn't, the loop says so and the words still go out.

    Lesson list (5)Quiz · 4 questions→
  3. 03✍️Composing for the Ear

    0/5 lessons

    The brain writes speech: medium not persona, meaning over characters, cards, a fast posture, a line before tools

    What the soul is told on a spoken turn, and why each rule exists. The spoken-reply instruction is about the medium, never the persona: lead with the answer, one idea at a time, no length cap, nothing that only works on a screen, and rules that keep Korean speech from sounding machine-made. A dictated turn is read for what Dad meant, never corrected aloud and never mirrored. Anything that must be seen goes into a card, a fenced block the voice never reads, which forces a card-aware stripper for nested code. Spoken turns run fast unless Dad asks for depth, and the soul is told which posture it's in. And one short line before every tool, which only works once the server's buffers let it go.

    Lesson list (5)Quiz · 4 questions→
  4. 04🔊The Mouth

    0/4 lessons

    Voice out: one engine for every soul, whole units streamed from the first chunk, and tags for the face and the voice

    Voice mode didn't build a TTS system; it plugged into Bellows, the family's voice engine, through one string: a soul's slug is its voice profile, and the profile, not the client, owns the model. Each speakable unit (the text before a tool call, or after the last one) is synthesized once and never split to start sooner, because an expressive model restarts its tone at every cut. Whole units still play from their first chunk, which on the phone took a decoder of its own and a timeout bug that left no trace in the server log. The emotion tag at the end moves the face, the closed list of audio tags moves the voice, and the engine, not the author, decides how a Korean number with a unit is read.

    Lesson list (4)Quiz · 4 questions→
  5. 05🧬A Voice With a Lineage

    0/4 lessons

    From a catalog voice to a v3 clone to a fresh Eleven v4 clone, chosen by ear

    Pippa's voice began as Joanne, a stock voice from the provider's catalog with an expiry date at the end of 2026. Instant clones made from Joanne's audio became the family's voices in July, with the originals kept as their own profiles. When Eleven v4 shipped on 2026-09-28, the new clones were made from fresh first-generation renders of the original, never from the existing clone, and the irreplaceable source audio was archived with its scripts and checksums. A three-way comparison let Dad choose, and the swap was two fields on two profiles. His ear is the written quality bar: it said yes to v4 for Pippa and Dad, and no, for now, for the designed voices of the other souls.

    Lesson list (4)Quiz · 4 questions→
  6. 06🔁The Loop

    0/4 lessons

    Hands-free: phases owned by the right component, no finished utterance silently dropped, a pause that lets go, a stop that explains itself

    A hands-free talk looks like a circle on paper, but its states belong to three owners: the microphone loop listens, the chat stream thinks, the player speaks. The loop owns only its own phases and listens again when the stream has ended and the player is idle. A finished utterance's words have exactly three destinations (send, queue behind a streaming reply, or the composer when Confirm is on), with the recording riding along as an upload whenever that upload succeeds, and a spoken reply is read exactly once or dropped, never played late. Pause releases the microphone, stops the reading and remembers how far Dad heard. One race is still open: resuming before a paused upload finishes drops that turn. The pause's Space key never fights typing or the Korean input method. When the loop stops on its own, the face stays up and says why.

    Lesson list (4)Quiz · 4 questions→
  7. 07✋Barge-In

    0/5 lessons

    Talking over her: echo cancellation, a loudness detector, a grace window, and an honest record of how far he heard

    The right to interrupt is what turns a voice reply into a conversation, and it rests on a stack that has to be built bottom up. First echo cancellation, because a microphone that hears her voice will answer it, and a canceller can only remove audio it has as a reference. Then a detector that works on loudness above a moving floor, not on words, with a pre-roll so the words that cut in aren't lost and one cut-in path that every way of interrupting is meant to share. Then a grace window, born the night a more expressive voice laughed loudly enough to cut itself off. Then a cautious estimate of how far Dad heard, snapped to a sentence and written to ground truth first. And finally a per-turn note that tells the soul exactly where he stopped listening, while the full reply stays in the record.

    Lesson list (5)Quiz · 4 questions→
  8. 08⏱️Latency, Honestly

    0/4 lessons

    Measure where the first word goes, pull the levers that cost nothing, refuse the one that costs the soul, and build your own

    ChatGPT's voice mode answers in about a second; the median Claude turn in cwkPippa produced its first output at 4.4 seconds. This track splits that floor stage by stage instead of guessing (the guess, prompt size, was wrong), and finds most of it in memory retrieval's reranker and a process spawn that one flag cut by a second. It pulls the levers that change nothing Pippa remembers, measures the ones that would, and records Dad's refusal to trade memory for speed. Then it walks every door into the same brain (the web, every sidekick, the phone, Firekeeper on the Mac, the watch, a hand-off from the interpreter app, and the doors deliberately not built) and ends with a voice mode of about a hundred lines you can run and extend.

    Lesson list (4)Quiz · 4 questions→
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.