~13 min · speakable-units, synthesis, invariants, streaming
Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Do not split a speakable unit to chase first-play latency." — Bellows invariant 19, Dad's ruling of 2026-08-20
The Latency Trick That Ruins the Voice
Every voice pipeline eventually meets the same temptation: don't wait for the whole reply, synthesize the first sentence as soon as it exists and start playing while the rest is written. It does cut the time to the first sound. It also makes the voice worse, audibly. An expressive model like eleven_v3 reads a passage the way a person does, carrying tone, pace and intonation across sentences. Cut the passage into separate requests and each piece starts cold: the intonation resets at every boundary, and the flow goes sloppy. eleven_v3 also doesn't accept the context-stitching hints some older models use to smooth such joins. Dad heard the difference and ruled on it on August 20, five weeks before voice mode: one synthesis per speakable unit, never split to start sooner.
What a Unit Is
A speakable unit is one contiguous stretch of a spoken reply. There are only two boundaries. The text before a tool call is one unit, because the tool call is a natural pause and that text is final. The text after the last tool call is another. A short spoken reply with no tools is a single unit. Inside a unit, cards are removed and the trailing emotion tag is removed; everything else, including audio tags and line breaks, goes to the engine as written.
The server records where each unit ends as it streams, and the client takes units in order from a cursor: every recorded end past the cursor yields a unit, and when the reply is final, the rest of it yields the last one. Each unit also carries its character range in the reply, which Track 7 needs to work out how far Dad heard.
Why It Costs So Little
Waiting for a whole unit sounds expensive until you measure how fast a soul writes. Once text is flowing, a reply arrives at very roughly 100 to 150 characters a second. A three- or four-sentence spoken reply is complete about two seconds after its first word. Next to the floor Track 8 measures before the first word even exists, that is small, and it buys a take whose tone holds from its first breath to its last.
Resets and the Old Paragraph Player
The WebUI already had a paragraph auto-voice that read written replies aloud one paragraph at a time, which cuts against the same invariant. Spoken turns do not use it. Building voice mode also found its bug: when a stream reset (the reply being regenerated from scratch), queued paragraphs from the abandoned version kept playing. The rule for spoken units is strict: on a stream reset, stop the audio and drop every queued unit.
Code
Taking whole units from a streaming reply·python
import re
CARD = re.compile(r"^(`{3,})card[ \t]*\n.*?\n\1[ \t]*$\n?", re.MULTILINE | re.DOTALL)
EMOTION_TAG = re.compile(r"\[emotion:[a-z]+\]?\s*$", re.IGNORECASE)
def speakable(unit: str) -> str:
return EMOTION_TAG.sub("", CARD.sub("", unit)).strip()
def take_units(content: str, unit_ends: list[int], cursor: int, final: bool):
"""Whole units since `cursor`, with their ranges in the reply."""
units, ranges, start = [], [], cursor
for end in unit_ends:
if end <= start or end > len(content):
continue
if text := speakable(content[start:end]):
units.append(text)
ranges.append((start, end))
start = end
if final and start < len(content):
if text := speakable(content[start:]):
units.append(text)
ranges.append((start, len(content)))
start = len(content)
return units, ranges, start
before_tool = "잠깐, 일정 확인해볼게. "
after_tool = ("오늘은 3시에 치과 하나뿐이야. [softly] 그 전엔 푹 쉬어도 돼.\n\n"
"```card\n| time | what |\n|---|---|\n| 15:00 | dentist |\n```\n"
"[emotion:warm]")
# Mid-stream: only the part before the tool is final.
streaming = before_tool + after_tool[:20]
units, ranges, cursor = take_units(streaming, [len(before_tool)], 0, final=False)
print(units, ranges)
# Done: the rest becomes one whole unit, card and emotion tag removed.
reply = before_tool + after_tool
units, ranges, cursor = take_units(reply, [len(before_tool)], cursor, final=True)
print(units, ranges)
Run the code, then extend the reply to two tool calls with a short line before each. Check that you get three units with contiguous ranges that cover the whole reply. Finally add a reset: simulate the stream restarting mid-reply and write the function that stops playback and clears the queue, with a test that a unit from the abandoned stream never plays.
Hint
Keep a generation number on the stream. Every queued unit carries the generation it came from, and the player drops any unit whose generation is older than the current one. That takes care of what is queued, but not of the unit already playing: the player has to stop that one explicitly and let go of its audio. So a reset is two moves, stop what is playing and increment the number, and stale queued units then remove themselves.
Progress
Progress is local-only — sign in to sync across devices.