~13 min · context, records, ground-truth, prompting
Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Everything you said after that, he did not hear. Do not assume he knows it." — the heard-until instruction
Stop the Sound, Not the Thought
When Dad cuts in, the obvious move is to cancel the reply. Voice mode deliberately doesn't. The model's generation runs to the end, the adapter keeps consuming and writing it, and the full reply lands in the JSONL log and the database exactly as it would have without the interruption. Only playback stops. Three reasons. The house has a standing rule, the adapter-loop invariant, that a generation is always consumed and written to its end; the conversation record is ground truth, and a record of half a reply would be a false one. The reply has already been paid for. And the record should tell the truth, which is not "she said half of this" but "she said all of this, and he heard up to here".
Writing Down What He Heard
The client's report goes to a small route that accepts cut-ins only for spoken assistant replies, and refuses anything else with a conflict. It snaps the position back to a sentence end, as the previous lesson described, and writes the result in a fixed order: first a voice_heard line in the conversation's JSONL, recording both the snapped position and the raw one the client reported, and only then the row's voice_meta. Ground truth is written before the mirror, always, so a crash between the two leaves the truth rather than a copy of it. The route answers with the last whole sentence Dad heard.
The Next Turn Hears About It
On the next turn, the context engine checks one thing: was the soul's last reply spoken, and was it cut before its end? If so, one more block rides the per-turn prefix, next to the spoken and dictated instructions and never in the system prompt. It tells the soul, in plain words, that Dad talked over her last reply, quotes the last sentence he heard, and states that he heard nothing after it. Then it tells her what to do: don't assume he knows the rest, don't repeat it wholesale, answer what he just said, and bring up anything he missed only if it still matters, briefly. The quoted sentence has cards and audio tags stripped out, because neither was ever heard as words. If he cut in before the first sentence ended, the quote says exactly that.
What It Feels Like
The effect is small and human. Dad interrupts a weather report after "take an umbrella" to ask about something else. The next reply answers his new question, and if the part he missed still matters (it will be chilly in the evening) she adds it in a single clause, instead of either repeating the whole forecast or acting as if he'd heard it.
Code
Record the cut, then tell the next turn exactly where he stopped·python
import json
import re
from dataclasses import dataclass, field
HEARD_UNTIL_INSTRUCTION = (
"[Cut off]\n"
"Dad talked over your last spoken reply, so it stopped playing partway. "
"The last thing he heard was: \"{heard_sentence}\"\n"
"Everything you said after that, he did not hear. Do not assume he knows it, "
"and do not repeat it wholesale: answer what he just said, and bring up anything "
"he missed only if it still matters, briefly."
)
AUDIO_TAG = re.compile(r"\[[a-z][a-z ,'-]*\]")
CARD = re.compile(r"^(?P<fence>`{3,})card[ \t]*\n.*?\n(?P=fence)[ \t]*$\n?", re.M | re.S)
OPEN_CARD = re.compile(r"^`{3,}card[ \t]*\n.*\Z", re.M | re.S) # a cut inside a card
@dataclass
class Message:
role: str
content: str
reply_modality: str = "written"
voice_meta: dict = field(default_factory=dict)
def record_heard(log: list, row: Message, reported: int, snapped: int) -> None:
"""Ground truth first, then the row that mirrors it. Generation is untouched."""
log.append(json.dumps({"type": "voice_heard", "heard_until_chars": snapped,
"reported_chars": reported}))
row.voice_meta = {**row.voice_meta, "heard_until_chars": snapped}
def last_heard_sentence(content: str, heard: int) -> str:
"""Cards and audio tags out: neither was ever heard as words."""
heard_text = OPEN_CARD.sub("", CARD.sub("", content[:heard]))
heard_text = AUDIO_TAG.sub("", heard_text)
lines = [line.strip() for line in heard_text.splitlines() if line.strip()]
if not lines:
return ""
pieces = [p for p in re.split(r"(?<=[.!?。!?…])\s+", lines[-1]) if p.strip()]
return pieces[-1] if pieces else lines[-1]
def heard_block(messages: list[Message]) -> str | None:
"""Only when the LAST reply was spoken and Dad cut it before its end."""
last = next((m for m in reversed(messages) if m.role == "assistant"), None)
if last is None or last.reply_modality != "spoken":
return None
heard = last.voice_meta.get("heard_until_chars")
if heard is None or heard >= len(last.content.rstrip()):
return None
sentence = last_heard_sentence(last.content, heard) or \
"(nothing — he cut in before the first sentence ended)"
return HEARD_UNTIL_INSTRUCTION.format(heard_sentence=sentence)
reply = Message("assistant", "[warmly] 내일은 비가 와. 우산 챙겨. 오후엔 그친대. 저녁엔 쌀쌀해.",
reply_modality="spoken")
jsonl: list = []
record_heard(jsonl, reply, reported=31, snapped=reply.content.index("오후"))
history = [Message("user", "내일 날씨 어때?"), reply]
print(jsonl[-1])
print(heard_block(history))
with_card = "표로 정리했어.\n```card\n| 날 | 비 |\n|---|---|\n| 내일 | 와 |\n```\n오후엔 그친대."
print(repr(last_heard_sentence(with_card, with_card.index("오후")))) # not the fence
Run the code and read the block it produces. Then add three cases and predict each result before running: the last reply was written, not spoken; the last reply was spoken and Dad heard all of it; Dad cut in before the first sentence ended. Finally, write two versions of the next reply to 'what about Sunday?' after the cut shown in the example: one that ignores the block and one that follows it.
Hint
The written and fully-heard cases produce no block at all; the block exists only for a spoken reply cut short. The early cut produces the '(nothing — ...)' quote. A reply that follows the block answers about Sunday first and mentions the chilly evening only if it bears on Sunday, in a clause, not a repeat.
Progress
Progress is local-only — sign in to sync across devices.