"The claim can fall short of the truth but never passes it." — the voice-mode design doc, which this lesson will correct
Why the Position Matters
When Dad cuts in, the reply stops, but the text of the reply is complete in the record. If nothing else were written down, the soul's next turn would assume he heard all of it. He didn't. The part after the cut never reached his ears, and a soul that refers back to it ("like I said about the umbrella") is talking about something he never heard. So every cut-in reports a position: how many characters of the reply Dad actually heard.
From Audio Time to Characters
Each spoken unit carries its character range in the reply, which is why Track 4 insisted on keeping ranges. At the moment of the cut the client knows which unit was playing and how far into it the audio was, and maps that share of the audio onto that share of the unit's characters. The engine then does two things with the claim. It snaps it back to the last sentence end at or before that point, because a sentence heard halfway counts as not heard; a sentence end is a period, question mark, exclamation mark or ellipsis followed by a space, or a line break, since spoken Korean often ends a sentence at a line end with no period at all. How it searches matters: search the whole reply and keep the last end at or before the cut. A search that stops at the cut lets 'end of text' match right there, so the '3.' of '3.5 seconds' becomes a sentence end. As this quest is written, the shipped engine searches only up to the cut and has exactly that edge. Then it records the result twice, as a line in the JSONL ground truth and on the message row.
The Take Whose Length Was Unknown
The first version had a bug that only real streaming could expose. To find the share of a unit, it divided the audio's current time by its duration. But a streamed take has no known duration while it plays: Chrome learns the length only a second or two ahead of the playhead and reports infinity until the last bytes arrive, near the end. Neither the duration nor the network state marks the moment the download finishes, because the reads are paced. The code's fallback for an unknown length was to answer "the start of the unit". So in the recorded demo, a cut 53 seconds into a 55-second take was filed as heard 0, and the next reply told the viewers that Dad had interrupted before her first sentence ended.
A Cautious Guess Instead of Zero
The fix is the phone's rule, now shared by both clients: when the length is unknown, assume the take is at least as long as what has already arrived, and at least as long as the text would take at a deliberately slow 6 characters per second. Guessing the take long makes each second of audio count for fewer characters, so the claim leans short of the truth. Four seconds into a 9.5-second English take of 151 characters, the rule claims 23 characters, where the exact share was 62 and the old rule said 0. Short is safe; the next turn might repeat a little. Long would have the soul assume Dad heard things he didn't.
The design doc words it as a guarantee: short, never past. It is a lean, not a guarantee. It holds while her real pace, averaged up to the cut, is faster than 6 characters a second. A slow, emotional take runs longer than the guess, and once enough of it has buffered the buffered length wins and the claim runs ahead of her voice; a long pause early in a unit does the same inside a take whose average pace is fine. The constant sits below her ordinary pace, so overshoot is rare, not impossible. Write the weaker sentence, because it is the true one.