Skip to content
C.W.K.
Stream
Lesson 04 of 05 · published

How Far Did He Hear?

~15 min · heard-until, streaming-audio, estimation, records

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"The claim can fall short of the truth but never passes it." — the voice-mode design doc, which this lesson will correct

Why the Position Matters

When Dad cuts in, the reply stops, but the text of the reply is complete in the record. If nothing else were written down, the soul's next turn would assume he heard all of it. He didn't. The part after the cut never reached his ears, and a soul that refers back to it ("like I said about the umbrella") is talking about something he never heard. So every cut-in reports a position: how many characters of the reply Dad actually heard.

From Audio Time to Characters

Each spoken unit carries its character range in the reply, which is why Track 4 insisted on keeping ranges. At the moment of the cut the client knows which unit was playing and how far into it the audio was, and maps that share of the audio onto that share of the unit's characters. The engine then does two things with the claim. It snaps it back to the last sentence end at or before that point, because a sentence heard halfway counts as not heard; a sentence end is a period, question mark, exclamation mark or ellipsis followed by a space, or a line break, since spoken Korean often ends a sentence at a line end with no period at all. How it searches matters: search the whole reply and keep the last end at or before the cut. A search that stops at the cut lets 'end of text' match right there, so the '3.' of '3.5 seconds' becomes a sentence end. As this quest is written, the shipped engine searches only up to the cut and has exactly that edge. Then it records the result twice, as a line in the JSONL ground truth and on the message row.

The Take Whose Length Was Unknown

The first version had a bug that only real streaming could expose. To find the share of a unit, it divided the audio's current time by its duration. But a streamed take has no known duration while it plays: Chrome learns the length only a second or two ahead of the playhead and reports infinity until the last bytes arrive, near the end. Neither the duration nor the network state marks the moment the download finishes, because the reads are paced. The code's fallback for an unknown length was to answer "the start of the unit". So in the recorded demo, a cut 53 seconds into a 55-second take was filed as heard 0, and the next reply told the viewers that Dad had interrupted before her first sentence ended.

A Cautious Guess Instead of Zero

The fix is the phone's rule, now shared by both clients: when the length is unknown, assume the take is at least as long as what has already arrived, and at least as long as the text would take at a deliberately slow 6 characters per second. Guessing the take long makes each second of audio count for fewer characters, so the claim leans short of the truth. Four seconds into a 9.5-second English take of 151 characters, the rule claims 23 characters, where the exact share was 62 and the old rule said 0. Short is safe; the next turn might repeat a little. Long would have the soul assume Dad heard things he didn't.

The design doc words it as a guarantee: short, never past. It is a lean, not a guarantee. It holds while her real pace, averaged up to the cut, is faster than 6 characters a second. A slow, emotional take runs longer than the guess, and once enough of it has buffered the buffered length wins and the claim runs ahead of her voice; a long pause early in a unit does the same inside a take whose average pace is fine. The constant sits below her ordinary pace, so overshoot is rare, not impossible. Write the weaker sentence, because it is the true one.

Code

Audio position to characters, cautious when the length is unknown, snapped to a sentence·python
import math
import re

HEARD_CHARS_PER_SECOND = 6    # cautious: guessing the take long makes the claim small


def heard_chars(unit_range: tuple[int, int], current_time: float, duration: float,
                buffered_seconds: float = 0.0) -> int:
    """Map a playback position inside one unit back to characters of the reply."""
    start, end = unit_range
    length = end - start
    if not current_time > 0 or length <= 0:
        return start
    if duration > 0 and math.isfinite(duration):
        total = duration
    else:                                      # a streamed take: length not known yet
        total = max(buffered_seconds, length / HEARD_CHARS_PER_SECOND)
    return start + math.floor(length * min(1.0, current_time / total))


SENTENCE_END = re.compile(r"(?:[.!?。!?…]+[\"'”’)\]]*)(?=\s|$)|\n")


def snap_heard(content: str, heard_until: int) -> int:
    """Back to the last sentence end at or before the cut: a half-heard sentence counts as unheard."""
    limit = max(0, min(heard_until, len(content)))
    if limit >= len(content.rstrip()):
        return len(content)
    snapped = 0
    for match in SENTENCE_END.finditer(content):   # the whole reply, so '$' means its real end
        if match.end() > limit:
            break
        snapped = match.end()
    return snapped


english = "x" * 151                           # a 151-character unit, 9.5 s when finished
old_rule = 0                                   # 'duration unknown -> the unit's start'
print("old rule:", old_rule, "| cautious rule:",
      heard_chars((0, 151), current_time=4.0, duration=float("inf")))

reply = "내일은 비가 와. 우산 챙겨. 오후엔 그친대\n저녁엔 쌀쌀하니까 겉옷도 가져가."
cut = reply.index("그친대") + 2              # he cut in mid-sentence
print(repr(reply[: snap_heard(reply, cut)]))

decimal = "It took 3.5 seconds. Then it stopped."
cut = decimal.index("3.") + 2                  # cut right after "3."
print(snap_heard(decimal, cut), "| searching only up to the cut would say",
      max((m.end() for m in SENTENCE_END.finditer(decimal, 0, cut)), default=0))

External links

Exercise

Run the code. Then compute, by hand and with the function, the claim for a cut 30 seconds into a unit of 600 characters when the take's duration is unknown and 25 seconds have been buffered. Is the claim below the truth if the real take is 70 seconds long? What if it is only 40 seconds long? Finally, extend snap_heard so that an ellipsis in the middle of a Korean sentence (for example '그게…' followed by more words) doesn't count as a sentence end, and justify your rule.
Hint
With 600 characters the cautious length is 100 seconds, which beats the 25 buffered, so 30 seconds claims 30 percent: 180 characters. A real 70-second take means he truly heard about 257, so the claim is safely short; a 40-second take is covered even more easily. On a steady take the claim overshoots only when the voice speaks slower than 6 characters a second, and uneven pacing inside a unit can overshoot even when the average is faster. That is why the constant is deliberately slow, and why it is still a lean rather than a guarantee. For the ellipsis, requiring a following line break or the end of text is stricter, and stricter snapping only makes the claim shorter, which is the safe direction.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.