Skip to content
C.W.K.
Stream
Lesson 04 of 05 · published

A Commit Is Not the End of Speech

~15 min · forced-commit, debugging, measurement, batch-retranscription

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"There's a small bug in the cut-in feature. It tends to cut me off anywhere." — Dad, 2026-09-27

The Wrong Diagnosis First

In a recorded demo, Dad's words went out mid-sentence twice, each ending on a half-phrase. The first report, filed by another Pippa instance, called it a false barge-in: surely the soul's voice had been cut by noise. The logs disagreed. Every reply had played to its full length, and both "barge-ins" were Dad's own voice. The cut was on the way in, not on the way out.

35.84 Seconds

Both clipped recordings were exactly 35.84 seconds long, and in both Dad was still talking in the last chunk. That coincidence is the whole clue. Replaying the two recordings joined together through the same realtime socket showed it plainly: commits at 35.81 and 71.56 seconds in the middle of speech, then the ordinary silence commit 3.36 seconds after the speech really ended, with the next partial arriving 1.0 to 1.1 seconds after each forced commit. Scribe's realtime model commits on its own after roughly 35.8 seconds of continuous audio, whether or not anyone paused. The loop had treated every commit as the end of speech, so at the 3-second default each forced commit sent the turn immediately and the rest of Dad's sentence was lost.

Fix One: Was It Quiet Before the Commit?

A silence commit arrives after the provider's silence window of quiet; a forced commit arrives while someone is talking. The audio can tell them apart. Take the utterance's quietest tenth as the floor (the room between words). A chunk above three times the floor and above 300 RMS is speech. If a third or more of the last provider-silence window was loud, he is still talking: keep the committed piece and keep listening. Replayed over four days of Dad's 76 recordings, every normal ending was quiet (at most a few loud chunks out of thirty) and all three 35.84-second cuts were loud (16 to 20 of 23).

Fix Two: Don't Wait on Words Alone

The first version then waited for the next words to arrive. Dad's retest broke it: he counted aloud for 75 seconds, and after the second forced commit the transcriber's next words came 4.9 seconds late, so a 4-second wait sent the turn while he was still counting. Now the loop waits the window plus a second, then asks its own ears, and sends only once the microphone has been quiet for the provider's window. Loud with no words for 30 seconds goes out too, because a room's noise is not a talk.

Fix Three: The Realtime Model Can Drop Words

His next test ran 132 seconds without a cut, and twenty seconds of numbers in the second minute were simply missing from the transcript. The recording was continuous speech through that stretch, so every chunk had reached the provider; the realtime model just never wrote those words. The batch model wrote every word of the same file in 3.3 seconds: 645 characters against the live 535. So any utterance of 30 seconds or more is now transcribed again, whole, by the batch path before it goes out, in the same call that keeps the recording. If that call fails, the live words go out instead. The cost is a few seconds after a long talk, and the batch call carries no keyterms, so a soul's name can come back spelled differently. Losing twenty seconds of what Dad said is worse.

Code

Did the commit follow silence, or did the provider force it mid-speech?·python
def commit_follows_silence(levels: list[float], chunk_s: float,
                           provider_silence_s: float) -> bool:
    """levels: every chunk's RMS since the socket opened, oldest first."""
    if not levels or chunk_s <= 0 or provider_silence_s <= 0:
        return True                           # nothing heard: the commit stands
    ordered = sorted(levels)
    floor = ordered[int(len(ordered) * 0.1)]  # the quietest tenth: the room
    line = max(floor * 3, 300)                 # above this, a chunk is speech
    covered = loud = 0.0
    for level in reversed(levels):            # walk back over the silence window
        if covered >= provider_silence_s:
            break
        covered += chunk_s
        if level > line:
            loud += chunk_s
    return loud * 3 < covered                 # a third or more loud: still talking


LONG_UTTERANCE_S = 30.0


def final_transcript(live_words: str, seconds: float, batch) -> str:
    """Long talks get a second, whole pass; the live words are the fallback."""
    if seconds < LONG_UTTERANCE_S:
        return live_words
    try:
        return batch() or live_words
    except Exception:
        return live_words


CHUNK = 0.128
ROOM = 150.0


def talk(chunks: int) -> list[float]:
    """Speech with the short dips between words that real talk always has."""
    return [ROOM if i % 5 == 4 else 2400.0 for i in range(chunks)]


natural_end = talk(200) + [ROOM] * 24   # he stopped; ~3 s of room tone
forced_cut = talk(280)                  # ~35.8 s in, still mid-sentence
print(commit_follows_silence(natural_end, CHUNK, 3.0))   # True  -> end of speech
print(commit_follows_silence(forced_cut, CHUNK, 3.0))    # False -> keep listening
print(final_transcript("...스물, 스물하나", 132.0, lambda: "...스물, 스물하나, 스물둘 ... 서른아홉"))

External links

Exercise

Run the code block, then build a harder case: speech that pauses for exactly one second every four seconds for 40 seconds, with a forced commit landing just after one of those pauses. Does the rule call it silence or speech? Adjust the synthetic data until you find the boundary, and write down what that tells you about the one-third threshold.
Hint
The rule lets a commit stand only when less than a third of the last three seconds was loud, which means roughly two seconds of quiet. A one-second pause leaves about two loud seconds in the window, so it reads as speech, well clear of the line; you have to stretch the trailing quiet to about 1.9 seconds (a little under two, because talk's own dips count as quiet) before the commit stands. The threshold was not chosen from theory; it was checked against 76 real recordings where normal endings sat near zero loud and the forced cuts at 16 to 20 loud chunks out of 23. Your boundary case is the reason the replay mattered. Also try the rule on speech with no dips at all: the quietest tenth is then speech, and the rule stops working. Real talk always has gaps between words, and that assumption is worth writing down.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.