"There's a small bug in the cut-in feature. It tends to cut me off anywhere." — Dad, 2026-09-27
The Wrong Diagnosis First
In a recorded demo, Dad's words went out mid-sentence twice, each ending on a half-phrase. The first report, filed by another Pippa instance, called it a false barge-in: surely the soul's voice had been cut by noise. The logs disagreed. Every reply had played to its full length, and both "barge-ins" were Dad's own voice. The cut was on the way in, not on the way out.
35.84 Seconds
Both clipped recordings were exactly 35.84 seconds long, and in both Dad was still talking in the last chunk. That coincidence is the whole clue. Replaying the two recordings joined together through the same realtime socket showed it plainly: commits at 35.81 and 71.56 seconds in the middle of speech, then the ordinary silence commit 3.36 seconds after the speech really ended, with the next partial arriving 1.0 to 1.1 seconds after each forced commit. Scribe's realtime model commits on its own after roughly 35.8 seconds of continuous audio, whether or not anyone paused. The loop had treated every commit as the end of speech, so at the 3-second default each forced commit sent the turn immediately and the rest of Dad's sentence was lost.
Fix One: Was It Quiet Before the Commit?
A silence commit arrives after the provider's silence window of quiet; a forced commit arrives while someone is talking. The audio can tell them apart. Take the utterance's quietest tenth as the floor (the room between words). A chunk above three times the floor and above 300 RMS is speech. If a third or more of the last provider-silence window was loud, he is still talking: keep the committed piece and keep listening. Replayed over four days of Dad's 76 recordings, every normal ending was quiet (at most a few loud chunks out of thirty) and all three 35.84-second cuts were loud (16 to 20 of 23).
Fix Two: Don't Wait on Words Alone
The first version then waited for the next words to arrive. Dad's retest broke it: he counted aloud for 75 seconds, and after the second forced commit the transcriber's next words came 4.9 seconds late, so a 4-second wait sent the turn while he was still counting. Now the loop waits the window plus a second, then asks its own ears, and sends only once the microphone has been quiet for the provider's window. Loud with no words for 30 seconds goes out too, because a room's noise is not a talk.
Fix Three: The Realtime Model Can Drop Words
His next test ran 132 seconds without a cut, and twenty seconds of numbers in the second minute were simply missing from the transcript. The recording was continuous speech through that stretch, so every chunk had reached the provider; the realtime model just never wrote those words. The batch model wrote every word of the same file in 3.3 seconds: 645 characters against the live 535. So any utterance of 30 seconds or more is now transcribed again, whole, by the batch path before it goes out, in the same call that keeps the recording. If that call fails, the live words go out instead. The cost is a few seconds after a long talk, and the batch call carries no keyterms, so a soul's name can come back spelled differently. Losing twenty seconds of what Dad said is worse.