"One second is definitely too short. It makes me rush... It needs to be three. Make three the default and widen the range to one to five." — Dad, 2026-09-25
The Hardest Question in the Loop
A hands-free loop has to decide, with no button, that the person has finished. Decide too early and you send half a thought; the soul answers the wrong question and the person has to start over. Decide too late and every turn opens with dead air. There is no right answer in general. Korean speakers in particular pause mid-sentence while choosing the next word, and those pauses can be long.
Start With the Provider's Silence
Scribe's realtime socket, with the vad commit strategy, commits a segment after a stretch of silence you choose. The design started at 1.0 second, the aggressive end, on the theory that a quick reply feels natural. Dad tried it on his phone, and his verdict was the quote at the top of this page: one second made him rush for fear of being cut off. He set the default to 3.0 seconds and asked for a range of 1 to 5. Each device can set its own window in half-second steps, and a house setting in Admin covers the web and any phone that hasn't chosen. A device's own choice always wins.
The Provider Has a Ceiling
Then the measurement. Scribe refuses a silence threshold above 3.0 seconds, with an explicit error that the value must lie between 0.3 and 3.0. Dad's range goes to 5. So the socket gets min(window, 3.0), the session response names both numbers, and the client owns the difference. After a commit, it waits out the rest of Dad's window before sending, so if he starts again in that gap, the turn stays open for him.
Listening Inside the Wait
Waiting is not enough on its own, because Scribe's words trail the speech by about a second. If Dad starts talking again two seconds into the wait, the transcriber may not produce a single word before the timer fires. So during the wait the client listens to loudness itself. It takes the committed silence as the room's floor, and two chunks in a row above three times that floor, and above an absolute minimum, mean he is talking again. The timer is cancelled and the turn stays open. If words then arrive, listening simply continues to the next commit. If no words confirm the sound within 2.5 seconds (a cough, a door, a cup set down), the turn goes out with what was already said.
One edge stays open by construction. Two loud chunks take about a quarter of a second to arrive, so speech that begins in the last quarter-second of the wait reaches the deadline before it is confirmed, and the first part goes out as its own turn. The fourth scenario in the code is exactly that case. A wider window moves the edge; nothing removes it.
The Bench, Then the Ear
The bench test was one recording: 4.4 seconds of speech, 3.3 seconds of silence, the same speech again. With a 5-second window the two halves became one turn. With the 3-second default the first sentence went out alone, which is correct, because 3.3 seconds of silence is longer than the window. Dad's verdict once the build carrying it reached his phone: three seconds is the most natural, and the conversation flows. A window that adapts to the talk is written down as a todo, to be built only when he asks. The measurement that matters here is his ear, and it has spoken.