Skip to content
C.W.K.
Stream
Lesson 03 of 05 · published

When Is He Done Talking?

~14 min · end-of-speech, vad, turn-taking, settings

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"One second is definitely too short. It makes me rush... It needs to be three. Make three the default and widen the range to one to five." — Dad, 2026-09-25

The Hardest Question in the Loop

A hands-free loop has to decide, with no button, that the person has finished. Decide too early and you send half a thought; the soul answers the wrong question and the person has to start over. Decide too late and every turn opens with dead air. There is no right answer in general. Korean speakers in particular pause mid-sentence while choosing the next word, and those pauses can be long.

Start With the Provider's Silence

Scribe's realtime socket, with the vad commit strategy, commits a segment after a stretch of silence you choose. The design started at 1.0 second, the aggressive end, on the theory that a quick reply feels natural. Dad tried it on his phone, and his verdict was the quote at the top of this page: one second made him rush for fear of being cut off. He set the default to 3.0 seconds and asked for a range of 1 to 5. Each device can set its own window in half-second steps, and a house setting in Admin covers the web and any phone that hasn't chosen. A device's own choice always wins.

The Provider Has a Ceiling

Then the measurement. Scribe refuses a silence threshold above 3.0 seconds, with an explicit error that the value must lie between 0.3 and 3.0. Dad's range goes to 5. So the socket gets min(window, 3.0), the session response names both numbers, and the client owns the difference. After a commit, it waits out the rest of Dad's window before sending, so if he starts again in that gap, the turn stays open for him.

Listening Inside the Wait

Waiting is not enough on its own, because Scribe's words trail the speech by about a second. If Dad starts talking again two seconds into the wait, the transcriber may not produce a single word before the timer fires. So during the wait the client listens to loudness itself. It takes the committed silence as the room's floor, and two chunks in a row above three times that floor, and above an absolute minimum, mean he is talking again. The timer is cancelled and the turn stays open. If words then arrive, listening simply continues to the next commit. If no words confirm the sound within 2.5 seconds (a cough, a door, a cup set down), the turn goes out with what was already said.

One edge stays open by construction. Two loud chunks take about a quarter of a second to arrive, so speech that begins in the last quarter-second of the wait reaches the deadline before it is confirmed, and the first part goes out as its own turn. The fourth scenario in the code is exactly that case. A wider window moves the edge; nothing removes it.

The Bench, Then the Ear

The bench test was one recording: 4.4 seconds of speech, 3.3 seconds of silence, the same speech again. With a 5-second window the two halves became one turn. With the 3-second default the first sentence went out alone, which is correct, because 3.3 seconds of silence is longer than the window. Dad's verdict once the build carrying it reached his phone: three seconds is the most natural, and the conversation flows. A window that adapts to the talk is written down as a todo, to be built only when he asks. The measurement that matters here is his ear, and it has spoken.

Code

After the provider commits: wait out the rest of Dad's window·python
from dataclasses import dataclass

CHUNK_S = 0.128          # 2048 samples at 16 kHz
RESUME_CONFIRM_S = 2.5   # sound with no words for this long is not speech


@dataclass
class Event:
    t: float     # seconds after the provider's commit
    kind: str    # "level" (RMS of one chunk) or "words" (a partial transcript)
    value: float = 0.0


def after_commit(events: list[Event], floor: float, window: float,
                 provider_silence: float) -> str:
    """Decide what the provider's commit means for the turn."""
    extra = max(0.0, window - provider_silence)   # the part the client owns
    loud_run = 0
    resume_at: float | None = None
    for event in sorted(events, key=lambda e: e.t):
        if resume_at is None and event.t >= extra:
            return f"send at +{extra:.1f}s: the window passed in silence"
        if resume_at is not None and event.t >= resume_at + RESUME_CONFIRM_S:
            return "send: sound came back but no words confirmed it"
        if event.kind == "words":
            return f"keep listening: he is talking again (+{event.t:.1f}s)"
        if event.kind == "level" and resume_at is None:
            loud_run = loud_run + 1 if event.value > max(floor * 3, 300) else 0
            if loud_run >= 2:
                resume_at = event.t          # hold the turn open, await words
    return "send: nothing more happened"


quiet = [Event(i * CHUNK_S, "level", 120) for i in range(40)]
print(after_commit(quiet, floor=110, window=5.0, provider_silence=3.0))

resumes = quiet[:8] + [Event(1.1, "level", 900), Event(1.23, "level", 950),
                       Event(2.2, "words", 0)]
print(after_commit(resumes, floor=110, window=5.0, provider_silence=3.0))

cough = quiet[:8] + [Event(1.1, "level", 900), Event(1.23, "level", 950),
                     Event(4.0, "level", 120)]
print(after_commit(cough, floor=110, window=5.0, provider_silence=3.0))

late = quiet[:15] + [Event(1.9, "level", 900), Event(2.03, "level", 950),
                     Event(2.5, "words", 0)]
print(after_commit(late, floor=110, window=5.0, provider_silence=3.0))   # the edge

External links

Exercise

Run the four scenarios in the code block. Then add a fifth: Dad pauses 4 seconds with a 5-second window and a 3-second provider threshold, and says nothing more. Predict the output before running it. Finally, change the window to 3.0 and rerun all five; explain which ones change and why.
Hint
With the window equal to the provider's threshold, the client owns zero extra seconds, so the first event at or after +0.0 s sends the turn and the loudness logic never gets a chance. That is exactly the 3-second default's behavior: the provider's commit is the end of speech, unless the next lesson's forced-commit check says otherwise.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.