Skip to content
C.W.K.
Stream
Lesson 02 of 05 · published

Hearing Dad Over Her Voice

~14 min · barge-in, detection, thresholds, pre-roll

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Talking over the soul stops her, and she learns how far Dad heard." — step 3 of the voice mode build, 2026-09-26

Loudness, Not Words

Once the canceller has removed most of her voice from the microphone, what's left while she speaks is a low floor: the residue of her echo and the room. Dad talking over her shows up as a jump well above that floor. The detector works on loudness alone, not on words, because words arrive from the transcriber about a second late and a cut-in that waits a second feels like being ignored.

The Web Detector

On the web, the microphone delivers chunks of 2048 samples at 16 kHz, 128 milliseconds each. The detector spends the first five chunks after she starts speaking learning the floor, taking their median so one plosive can't skew it. After that, three chunks in a row above the larger of 3.5 times the floor and an absolute 700 RMS count as Dad: that's about 0.4 seconds of sustained sound, long enough to ignore a cough or a clink and short enough to feel immediate. The floor never freezes; every quiet chunk nudges it toward the current level, so it follows her voice getting louder or softer through a long reply. The phone runs the same idea with its own tuning, four chunks of 100 milliseconds.

Don't Lose the Words That Cut In

By the time the detector is sure, Dad has already said a word or two, and the listening socket isn't open yet: while she speaks, audio goes nowhere but the detector. So the loop keeps a rolling pre-roll of about a second while she talks, and everything that arrives while the socket reopens. When a cut-in fires, that audio goes to the transcriber first, and his first words aren't lost.

One Cut-In Path

There are several ways to interrupt her: talk over her, tap her face, press a microphone. Each client routes them into one cut-in function, so every interruption does the same things in the same order: stop playback, report how far Dad heard, and start listening. A cut that lands while the reply is still streaming waits for the reply's real message id before reporting, so the heard position is filed against the right message.

The design lists a fourth way, sending a message while she's speaking, and as this quest is written the web doesn't route it there yet. Its send path stops her and sends, and reports nothing, so that reply's record says Dad heard all of it. One missed door into the shared path is exactly how a single rule turns into two behaviors; the way to catch it is a test that drives every interruption and asserts each one reports a position.

Thresholds Are First Guesses

Every number above was chosen on the bench, and the design says so: the thresholds are first guesses until Dad's own speakers and rooms say otherwise. Each client therefore has an off switch: a "Cut in" toggle on the web's face screen, and a setting on the phone for cutting in by talking over Pippa. A detector that fires on its own speaker is worse than none, and the person living with it needs to be able to turn it off without a deploy.

Code

A loudness detector with a moving floor, and a pre-roll that keeps the first words·python
from collections import deque
from dataclasses import dataclass, field
from statistics import median


@dataclass
class BargeInDetector:
    """Dad talking over her, heard as loudness on the open, echo-cancelled mic."""
    calibrate: int = 5        # chunks that learn the floor: her echo + the room
    sustain: int = 3          # loud chunks in a row that count as Dad
    ratio: float = 3.5        # how far above the floor
    minimum: float = 700.0    # and never below this RMS
    seen: list = field(default_factory=list)
    floor: float | None = None
    loud: int = 0

    def reset(self) -> None:
        self.seen, self.floor, self.loud = [], None, 0

    def feed(self, level: float) -> bool:
        if self.floor is None:
            self.seen.append(level)
            if len(self.seen) >= self.calibrate:
                self.floor = median(self.seen)
            return False
        if level > max(self.floor * self.ratio, self.minimum):
            self.loud += 1
            return self.loud >= self.sustain
        self.loud = 0
        self.floor = self.floor * 0.9 + level * 0.1   # keep following the quiet
        return False


PREROLL_CHUNKS = 8            # about a second at 128 ms per chunk
preroll: deque = deque(maxlen=PREROLL_CHUNKS)
detector = BargeInDetector()

her_echo = [320, 410, 380, 350, 400, 900, 360, 390]     # one loud blip: a plosive
dad = [2400, 2800, 2600, 2500]
for t, level in enumerate(her_echo + dad):
    preroll.append((t, level))                          # the words that cut in, kept
    if detector.feed(level):
        print(f"cut in at chunk {t} ({t * 0.128:.2f}s): stop her, report heard, listen")
        print("  preroll sent to the transcriber first:", [c for c, _ in preroll])
        break
else:
    print("no cut-in")

External links

Exercise

Run the code and see where the cut-in fires and which chunks the pre-roll hands over. Then make the echo louder (raise every value in her_echo by 3x) and rerun. Does the detector still fire on Dad? Does it ever fire on her? Adjust ratio and minimum until it fires on Dad within 0.5 seconds and never on the echo, and write down what you gave up.
Hint
Raising the minimum protects against a loud echo but makes a soft-spoken interruption harder to catch; raising sustain protects against blips but delays every cut-in. There is no setting that wins everywhere, which is exactly why the thresholds are first guesses, the room decides, and the off switch exists.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.