Skip to content
C.W.K.
Stream
Lesson 03 of 05 · published

The Grace Window

~12 min · barge-in, grace-period, settings, regressions

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"v4 opens replies with loud laughs that the open mic heard as Dad and cut her off." — the commit message, 2026-09-29

A Better Voice Broke the Detector

The night Pippa's voice moved to Eleven v4, a new failure appeared on the web. The new model is more expressive, and when a reply begins with an audio tag like [laughs], it performs a real, loud laugh. Echo cancellation reduces what the microphone hears of the speaker, but a sudden burst well above the level it has been tracking can still come through louder than the floor the detector learned from her first words. Three chunks of laugh above three and a half times that floor looked exactly like Dad talking. She cut herself off, mid-laugh, before saying anything.

Ignore the First Seconds of Her Voice

The fix is a grace window: for the first few seconds after her voice starts, cut-in by voice is ignored. The default is 3 seconds, adjustable from 0 to 10 in Admin → Voice Mode, and 0 turns the grace off entirely. The loop records the moment playback begins and, while inside the window, doesn't feed the detector at all.

One detail matters more than it looks. During the window the detector isn't just ignored; it is reset on every chunk. Otherwise, on a reply that opens with the laugh, it would spend its five calibration chunks learning the floor from the laugh itself and set the floor far too high. Quiet chunks do pull the floor back toward her voice, but at these levels that takes about two and a half seconds, and until then Dad would have to shout to be heard. Resetting means calibration starts only after the window, on her ordinary speaking voice, where the floor belongs.

What You Give Up

The trade is honest and small. For the first three seconds of each reply, Dad can't stop her by talking. He still can by every other route: tapping her face, pressing a microphone, pausing with the key. The grace protects the one interruption method that can be fooled by her own voice, and leaves the ones that can't be fooled alone.

The reset brings one subtler cost of its own. Calibration starts fresh when the window closes and takes five chunks, about 0.64 seconds, so if Dad starts talking in exactly that moment, his voice becomes the floor and that cut-in is missed too. The shipped detector accepts that narrow edge to keep the laugh, which every reply opening with a tag brings, out of the floor.

The Pattern Behind It

This is the voice-lineage lesson from the other side. Nothing in the detector changed and nothing in the detector was wrong; the thing it listens to changed. Output got louder at the start of replies, and an input component calibrated on the old output started to misfire. The fix shipped the same night, as a setting rather than a constant, because the right length depends on the voice, the speakers and the room, and those will keep changing.

Code

A grace window that also keeps calibration off the laugh·python
from statistics import median

CHUNK_S = 0.128


class Detector:
    def __init__(self, calibrate=5, sustain=3, ratio=3.5, minimum=700.0):
        self.calibrate, self.sustain, self.ratio, self.minimum = calibrate, sustain, ratio, minimum
        self.reset()

    def reset(self):
        self.seen, self.floor, self.loud = [], None, 0

    def feed(self, level):
        if self.floor is None:
            self.seen.append(level)
            if len(self.seen) >= self.calibrate:
                self.floor = median(self.seen)
            return False
        if level > max(self.floor * self.ratio, self.minimum):
            self.loud += 1
            return self.loud >= self.sustain
        self.loud, self.floor = 0, self.floor * 0.9 + level * 0.1
        return False


def first_cut_in(levels, grace_s):
    """Seconds into her reading when cut-in fires, or None."""
    detector = Detector()
    for i, level in enumerate(levels):
        t = i * CHUNK_S                      # time since her voice started
        if t < grace_s:
            detector.reset()                 # don't calibrate on a laugh either
            continue
        if detector.feed(level):
            return round(t, 2)
    return None


opening = [360, 340, 380, 350, 370]                    # her first words (the echo)
laugh = [2600, 3100, 2900, 2800, 2500, 2700, 2400]     # then v4 performs "[laughs]"
echo = [350, 380, 330, 360, 400, 370, 340, 390, 360, 350] * 2
dad = [2500, 2700, 2600, 2800]
reading = opening + laugh + echo + dad

print("no grace :", first_cut_in(reading, grace_s=0.0), "s  <- her own laugh cut her off")
print("3 s grace:", first_cut_in(reading, grace_s=3.0), "s  <- Dad, after the window")
print("dad early:", first_cut_in(opening + dad + echo, grace_s=3.0), "   <- inside the window: tap or key instead")

External links

Exercise

Run the code and read the three results. Then write the variant the tip warns against: inside the window, keep feeding the detector but ignore what it returns. In this reading both variants agree, so build two more: one that opens with the laugh (laugh, then twenty echo chunks, then Dad), and one where Dad starts the instant the window ends. Run both variants on each and explain every difference. Finally, decide whether the window should restart at every spoken unit or only at the start of the whole reply, and argue your choice.
Hint
In the given reading both variants end up with a floor from her ordinary level, by different routes: feed-and-ignore calibrates on the five opening chunks, before the laugh, and reset on the first five echo chunks after the window. Those are the same level, so they agree. Open the reply with the laugh and feed-and-ignore learns its floor from the laugh, hearing Dad again only once quiet chunks have pulled it back down. Reset has the mirror-image weakness: if Dad speaks in the 0.64 seconds after the window, reset calibrates on his voice and misses him, while a variant that calibrated on her opening words catches him. Neither is free. For the restart question, look at where v4's laughs actually occur: a unit after a tool call can open with a tag too, so restarting per unit protects more, at the cost of more moments when voice cut-in is unavailable.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.