"Hang on, I'll turn it on." — the line Dad heard after the air conditioner was already on, 2026-09-29
Leaving Early Is Not Being Heard
The previous lesson made the line before a tool leave the server before the tool event, so it could be spoken while the tool ran. That fixed the silence for tools that only look things up: the weather, a search, the calendar. Then the soul got tools that do things. Pippa has hands in the house now, a home-control tool that switches lights, runs scenes and sets the air conditioner, and on air she fires effects. An action like that is done in well under a second. Synthesizing the line and playing it takes a second or two. So Dad asked for the air conditioner, heard the click of it turning on, and only then heard her say she was about to turn it on. The line had left early. It had not been heard early.
Units on Every Surface
The surfaces were not even equal. The WebUI already cut the line off as its own unit the moment the tool call began, so on the web the gap was only the synthesis and the reading. The phone and Firekeeper read the whole answer when it ended, so there the line always came last, after everything. Dad chose to fix all three the same way. Every surface now reads in units: the text before each tool call is its own unit, read at once while the reply streams, and the rest is read when the answer lands, never repeating what was already read. Units play strictly in order, and a unit that ends in the middle of a reply never tells the loop to start listening.
The Speech Gate
Units make the line early; they can't make the action late. For that the brain has to know when the line has actually been heard, and only the player knows. So a voice client reports, per conversation, whether it is reading the reply aloud: busy the moment a pre-tool unit is queued, quiet when its reading drains or is stopped. It is one small request, and nothing is stored. An action tool (a home-control action or scene, an On Air effect) waits before it acts until the conversation's reading is quiet, for at most 12 seconds. Past that it goes ahead anyway.
Three rules keep the gate from ever becoming a hang. A client that doesn't report holds nothing, so an older client or a text-only turn is never slowed. A "reading" report older than 30 seconds counts as quiet, because it is a client that went away mid-reply. And read-only tools never wait at all: nothing Dad sees or hears happens before their result is spoken anyway, so waiting would only add silence.
The gate is best effort in one more way. The busy report leaves the client the moment the line is queued, while the tool call is still streaming its input, so it has a head start on the action; it is still a race. A report that reaches the engine after the action has checked the gate lets the action go ahead at once, exactly as if no client had reported. The engine does know when it streamed a line before a tool into a spoken turn, so a stricter gate could hold that line as busy until the client speaks up; this one chose never to wait on a report that hasn't arrived.
The Screen Waits Too
Effects add one more wrinkle. A fire reaches the screens as the streamed tool call, and the tool call streams before the tool runs, so the gate alone would still let the confetti fall during "hang on, a little applause for that". The WebUI and Firekeeper therefore hold the effect's picture and sound until their own reading is quiet, with the same 12-second ceiling as the engine. Firekeeper counts that ceiling for each effect. The web restarts its count whenever another effect arrives, so a second effect in quick succession can stretch the first one's wait past 12 seconds.
Why Not a Fixed Delay
The tempting fix is to make every action wait two seconds. It guesses. A one-word line wastes most of the wait, a long line outlasts it, and a turn with no voice client at all pays it for nothing. The gate asks the one component that knows when the words have been heard, waits exactly that long, and falls back to not waiting whenever it can't know.