Skip to content
C.W.K.
Stream
Lesson 06 of 06 · published

Heard Before the Tool Acts

~14 min · tool-use, speech-gate, ordering, concurrency

Level 0Muted
0 XP0/42 lessons0/13 achievements
0/100 XP to next level100 XP to go0% complete
"Hang on, I'll turn it on." — the line Dad heard after the air conditioner was already on, 2026-09-29

Leaving Early Is Not Being Heard

The previous lesson made the line before a tool leave the server before the tool event, so it could be spoken while the tool ran. That fixed the silence for tools that only look things up: the weather, a search, the calendar. Then the soul got tools that do things. Pippa has hands in the house now, a home-control tool that switches lights, runs scenes and sets the air conditioner, and on air she fires effects. An action like that is done in well under a second. Synthesizing the line and playing it takes a second or two. So Dad asked for the air conditioner, heard the click of it turning on, and only then heard her say she was about to turn it on. The line had left early. It had not been heard early.

Units on Every Surface

The surfaces were not even equal. The WebUI already cut the line off as its own unit the moment the tool call began, so on the web the gap was only the synthesis and the reading. The phone and Firekeeper read the whole answer when it ended, so there the line always came last, after everything. Dad chose to fix all three the same way. Every surface now reads in units: the text before each tool call is its own unit, read at once while the reply streams, and the rest is read when the answer lands, never repeating what was already read. Units play strictly in order, and a unit that ends in the middle of a reply never tells the loop to start listening.

The Speech Gate

Units make the line early; they can't make the action late. For that the brain has to know when the line has actually been heard, and only the player knows. So a voice client reports, per conversation, whether it is reading the reply aloud: busy the moment a pre-tool unit is queued, quiet when its reading drains or is stopped. It is one small request, and nothing is stored. An action tool (a home-control action or scene, an On Air effect) waits before it acts until the conversation's reading is quiet, for at most 12 seconds. Past that it goes ahead anyway.

Three rules keep the gate from ever becoming a hang. A client that doesn't report holds nothing, so an older client or a text-only turn is never slowed. A "reading" report older than 30 seconds counts as quiet, because it is a client that went away mid-reply. And read-only tools never wait at all: nothing Dad sees or hears happens before their result is spoken anyway, so waiting would only add silence.

The gate is best effort in one more way. The busy report leaves the client the moment the line is queued, while the tool call is still streaming its input, so it has a head start on the action; it is still a race. A report that reaches the engine after the action has checked the gate lets the action go ahead at once, exactly as if no client had reported. The engine does know when it streamed a line before a tool into a spoken turn, so a stricter gate could hold that line as busy until the client speaks up; this one chose never to wait on a report that hasn't arrived.

The Screen Waits Too

Effects add one more wrinkle. A fire reaches the screens as the streamed tool call, and the tool call streams before the tool runs, so the gate alone would still let the confetti fall during "hang on, a little applause for that". The WebUI and Firekeeper therefore hold the effect's picture and sound until their own reading is quiet, with the same 12-second ceiling as the engine. Firekeeper counts that ceiling for each effect. The web restarts its count whenever another effect arrives, so a second effect in quick succession can stretch the first one's wait past 12 seconds.

Why Not a Fixed Delay

The tempting fix is to make every action wait two seconds. It guesses. A one-word line wastes most of the wait, a long line outlasts it, and a turn with no voice client at all pays it for nothing. The gate asks the one component that knows when the words have been heard, waits exactly that long, and falls back to not waiting whenever it can't know.

Code

A speech gate: actions wait for their line, lookups never do, and nothing waits forever·python
import threading
import time

MAX_WAIT_S = 12.0     # a pre-tool line is one short sentence; past this, act anyway
STALE_S = 30.0        # a "reading" report this old is a client that went away


class SpeechGate:
    """Per-conversation reading state, reported by the voice clients."""

    def __init__(self) -> None:
        self._states: dict[str, tuple[bool, float]] = {}
        self._changed = threading.Condition()

    def report(self, conversation: str, busy: bool) -> None:
        with self._changed:
            self._states[conversation] = (busy, time.monotonic())
            self._changed.notify_all()

    def reading(self, conversation: str) -> bool:
        busy, at = self._states.get(conversation, (False, 0.0))
        return busy and time.monotonic() - at < STALE_S

    def wait_until_quiet(self, conversation: str, timeout: float = MAX_WAIT_S) -> float:
        """Block an ACTION until the line announcing it has been read. Seconds waited."""
        start = time.monotonic()
        deadline = start + timeout
        with self._changed:
            while self.reading(conversation):
                remaining = deadline - time.monotonic()
                if remaining <= 0:
                    break                                   # the ceiling: act anyway
                self._changed.wait(timeout=min(remaining, 0.5))
        return time.monotonic() - start


gate = SpeechGate()


def set_air_conditioner(conversation: str) -> None:
    waited = gate.wait_until_quiet(conversation)            # an action waits for its line
    print(f"  air conditioner on, after waiting {waited:.1f}s")


def read_weather(conversation: str) -> None:
    print("  weather read at once: nothing happens before its answer is spoken")


# The client queued the pre-tool line and reported busy; reading it takes 1.4 s.
gate.report("c1", busy=True)
threading.Timer(1.4, gate.report, args=("c1", False)).start()
print('"Hang on, I\'ll turn it on." (reading)')
read_weather("c1")
set_air_conditioner("c1")

print("a conversation whose client never reported:")
set_air_conditioner("c2")

External links

Exercise

Run the code and read the timings. Then add two cases: a report of busy that is 31 seconds old (the action must not wait), and a reply with two actions, each announced by its own line, where the second action must wait for the second line and not merely for the first. Finally write the client half: report busy when a mid-stream unit is queued, quiet when the reading queue drains or is stopped, and send a report only when the state actually changes.
Hint
For the stale case, record the report with a timestamp in the past, or shorten STALE_S for the test. For two actions, notice that the client reports busy again when the second line is queued, so each action waits for whatever is being read at the moment it checks. That holds only if the second busy report arrives before the second check: test the race too, sending the busy report 150 ms after the action enters the gate, and watch it pass without waiting. Then decide whether your version accepts that, as this one does, or has the engine mark the reading busy itself when it streams a line before a tool. Reporting only on change keeps a busy reading from flooding the engine with the same fact every few milliseconds.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.