Skip to content
C.W.K.
Stream
Lesson 05 of 05 · published

A Line Before the Tool

~14 min · tool-use, streaming, filters, spoken-units

Level 0Muted
0 XP0/35 lessons0/12 achievements
0/100 XP to next level100 XP to go0% complete
"Before you call a tool, say one short line first ('hang on, let me look'), so the silence has a reason." — the spoken-turn instruction

Four in Twelve Hundred

Tools make a soul useful: the weather, the calendar, a search, a file. They also take time, and in a written chat nobody minds, because the answer appears when it's ready. The design sweep on 2026-09-25 counted how often a soul said anything before calling a tool: 4 times in roughly 1,265 tool turns. In text, that silence is invisible. In speech it sounds exactly like a dropped call. So the spoken instruction asks for one short line before any tool, something like "hang on, let me check the fine dust in Ilsan", and the voice pipeline turns that line into its own spoken unit, played while the tools run.

The Line Was Written, and Nobody Heard It

Adding the instruction worked on the first try: the model wrote the line. Then it sat on the server. Two filters on the Claude route, both older than voice mode and both there for good reasons, were holding text back.

  • A leading-scaffold filter held up to the first 600 characters of every reply while it watched for a leaked "Assistant" role marker, a rare model glitch it exists to remove. It released its hold only when the reply was done.
  • A soft-thinking filter kept the last nine characters of every stream in case they were the start of a <thinking> tag.

Together they meant the line before the tool couldn't leave the server until the whole reply, including everything after the tools, was finished. The fix was not to remove either filter. It was to flush both, in the same order the end-of-reply path already used, at the moment a tool call begins. A tool call is a natural boundary: the text before it is final.

Measured After

On the live test after the fix, the line "hang on, let me check the fine dust in Ilsan" left the server 3.35 seconds into the turn, before the tool event. It was spoken while two shell calls ran, and the answer followed as its own unit at 12.8 seconds. That is the shape of a person who says "one sec" and then comes back with the answer.

The Measurement the Filter Faked

There is a second lesson here. An early census of streaming behavior found the median text delta was about 300 characters and concluded the model streams in big chunks. It doesn't. That number was the leading-scaffold filter releasing its whole hold at once. Written replies under 600 characters still appear in one piece at the end, which is now a display question about all turns, not a voice question. But every measurement taken downstream of a buffer measures the buffer first.

Code

A holding filter that releases at tool boundaries, and the units that result·python
from dataclasses import dataclass, field


@dataclass
class HoldingFilter:
    """Holds the start of a reply while it watches for a leaked role marker."""
    hold_chars: int = 600
    held: str = ""
    released: bool = False

    def feed(self, delta: str) -> str:
        if self.released:
            return delta
        self.held += delta
        if len(self.held) < self.hold_chars:
            return ""                       # still watching
        return self.flush()

    def flush(self) -> str:
        out, self.held, self.released = self.held.removeprefix("Assistant: "), "", True
        return out


@dataclass
class SpokenStream:
    flush_at_tools: bool
    filter: HoldingFilter = field(default_factory=HoldingFilter)
    sent: str = ""
    unit_ends: list[int] = field(default_factory=list)
    log: list[str] = field(default_factory=list)

    def on_event(self, t: float, kind: str, text: str = "") -> None:
        if kind == "text":
            out = self.filter.feed(text)
        elif kind == "tool_use":
            out = self.filter.flush() if self.flush_at_tools else ""
            if out or self.sent:
                self.unit_ends.append(len(self.sent) + len(out))   # a unit ends here
        else:  # done
            out = self.filter.flush()
        if out:
            self.sent += out
            self.log.append(f"{t:5.2f}s -> client: {out.strip()!r}")


events = [(2.9, "text", "잠깐, 일산 미세먼지 확인해볼게. "), (3.35, "tool_use"),
          (12.8, "text", "지금은 보통이야. 산책해도 괜찮아."), (13.1, "done")]
for flush in (False, True):
    stream = SpokenStream(flush_at_tools=flush)
    for event in events:
        stream.on_event(*event)
    print(f"flush at tools={flush}:", *stream.log, sep="\n  ")

External links

Exercise

Run the code and compare the two logs. Then add the second filter: one that always keeps the last nine characters back in case they start a thinking tag. Put it in front of the first filter, upstream, the way cwkPippa orders them, so the role-marker check reads text with any thinking block already removed. Make both flush at a tool call in the right order, and check that the pre-tool line still leaves at 3.35 s with its final nine characters intact.
Hint
Order matters because each filter can be holding text the next one hasn't seen. Flush the upstream thinking filter first and feed what it releases through the role-marker filter before flushing that one too, exactly as the end-of-reply path does. Reverse the placement and a role marker that follows a thinking block slips through, because the role-marker filter only looks at the start of what it is fed. If you flush them in the wrong order, the last few characters of the line arrive after the answer.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign in — Please sign in to comment.

No comments yet — be the first.