"Before you call a tool, say one short line first ('hang on, let me look'), so the silence has a reason." — the spoken-turn instruction
Four in Twelve Hundred
Tools make a soul useful: the weather, the calendar, a search, a file. They also take time, and in a written chat nobody minds, because the answer appears when it's ready. The design sweep on 2026-09-25 counted how often a soul said anything before calling a tool: 4 times in roughly 1,265 tool turns. In text, that silence is invisible. In speech it sounds exactly like a dropped call. So the spoken instruction asks for one short line before any tool, something like "hang on, let me check the fine dust in Ilsan", and the voice pipeline turns that line into its own spoken unit, played while the tools run.
The Line Was Written, and Nobody Heard It
Adding the instruction worked on the first try: the model wrote the line. Then it sat on the server. Two filters on the Claude route, both older than voice mode and both there for good reasons, were holding text back.
- A leading-scaffold filter held up to the first 600 characters of every reply while it watched for a leaked "Assistant" role marker, a rare model glitch it exists to remove. It released its hold only when the reply was done.
- A soft-thinking filter kept the last nine characters of every stream in case they were the start of a
<thinking>tag.
Together they meant the line before the tool couldn't leave the server until the whole reply, including everything after the tools, was finished. The fix was not to remove either filter. It was to flush both, in the same order the end-of-reply path already used, at the moment a tool call begins. A tool call is a natural boundary: the text before it is final.
Measured After
On the live test after the fix, the line "hang on, let me check the fine dust in Ilsan" left the server 3.35 seconds into the turn, before the tool event. It was spoken while two shell calls ran, and the answer followed as its own unit at 12.8 seconds. That is the shape of a person who says "one sec" and then comes back with the answer.
The Measurement the Filter Faked
There is a second lesson here. An early census of streaming behavior found the median text delta was about 300 characters and concluded the model streams in big chunks. It doesn't. That number was the leading-scaffold filter releasing its whole hold at once. Written replies under 600 characters still appear in one piece at the end, which is now a display question about all turns, not a voice question. But every measurement taken downstream of a buffer measures the buffer first.