Numbers That Earned a Gate
Late August 2026, same two-step filesystem task, fixed seeds, a local Qwen on a family daemon. Loop finished 3/3 in four steps and about seven thousand tokens. Plan invented six steps, hit the step cap, cost about six times as much, and left a run incomplete. After the plan learned an executable budget from that turn's own caps — a plan step costs about two model steps, and the instruction forbids splitting work one tool call finishes — planned steps fell to about 1.7, the cap vanished, success went 3/3. Plan was then correct and still about 2.8× the loop. Loop stayed the local default. Plan earns its cost when the numbered list is the deliverable.
Compaction was hijacked by conversational momentum: the summarizer, given the pending operator prompt, answered the task instead of summarizing. The fix excludes that trailing prompt, compares estimate against estimate so a fresh marker does not look stale, and falls through to an ordinary step after one failed attempt. A later staged crossing summarized once and recalled four planted facts, three of them only in the summary. The gate has a remaining blind spot for a single result that jumps past the window; that is named, not wished away.
Fidelity Is Per Release
The same small-model battery on daemon 0.33.0: a 27B Qwen 3/3 with no tool errors; the latest Qwen tag 3/3; a 31B Gemma 3/3 after the previous daemon had eaten retries; a 20B mixture-of-experts still 0/3 because the daemon 500s on a year-old blob that ships no tool capability. The upgrade cost nothing and bought competence back on the weaker arms. The model that still fails has not earned the bench by being on disk.
The mini door cost about twice the tokens on short tool tasks and changed competence not at all — 3/3 both arms. Attach it for the voice. Same-model delegate was 2/2 on a healthy pair and failed typed when the child named a broken model — no mutation, honest report. Those rows changed instructions, gates, and defaults. They did not become a leaderboard.