"An experiment proves a thing can be done once. An operation is the thing being done without you. The household keeps both, and labels which is which, because the label is what decides whether a machine is allowed to be turned off."
Six Questions
The code block is the line as a checklist, and it is the fleet track's three placement rules plus three about time. Does it start by itself after a reboot — a launchd job, not a terminal someone left open? Does it have exactly one owner process on one machine, the single-writer rule? Does something depend on it daily? Is a failure noticed by a monitor rather than by a person wondering why a thing is slow? Can it be rebuilt on a replacement Mac in an afternoon from records — the store's catalog, the engine's configuration, the quest workshop's briefs? And were its numbers measured after deployment, not only in the experiment that preceded it? An operation answers yes to all six. Applied to the household: the inference hub, the model store and the video-memory control plane are six for six. This quest's lab — four Macs, one script, one day — is two for six and was never meant to be more: it is an experiment whose records are the measurements companion, and it can be rerun by anyone with the files. A 1 TB model pooled over RDMA is zero for six on the household's own trial, which is what "not practical" meant.
Two Things in the Middle
The interesting rows are the fours. The music engine pins its source separator to the CPU: it starts by itself, has one owner, is depended on, and can be rebuilt — but nothing monitors it and its numbers were measured before deployment against a version of PyTorch that has since changed, which is how the journey track found it running ten times slower than it could. That is an operation with a debt: the sixth question unanswered since the day it was deployed. Ollama's MLX engine as a chat tier is the other four: it starts, it is owned, it is rebuildable, and the lab measured it after deployment — the speculative-decoding finding — but nothing depends on it daily yet and nothing watches it. It is an operation waiting for a dependent, which the tier lesson's setting could give it in a minute. The line is not a wall; it is a count, and the count says which question to answer next.
The Line for Big Models
Everything this track computed sorts by the checklist. GLM-5.3 at 4 bits on one Studio could be an operation: it fits, it would start from launchd, the hub could own it, and its numbers can be measured after deployment — and the household has not deployed it, because the sixth question's honest answer for a 753B mixture is single-digit tokens per second, which nothing in the family would depend on daily when the cloud tier answers faster and better. Qwen3.8-Flash-Next at 4 bits on a laptop is an experiment a laptop's owner can run this week and, at 6B active, might well become an operation. A pooled model is an afternoon. The household's rule for the whole track, stated as its verdict on RDMA and carried here: what runs in the house is what can be left alone, and the store keeps everything else for the day the answer changes. The edge track is about that day.