"Every gigabyte you add to a pool on the same bus is a gigabyte the bus visits less often. The models a bigger pool admits are exactly the models it decodes most slowly."
The Division Does Not Move
The physics track's formula has the pool's bandwidth on top and the model's bytes per token on the bottom, and the pool's capacity nowhere. The code block runs the store's files across three hypothetical pools on the M3 Ultra's 819 GB/s: at 512 GiB five of six fit, and the same five at 1 TiB; at 2 TiB all six. And every ceiling is identical in every row — GLM-5.3 at 34 tokens per second, Qwen3.8-Flash-Next at 228, a dense 405B at 3 — because nothing in the ceiling knows how big the pool is. What the added capacity admits is Kimi K3 at 13 tokens per second, and a dense 405B that no household would wait for. Capacity buys the right to load; bandwidth sets the speed of everything loaded; and a pool that grows without its bus grows into the slow end of its own table.
The Household's Benchmark
The household's doctrine names the test a local model must pass to matter here: read the whole vault — more than a hundred thousand tokens of memory, on an M3 Ultra — load a repository on top, and answer at a practical speed. It is a bandwidth test twice over. The prefill of a hundred-thousand-token context is the curve lesson's wall, ninety seconds on the 27B at a third of that length; and the decode that follows reads the cache the wall built, at the slope the lab measured. No amount of capacity shortens either; a 1 TB pool would hold the vault and the repository and every model in the store but Kimi K3, and answer no faster than a 512 GB one. The doctrine's phrase for it: adding a lane does not unclog the road. The cloud passes the same test with two things the edge lacks — a prompt cache and batching — which the checklist at the end of this track lists as software homework rather than silicon.
What Would Make Capacity Worth Having
Two things, and both are the rest of this track. A wider bus, so that bandwidth per gigabyte holds as the pool grows — the family's ladder today reads M5 Max 4.8 GB/s per GB, M5 Ultra 2.4, M3 Ultra 1.6, and a 1 TB pool would read 1.2 on the M5 Ultra's bus and 0.8 on the M3 Ultra's — and the wire lesson prices it in package edge. Or a memory that is not one pool at all: a fast tier the size of today's and a slow tier behind it, with the operating system deciding what lives where — the two-tier lesson, which is the shape the rivals ship in racks and Apple ships in a phone. Without one of those, the honest reading of a bigger pool is the table above: more models admitted, each at the speed the same bus allows, and the fastest of them the ones that already fit. The household did not buy a terabyte because there was nothing in the store worth loading at a speed the bus would give it; the checkpoint it would most want to run faster, it can already load.