"The rivalry is not a winner. It is a table with jobs down the side, and each row has an answer that the earlier lessons already computed."
The Table
Five lessons of arithmetic reduce to one question per job: where does the model land — fits, spills, splits — on each machine, and what is the stage the job spends its time in. The code block is that question as a function, with the household's own machine shapes and a rented card, and the table below is what it returns for the jobs a household actually has. Every cell is a ceiling from a vendor bandwidth or a boolean from a vendor capacity; the ranking is this quest's judgment, and the judgment column says so.
| Job | Decisive fact | Answer | Evidence |
|---|---|---|---|
| One user, a model that fits a card (≤ 27B at 4 bits) | bandwidth ratio 2.2× the Studio, 4× on an H100 | the card — 124 vs 57 tok/s ceiling; the Mac is not the fast choice here | derived from vendor bandwidth |
| One user, a model above 96 GB (a 235B mixture at 132 GB resident, a 405B at 228) | fits the 512 GB pool; the 235B also fits a pair of 96 GB cards, the 405B does not | the Mac Studio — the only single machine on the table that runs either, at 270 W and no fan you can hear | derived; vendor (power, Apple's test) |
| One user, a 70B at 4.5 bits per weight | fits a 96 GB workstation card at 45 tok/s, the Studio at 21, the Spark at 7; spills on a 5090 at 1.6 | the workstation card if you have the room and the power; the Studio if you want it on a desk | derived |
| Many users at once, any model that fits | batch turns decode into compute; tensor cores win compute | NVIDIA, and the larger the batch the more so — the batch lesson measured the Mac's aggregate at 7× single-stream, a card's goes further | measured (Mac), physics |
| Training or a full fine-tune at any real scale | 16 bytes per parameter; fifty-fold tensor compute | NVIDIA, rented or racked; the Mac for adapters and small models | physics, vendor throughput |
| Image generation, one user | compute stage; a card's tensor cores against the Mac's 22 TFLOP/s | the card is faster per image; the household runs it on the Mac because the same box serves everything else and the images are not urgent | measured (Mac), our-judgment |
| Always-on, in a room, for years | 9 W idle, 270 W max, silent; against 575–1,000 W and a case fan | the Mac — the fleet track is this row | vendor (Apple support), vendor (NVIDIA) |
| Speech and voice for the family | quality of the frontier cloud model, not local capability | the cloud, by choice — the fleet track's cloud-by-choice lesson | the household's ruling |
How the Household Answered
The fleet in Part 4 is this table, filled in. The inference hub is a Mac Studio because the models the household wants to run sit in the second row and the seventh. The image engine runs on the same Studio because the sixth row's speed difference did not matter to the way the family uses images. Nothing in the household trains at scale, and Apple's own models are trained elsewhere, so the fifth row has no local machine and does not need one. Voice went to the cloud on the last row's reasoning, which the household stated as a ruling: better quality, not missing capability. And every Mac in the house is an always-on machine in a room where people live, which is the row the rival rarely wins — the 140 W Spark is the one rival built for it — and the quest's last track is about. None of this makes the Mac the better computer. It makes it the right one for the rows this household is in — and the card the right one for the rows it is not, which is why the table is printed in a quest about Apple silicon.