"Apple shipped the feature. Five of the household's Macs qualify. The household tried it, and its verdict is the one this quest teaches: not practical, not the time. Here is the paper, the practice, and the arithmetic between them."
On Paper
Apple's technote is exact: "RDMA over Thunderbolt is available starting with macOS 26.2 on Macs with Apple silicon with Thunderbolt 5." It is enabled in Recovery — "Run rdma_ctl enable. Reboot" — and "can't route data": cabled pairs only. macOS 26.2's own notes name the use case, "distributed AI inference using MLX", and MLX's distributed backend added the RDMA path in December 2025, describing "latency an order of magnitude lower than the ring backend" and "only fully connected topologies". The fleet lesson's table qualifies five Macs by their Thunderbolt generation — the two 512 GB Studios, the M5 Max and M4 Max laptops, the M4 Pro mini — and disqualifies four on Thunderbolt 4, including both M2 Ultras. Every Mac ships the tool; every Mac reports it disabled. The only pair worth pooling is the Studios: 1,024 GB behind a cable that carries 10 GB/s against "over 2.5TB/s" (Apple's figure) inside each package. On paper, then, pooling exists, Apple supports it, and the wires lesson's ratio of 250 sits in the middle of it.
In Practice, Publicly
Two public results, read with the lab's checklist. A December 2025 test of four M3 Ultras — two 512 GB, two 256 GB — over RDMA reports a 235B mixture at 8 bits going from 19.5 tokens per second on one Mac to 26.2 on two and 31.9 on four, and a 671B mixture from 21.1 to 27.8 to 32.5, with link latency "from 300μs down to < 50μs"; the framework's own README claims "up to 1.8x speedup on 2 devices and 3.2x speedup on 4", which the charts did not reach: 1.3× at two Macs, 1.6× at four. The checklist adds a flag the chart did not: the 671B's 37B active parameters at 8 bits are 39 GB per token, a 21 tokens-per-second ceiling on one Mac's 819 GB/s, and the single-Mac figure is 21.1 — at the ceiling, above the 638 a kernel actually streams. So the row is either not 8 bits, not decode, or more than one token per pass, and the source does not say which. A January 2026 community run of five 512 GB Studios on a 1T-class mixture at 4 bits reports 14.45 tokens per second in a five-Mac pipeline, 14.49 in a four-Mac one and 14.82 with tensor parallelism: "only 2.3% difference", which is to say the fifth Mac bought nothing and the split bought 2%.
The Household's Verdict, and the Arithmetic Behind It
The household ran its own pooling experiments before this quest and stated the result at the plan gate as a ruling: not practical; tying the whole fleet together to load a 1 TB model does not mean it is usable; not the time. This quest documents RDMA and does not run it. The arithmetic agrees with the ruling on both counts. Capacity: two Studios reach 996 GB of working set, which admits a checkpoint between 498 and 996 GB — and the store holds none in that band that is worth the trouble; GLM-5.3 fits one Mac, Kimi K3 fits neither. Speed: a model split across the cable pays the wires lesson's per-layer crossing, and the public numbers put the whole gain at 1.3–1.6×, on models whose single-Mac rate was already reading speed. A 1 TB model that loads across two Studios and decodes at a few tokens per second is a demonstration, which is the last lesson's subject. The feature is real and the fleet is ready for it; the models that would justify it are not here, and the wire is the reason they would disappoint if they were.