Skip to content
C.W.K.
Stream
Quiz · 6 questions

🔬 Predict, Then Measure

The quest's own lab: the number written down before the run, four Macs of one household, and the gaps explained rather than excused

Level 0Spec-Sheet Skimmer
0 XP0/91 lessons0/19 achievements
0/100 XP to next level100 XP to go0% complete

Quiz

01Why does the lab compute a predicted decode ceiling from the safetensors headers before running a model?
02Over the dense models of one machine, seconds per token fit a straight line in bytes per token. What do the slope and intercept mean?
03On the 9B, the M3 Ultra has 8.2× the base M3's bandwidth and 8× its GPU cores. Measured prefill was 6.1× and decode 4.9×. What is the right reading?
04Why does the household's doctrine say the only benchmark that matters is decode and TTFT at 90% of the context window?
05Music (M2 Ultra, 800 GB/s, 76 cores) and office (M3 Ultra, 819 GB/s, 80 cores) ran the same ladder. Which stage did the newer chip win, and which did it lose?
06A neighbour quest on this site said, until its repair on 2026-09-15, that an M3 Ultra decodes a 70B INT4 model at about 95 tokens per second. What does the checker say?
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign inPlease sign in to comment.

No comments yet — be the first.