Quiz · 5 questions
⚛️ What an LLM Asks of Hardware
Two workloads, the decode ceiling, a cache that grows while you talk, the honest benchmark, experts, a batch of one, and quantization
Level 0Spec-Sheet Skimmer
0 XP0/91 lessons0/19 achievements
0/100 XP to next level100 XP to go0% complete
Quiz
01Why is decode bandwidth-bound while prefill is compute-bound, for the same model on the same machine?
02Which bytes count toward 'bytes per token' for decode on a Qwen3.5 4-bit checkpoint?
03Llama-3.2-1B and Qwen3.5-9B both carry 32 KB of KV cache per token, yet from 10% to 90% of a 32K window the 1B lost 39% of its decode speed and the 9B lost 17%. Why?
04The 35B-A3B mixture-of-experts model reads about 1.66 GB per token, giving a ceiling near 385 tokens per second on office at the 638 GB/s a kernel streams. It measured 89–110 across the fleet. What does the track conclude?
05A 4-bit checkpoint of a 27B model reads 14.4 GB per token. What did quantization buy, and what does the label on the quality cost say?
Comments 0
🔔 Reply notifications (sign in)Sign in — Please sign in to comment.
No comments yet — be the first.