"7 billion parameters" is an abstract number. Building intuition for what each scale means — capability, hardware, cost — is a real skill.
| Size | Examples | Capability tier | Hardware |
|---|---|---|---|
| 1-3B | Phi-3-mini, Llama 3.2 1B/3B, Gemma 3 1B | Basic tasks, mobile-friendly | Single GPU, smartphone |
| 7-8B | Llama 3 8B, Mistral 7B, Qwen 2.5-7B | Solid general capability | Single 16GB+ GPU |
| 13-14B | Phi-4, Gemma 3 12B | Strong reasoning, near-frontier on focused tasks | Single 24GB+ GPU |
| 27-32B | Gemma 3 27B, Qwen 2.5-32B | Approaches frontier quality on most tasks | 1-2 GPUs |
| 65-70B | Llama 3.3 70B | Frontier-quality dense model | 2-4 GPUs in FP16, 1 GPU in INT4 |
| 200-400B dense | Llama 3.1 405B | Top-tier quality | Cluster (8+ GPUs) |
| MoE 100-700B total | Mixtral 8×22B, DeepSeek-V3, Llama 4 | Top-tier quality, ~20-40B active | 4-8 GPUs (depending on quantization) |
The takeaway: sweet spots have moved over time. In 2023, 70B dense was the sweet spot for quality at reasonable cost. In 2026, MoE models with 17-40B active parameters often match or beat 70B dense on quality while serving cheaper. Picking the right point on this curve is half the deployment decision.