The shape of every modern Transformer is a four-tuple. Read these from any model's config and you know the architecture.
| Symbol | Meaning | Typical 2026 values |
|---|---|---|
| d_model | Hidden dimension. Width of every internal representation. | 768 – 12,288 |
| d_ff | FFN intermediate dimension. With SwiGLU, ≈ 8/3 × d_model. | 2,048 – 28,672 |
| n_heads (Q) | Number of query heads. | 12 – 96 |
| n_kv_heads | Number of KV heads (GQA). Often n_heads / 4 or /8. | 4 – 96 |
| n_layers | Number of stacked Transformer blocks. | 12 – 126 |
| d_head | = d_model / n_heads. Modern default 128. | 64 – 128 |
The original Transformer Base used d_model=512, d_ff=2048, n_heads=8, n_layers=6, d_head=64. Llama 3.3 70B uses d_model=8192, d_ff~28000, n_q=64, n_kv=8 (GQA), n_layers=80, d_head=128. Same architecture, ~1000× scale, different rectangles.