Different models pick different head counts based on their d_model and design philosophy. The pattern: d_head has stabilized at 64 or 128, while d_model and head count grow together.
| Model | d_model | Q heads | KV heads | d_head |
|---|---|---|---|---|
| Transformer Base (2017) | 512 | 8 | 8 | 64 |
| BERT-base | 768 | 12 | 12 | 64 |
| GPT-2 | 768 | 12 | 12 | 64 |
| GPT-3 | 12,288 | 96 | 96 | 128 |
| Llama 3 (8B) | 4,096 | 32 | 8 (GQA) | 128 |
| Llama 3.3 (70B) | 8,192 | 64 | 8 (GQA) | 128 |
| Mixtral 8×22B | 6,144 | 48 | 8 (GQA) | 128 |
| Qwen 2.5-7B | 3,584 | 28 | 4 (GQA) | 128 |
The trend is unmistakable: modern models have d_head=128 and use GQA to keep KV heads small. The decoupling of Q heads (representational capacity) from KV heads (cache memory) is one of the most consequential design moves of the past three years.