The KV cache is the dragon of long-context inference. Standard MHA stores one K and V per query head, scaling linearly with both context length and head count. Grouped-Query Attention (GQA) and its extreme cousin Multi-Query Attention (MQA) attack this directly by sharing K and V across multiple Q heads.
The shapes
| Variant | Q heads | KV heads | KV cache reduction | Used by |
|---|---|---|---|---|
| MHA | H | H | 1× (baseline) | BERT, GPT-3, GPT-2 |
| GQA | H | G < H | H/G × | Llama 3, Mistral, Gemma 3, Mixtral |
| MQA | H | 1 | H × | Early PaLM, Falcon |
Llama 3.3 70B has 64 query heads and 8 KV heads — a GQA group size of 8. Each group of 8 query heads shares one K and one V projection. The KV cache shrinks by 8x. Quality stays close to full MHA because GQA has more KV diversity than MQA but most of the cache savings.
Why this is now the default
Empirically, going from MHA to GQA with 8 KV heads loses very little quality but saves enormous memory at long context. MQA is even more aggressive but shows quality degradation in some tasks, so the field settled on GQA as the default. You can also uptrain an MHA checkpoint into GQA: average the KV heads within each group and continue training for ~5% of the original pretraining compute. Llama 2 70B-chat → GQA was done this way.