Mistral AI, founded by ex-Meta and ex-Google researchers, focuses on efficient open-weight models. The lineage is short but consequential.
| Model | Total params | Active params | Architecture | Context |
|---|---|---|---|---|
| Mistral 7B (2023) | 7B | 7B | Dense + sliding-window attention | 32K |
| Mixtral 8×7B (2024) | 47B | 13B | MoE, top-2 of 8 experts | 32K |
| Mixtral 8×22B (2024) | 141B | 39B | MoE, top-2 of 8 experts | 64K |
| Mistral Small 3 (24B) | 24B | 24B | Dense | 32K |
| Mistral Large 3 (2024) | ~675B | ~41B | MoE | — |
| Mistral Small 4 (2025) | 119B | 6B | MoE, 128 experts top-4 | 256K |
Mixtral 8×22B specifics: 56 layers, d_model=6144, 48 Q heads with 8 KV heads (GQA), SwiGLU, RoPE, multilingual (English, French, Italian, German, Spanish), Apache 2.0 license. The 8×22B and Mistral Small 4 specifically demonstrated that aggressive MoE (low active / high total) could match much larger dense models.