Mixtral — the model that opened the door
Mixtral 8×7B (December 2023, Mistral AI) was the first capable open-weight MoE the broader community could actually run. 46.7B total / 12.9B active, 8 experts top-2, no shared experts. It matched or beat Llama 2 70B on most benchmarks at roughly 3× lower per-token compute cost. Mixtral 8×22B followed with 141B / 39B. After 2024 Mistral has continued the line into Mistral Small 4 and Mistral Large 3.
DeepSeek — the fine-grained-expert design
DeepSeek pioneered fine-grained expert segmentation — many small experts instead of a few large ones. Where Mixtral has 8 experts (28 possible 2-expert combinations), DeepSeek-V2 has 160 experts top-6 (over 10¹⁵ possible combinations). This dramatically increases the routing expressiveness and lets experts specialize on much narrower patterns.
| Model | Total | Active | Experts | Top-K | Innovation |
|---|---|---|---|---|---|
| DeepSeek-V2 | 236B | 21B | 160 + 2 shared | top-6 | Fine-grained, MLA attention |
| DeepSeek-V3 | 671B | 37B | 256 + 1 shared | top-8 | Aux-loss-free balancing, FP8, sigmoid routing |
| DeepSeek-R1 | 671B | 37B | 256 + 1 shared | top-8 | Same arch as V3 + GRPO reasoning RL |
Llama 4 — Meta's MoE pivot
Llama 4 (April 2025) is Meta's first MoE family, and it goes top-1 instead of top-2 — every token picks exactly one expert.
| Model | Total | Active | Experts | Top-K | Context |
|---|---|---|---|---|---|
| Llama 4 Scout | 109B | 17B | 16 | top-1 | 10M |
| Llama 4 Maverick | 400B | 17B | 128 | top-1 | 1M |
| Llama 4 Behemoth | ~2T | 288B | — | — | preview (teacher) |
Qwen3 MoE — many experts, no shared
Qwen3 (2025) ships two MoE variants: Qwen3 30B-A3B (128 experts top-8, no shared) and Qwen3 235B-A22B (128 experts top-8, no shared). Both support dual thinking/non-thinking modes in a single checkpoint. Qwen3 235B-A22B is one of the most capable open-weight MoE models for self-hosting in 2025–2026.
The trend at a glance
Across families and years: experts get more numerous and smaller, top-K rises from 1–2 to 6–8, sigmoid routing replaces softmax, shared experts come and go. The MoE design space is still actively evolving, and the 2026 best practice may not be the 2027 best practice.