What's actually in production
Mamba-family architectures are now in real production at multiple major labs:
AI21 Jamba 1.5 — 256K context, available on AWS Bedrock, Azure, GCP. The first Mamba-family model to hit major cloud marketplaces. Its 1.5 release was the moment hybrid-Mamba moved from research to enterprise procurement.
NVIDIA Nemotron-H 8B / 56B / 47B — 92% attention-replaced-by-Mamba-2. About 3× faster than Llama-3.1 70B at 65K context. The 56B was pretrained in FP8 on 6,144 H100s over 20T tokens — production scale by any measure.
IBM Granite 4.0 — 9:1 Mamba:attention ratio, 70%+ RAM reduction relative to comparable pure Transformers. IBM's enterprise customers see a smaller serving footprint at comparable quality.
TII Falcon Mamba 7B — beats Llama-3.1 8B on standard benchmarks while being a pure SSM (interesting precisely because it's pure — proves selectivity gets you most of the way at 7B scale).
Cartesia Llamba-8B — 12× throughput vs. its Llama 3.1 8B teacher, achieved by distilling a Transformer into a Mamba student. The distillation route is increasingly important: instead of training a Mamba from scratch, take a strong existing Transformer and convert it.
The honest limitations don't go away
Production validation doesn't repeal physics. Mamba-family models still have:
- Five-shot MMLU gap — pure SSMs noticeably underperform on few-shot in-context learning. Hybrids close most of this; pure SSMs do not.
- Asymmetry bias from nonlinear convolution — Mamba has a slight bias toward tokens earlier in the sequence due to how the selective scan accumulates information.
- Associative recall failures — structural, as established in the previous track.
- Narrow learning rate windows — production teams have to spend more on hyperparameter sweeps relative to Transformer recipes.
- Tooling gap — improving fast, but still measurably behind Transformer in 2026.
None of these are deal-breakers for the right workload. All of them are reasons to default to a Transformer (or a Transformer-heavy hybrid) when the workload doesn't specifically reward Mamba's strengths.