Independent convergence
Once Jamba shipped, multiple teams independently converged on hybrid architectures with their own twists. The diversity of approaches is itself the signal — there's no single "right" hybrid; there are many plausible designs that all work. Here's a tour of the most important ones in 2026:
Griffin / RecurrentGemma (Google, 2024)
Combines a real-gated linear recurrent unit (RG-LRU) with sliding window attention. About 6K tok/s vs. ~2K for Gemma at comparable scale. RecurrentGemma is the open-weights distillation; the architecture is one of Google's productionized alternatives. The recurrent component here is RWKV-flavored more than Mamba-flavored.
Hymba (NVIDIA, 2024)
The most architecturally novel hybrid. Instead of alternating SSM and attention layers, Hymba runs them in parallel within the same layer — attention heads and Mamba heads in parallel, fused at the output. About 11.67× cache reduction. The parallel fusion approach is interesting because it gets per-layer benefits from both operations rather than sequencing them.
Bamba (IBM, 2024)
9B model, 29 Mamba2 layers + 3 attention layers. 2.5× throughput over a comparable Transformer. Notable for day-0 vLLM support — the integration with serving infrastructure was ready at release. That's the kind of detail that matters more than benchmark numbers for adoption.
Zamba2 (Zyphra, 2024)
A shared-attention-backbone design with Mamba2 blocks. 4× faster generation. Zyphra's approach is interesting because the attention backbone is shared across the stack, which is a different way to balance recall and efficiency than pure layer alternation.
SAMBA
Equal parts Mamba + sliding window attention. SAMBA is interesting as a counter-data-point: it's a 1:1 ratio rather than 1:7, and works well in its specific configuration. Reminder that the optimal ratio depends on the design space — SAMBA's sliding-window flavor of attention can do more per-layer than full attention.