C.W.K.
Stream
Lesson 03 of 05 · published

The Hybrid Zoo — Griffin, Hymba, Bamba, Zamba2, SAMBA

~14 min · hybrid-zoo, griffin, hymba, bamba

Level 0Observer
0 XP0/50 lessons0/14 achievements
0/100 XP to next level100 XP to go0% complete

Independent convergence

Once Jamba shipped, multiple teams independently converged on hybrid architectures with their own twists. The diversity of approaches is itself the signal — there's no single "right" hybrid; there are many plausible designs that all work. Here's a tour of the most important ones in 2026:

Griffin / RecurrentGemma (Google, 2024)

Combines a real-gated linear recurrent unit (RG-LRU) with sliding window attention. About 6K tok/s vs. ~2K for Gemma at comparable scale. RecurrentGemma is the open-weights distillation; the architecture is one of Google's productionized alternatives. The recurrent component here is RWKV-flavored more than Mamba-flavored.

Hymba (NVIDIA, 2024)

The most architecturally novel hybrid. Instead of alternating SSM and attention layers, Hymba runs them in parallel within the same layer — attention heads and Mamba heads in parallel, fused at the output. About 11.67× cache reduction. The parallel fusion approach is interesting because it gets per-layer benefits from both operations rather than sequencing them.

Bamba (IBM, 2024)

9B model, 29 Mamba2 layers + 3 attention layers. 2.5× throughput over a comparable Transformer. Notable for day-0 vLLM support — the integration with serving infrastructure was ready at release. That's the kind of detail that matters more than benchmark numbers for adoption.

Zamba2 (Zyphra, 2024)

A shared-attention-backbone design with Mamba2 blocks. 4× faster generation. Zyphra's approach is interesting because the attention backbone is shared across the stack, which is a different way to balance recall and efficiency than pure layer alternation.

SAMBA

Equal parts Mamba + sliding window attention. SAMBA is interesting as a counter-data-point: it's a 1:1 ratio rather than 1:7, and works well in its specific configuration. Reminder that the optimal ratio depends on the design space — SAMBA's sliding-window flavor of attention can do more per-layer than full attention.

External links

Exercise

Pick three hybrid architectures from this list (Jamba, Hymba, Bamba, Zamba2, Griffin, SAMBA) and read each paper's architecture section. Write a one-paragraph summary for each that describes (a) what they're hybridizing, (b) how the operations are combined (alternating vs parallel vs shared backbone), (c) what claim about quality/efficiency they make, and (d) what production status they have. The point is to develop a felt sense of the design space, not to pick a winner.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign inPlease sign in to comment.

No comments yet — be the first.