Quantization formats, calibrated
| Format | Bits | Speed | Quality | Notes |
|---|---|---|---|---|
fp16 / bf16 | 16 | baseline | baseline | What unquantized "base" usually means in 2026. |
fp8 | 8 | 1.3-1.5x | ~baseline | Hopper / Ada GPUs only. Requires a fp8-aware engine. |
bitsandbytes (8/4-bit) | 8 / 4 | 1-1.5x | ~baseline / minor | Fastest path to "fits on smaller GPU." Less optimal for serving. |
GPTQ | 4 / 3 | 2x | small drop | Per-channel quantization. Variant set on Hub: {model}-GPTQ. |
AWQ | 4 | 2-3x | small drop | Activation-aware quantization. Currently the strongest 4-bit serving option. |
GGUF | 2-8 | n/a | varies | llama.cpp / Ollama format. Not for TGI / vLLM. Covered in ops track. |
The decision rule
Production serve target on modern GPU: AWQ-4bit if a variant exists, else fp8 if your GPU supports it, else bnb-nf4 as a fallback. Avoid GPTQ in 2026 unless you already have it — it's not strictly worse, but AWQ has eaten its lunch.