How big should your vocabulary be? The trend over the past five years has been clearly upward, and it's worth understanding why.
| Tokenizer / vocab | Size | Used by |
|---|---|---|
| BERT WordPiece | 30,522 | BERT, DistilBERT |
| Llama 1/2 SentencePiece | 32,000 | Llama 1, Llama 2 |
| tiktoken p50k_base | 50,256 | GPT-2, GPT-3 |
| tiktoken cl100k_base | 100,256 | GPT-3.5, GPT-4 |
| Llama 3 BPE | 128,000 | Llama 3, 3.1, 3.3 |
| tiktoken o200k_base | 200,019 | GPT-4o |
| tiktoken o200k_harmony | 201,088 | GPT-5 |
| Gemma 3 SentencePiece | 262,144 | Gemma 3 (140+ langs) |
The economics
Bigger vocabulary → fewer tokens per sentence → cheaper inference per "useful" message. The cost is a larger embedding matrix (vocab × d_model parameters) and a larger output projection. For a 70B model with d_model=8192, going from 32K to 200K vocab adds ~1.4B parameters — under 2% of total. Trivial cost for several percent token-count savings on every query forever after.
The other reason vocabularies grew: multilingual coverage. A 32K English-centric vocab tokenizes Korean or Chinese 3-4× less efficiently than English. Gemma 3's 262K vocab targeting 140+ languages closes most of that gap.