Output tokens cost more than input tokens
Gemini, like every other major LLM provider, charges asymmetrically: output tokens cost roughly 4–8× more than input tokens. Generation is autoregressive — the model runs a forward pass for every output token, so producing a 500-token reply takes 500× more compute than reading a 500-token prompt. Internalize this number and your cost intuition will line up with reality.
Gemini 2.5 pricing (per 1M tokens, as of mid-2026)
| Model | Input ≤ 200K | Input > 200K | Output ≤ 200K | Output > 200K | Cached input |
|---|---|---|---|---|---|
| 2.5 Pro | $1.25 | $2.50 | $10.00 | $15.00 | $0.125 |
| 2.5 Flash | $0.30 | $0.30 | $2.50 | $2.50 | $0.03 |
| 2.5 Flash-Lite | $0.10 | $0.10 | $0.40 | $0.40 | — |
Two things to notice in this table. First, Pro doubles the input rate at the 200K-token boundary. If your typical request hovers around 195K, you're one moderately-sized PDF away from 2× billing. Second, Flash-Lite is roughly 22.5× cheaper than Pro per output token. Routing 70% of traffic away from Pro is the single biggest cost lever you have.
Free tier, batch API, caching
Three tools shrink your bill that are mechanical, not creative:
- Free tier: 5–15 RPM and 100–1,000 RPD per model with 250K TPM shared. Enough for prototyping, never enough for production. EU/UK/CH excluded.
- Batch API: 50% off list price for offline jobs that don't need sub-minute turnaround (translations, embeddings backfill, summary pipelines).
- Context caching: ~90% reduction on the input tokens of cached content. Pay once to cache a 100K-token doc, then ask many questions against it for cents instead of dollars.