How much data you actually need
| Quality level | Examples | What you can expect |
|---|---|---|
| Minimum viable | 50–100 | Noticeable improvement on narrow tasks; risk of overfitting. |
| Good | 500–1,000 | Reliable performance on your domain. Most projects target this. |
| Strong | 1,000–10,000 | High-quality, consistent behavior across edge cases. |
| Production-grade | 10,000+ | Near-expert level on complex, multi-faceted tasks. |
Quality dominates quantity. 500 perfectly curated, diverse examples will outperform 50,000 noisy, repetitive ones every time. The single best place to spend an extra week is on data quality, not on hyperparameter sweeps.
The 2025–2026 cost landscape
| Path | Approximate cost | What you need |
|---|---|---|
| OpenAI managed fine-tuning | $0.80–$25 per 1M training tokens | Just data + an API key. |
| Google Colab (free tier) | $0 (limited GPU time) | T4 GPU (~15 GB VRAM). Fine for QLoRA on 7B models. |
| RunPod / Lambda Labs / vast.ai | $0.30–$3.00 per GPU-hour | A100 / H100 rental. Pay-as-you-go. |
| Your own consumer GPU | $700–$2,000 upfront | RTX 3090 / 4090 with 24 GB VRAM. |
| Apple Silicon (MLX) | $0 if you already own the Mac | M2/M3/M4 Ultra with 64 GB+ unified memory. |
Honest budget for a typical project
A first real fine-tuning project — Llama 3.1 8B with QLoRA on ~1,000 examples — costs roughly $0 (Colab free) to $10 (RunPod RTX 4090, ~3 hours). The expensive failure mode is not the GPU bill; it is the time you spend curating data, the iteration on hyperparameters, and the eval suite. Plan in days, not GPU-hours.