Why local at all
Cloud APIs are easy. Local AI is different: privacy, cost, latency, offline, and freedom from vendor terms. Pick local when one of those properties is load-bearing for the system you're building. If none of them matter, the cloud will out-quality you for less effort. The decision is operational, not ideological.
Five concrete reasons local wins
- Privacy. Patient images, internal contracts, journal entries, and unreleased product code never leave the machine. No vendor logging, no future training-data risk.
- Cost. Zero marginal cost per token after hardware. Million-token batch jobs that would cost real money on the cloud are free locally.
- Latency floor. First-token latency drops below 100 ms on a warm model on Apple Silicon. No round-trip variance.
- Offline / outage tolerance. Plane, rural site, AWS incident, vendor account suspension — local keeps working.
- No rate limit, no key rotation. Run the model as hard as the hardware allows.
What you give up
Frontier quality. A 70B local model is excellent, but Claude Opus, GPT-5, and Gemini 3 are still ahead on hardest reasoning. You also own the operational surface: model updates, hardware sizing, quantization choice, KV cache budget. The right pattern is almost never only local — it's local-first with cloud fallback for the cases the local model can't handle.