Where local vision wins
- Privacy is paramount. Medical images, legal documents, surveillance footage, employee data — these often cannot go to the cloud. Local is the only option.
- Bulk batches. "Process 50,000 product photos" is free locally, expensive on the cloud.
- Offline. Field work, air-gapped sites, edge deployments.
- Latency-sensitive UI. Local vision can answer in 1–2 seconds; cloud round-trips add 3–5+ seconds.
Where cloud still wins
- Hardest reasoning over images. "Read this diagram and explain the architecture" — Claude 3.5 Sonnet and GPT-4V are still ahead.
- Very high resolution. Cloud models handle 4K+ images more gracefully.
- Spatial reasoning at the frontier. Counting, occlusion, relative position — cloud is better.
The Pippa pattern
cwkPippa uses local vision (Gemma 3 / Qwen 2.5-VL) for routine image tasks (avatar generation prompts, screenshot debugging, OCR of receipts) and falls back to Claude vision for hardest cases. Local-first with cloud-fallback applies to vision identically to text.