OpenAI's model catalog in 2026 spans frontier reasoning models, efficient mini/nano variants, and specialized models for images, audio, and realtime interactions. The primary API is the Responses API, though Chat Completions remains fully supported.
Flagship Models (GPT-5.x Series)
| Model | Context | Max Output | Input $/1M | Output $/1M | Key Feature |
|---|---|---|---|---|---|
gpt-5.4 | 1.05M | 128K | $2.50 | $15.00 | Flagship, reasoning none→xhigh |
gpt-5.4-mini | 400K | 128K | $0.75 | $4.50 | Fast & affordable |
gpt-5.4-nano | 400K | 128K | $0.20 | $1.25 | Ultra-lightweight |
GPT-4.x & o-Series
| Model | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
gpt-4.1 | 1M | $2.00 | $8.00 | Best non-reasoning model |
gpt-4o | 128K | $2.50 | $10.00 | Multimodal (text+vision) |
o3 | 200K | $2.00 | $8.00 | Advanced reasoning |
o4-mini | 200K | $1.10 | $4.40 | Efficient reasoning |
Key notes: gpt-5.4 defaults to reasoning_effort: "none" — set it to "low", "medium", "high", or "xhigh" to enable chain-of-thought reasoning. Prompts exceeding 272K input tokens are charged at 2× input and 1.5× output for the entire session. gpt-4.1 excels at instruction following and tool calling with its 1M token context window.