OpenAI enforces rate limits at two levels: Requests Per Minute (RPM) and Tokens Per Minute (TPM). Your limits increase automatically as you move through usage tiers based on cumulative spend.
Tier System
| Tier | Qualification | Monthly Limit |
|---|---|---|
| Free | Allowed geography | $100 |
| Tier 1 | $5 paid | $100 |
| Tier 2 | $50 paid + 7 days | $500 |
| Tier 3 | $100 paid + 7 days | $1,000 |
| Tier 4 | $250 paid + 14 days | $5,000 |
| Tier 5 | $1,000 paid + 30 days | $200,000 |
Rate Limit Headers
Every API response includes rate limit information:
The SDK auto-retries 429 responses with exponential backoff (2 retries by default). For manual control, parse the Retry-After header. Use max_completion_tokens as tight as possible — rate limit calculations use the estimated maximum.