C.W.K.
Stream
Lesson 06 of 08 · published

Rate Limits

~22 min · rate-limits, tpm, rpm, backoff

Level 0Tokenizer
0 XP0/54 lessons0/10 achievements
0/120 XP to next level120 XP to go0% complete

OpenAI enforces rate limits at two levels: Requests Per Minute (RPM) and Tokens Per Minute (TPM). Your limits increase automatically as you move through usage tiers based on cumulative spend.

Tier System

TierQualificationMonthly Limit
FreeAllowed geography$100
Tier 1$5 paid$100
Tier 2$50 paid + 7 days$500
Tier 3$100 paid + 7 days$1,000
Tier 4$250 paid + 14 days$5,000
Tier 5$1,000 paid + 30 days$200,000

Rate Limit Headers

Every API response includes rate limit information:

The SDK auto-retries 429 responses with exponential backoff (2 retries by default). For manual control, parse the Retry-After header. Use max_completion_tokens as tight as possible — rate limit calculations use the estimated maximum.

Code

Reading rate-limit headers·text
x-ratelimit-limit-requests: 5000
x-ratelimit-limit-tokens: 5000000
x-ratelimit-remaining-requests: 4999
x-ratelimit-remaining-tokens: 4999979
x-ratelimit-reset-requests: 12ms
x-ratelimit-reset-tokens: 362ms

External links

Exercise

Write a small wrapper that retries on 429 with exponential backoff (base 1s, factor 2, jitter ±20%, max 5 retries). Honor the Retry-After header when present. Log the actual wait at each retry.

Progress

Progress is local-only — sign in to sync across devices.
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign inPlease sign in to comment.

No comments yet — be the first.