Why can an OpenAI call return 429 on TPM while RPM is barely used?
answer
- More than one budget applies at once
- Per model, per project buckets
- Big prompts, few requests
- The cap you ask for is reserved
- Capacity trickles back, not resets
basics
~20 sOpenAI enforces several limits at once — requests per minute, tokens per minute, and daily caps — per model and per project. Exceeding any single one returns 429, so a handful of very large prompts can exhaust the token budget while the request count stays trivial.
solid answer
~50 sRate limits are multi-dimensional and applied per model within a project: requests per minute (RPM), tokens per minute (TPM), requests per day (RPD), plus extra dimensions such as images per minute on image endpoints and a separate enqueued-token limit for the Batch API. You are throttled by whichever dimension you hit first, so ten requests carrying 50k-token prompts each can blow the TPM budget while RPM is nearly idle. The token side is estimated *before* generation from your input tokens plus the completion cap you requested, which means an oversized `max_completion_tokens` reserves headroom you may never spend. Capacity also refills continuously rather than resetting on the minute boundary, so firing your whole minute's allowance in one burst gets rejected even though the per-minute total would have fit. Limits scale by usage tier, which the account moves up automatically as cumulative spend and account age grow.
go deeper
Know that OpenAI limits both how many requests and how many tokens you may send per minute, and that either one can cause a 429.
Explain that budgets are per model per project across RPM, TPM and daily caps, that the token estimate includes the completion cap you request, and that capacity refills continuously rather than resetting each minute.
Show you monitor x-ratelimit-remaining-tokens as a leading indicator, shape burst traffic deliberately, and know which lever (completion caps, Batch, project split, tier) applies to which symptom.
Own capacity across the organisation: how projects are partitioned so one workload cannot starve another, which tier the business needs, and how limit headroom is planned against forecast traffic growth.
## Limits are a set, not a number OpenAI does not give you "a rate limit". It gives you several budgets that apply simultaneously, and rejects the request the moment *any one* of them is exhausted: - **RPM** — requests per minute. - **TPM** — tokens per minute. - **RPD** — requests per day (prominent on free and low tiers). - **Endpoint-specific dimensions** — for example images per minute on image generation, and audio-input limits on speech endpoints. - **A separate batch queue limit** — the number of tokens you may have enqueued in the Batch API at once, which does not consume your synchronous TPM. Crucially these budgets are scoped **per model, per project**. Heavy use of a small model does not eat the flagship model's allowance, and two projects under one organisation each carry their own buckets — which is exactly why isolating batch pipelines from interactive traffic into separate projects is a real mitigation, not just tidiness. ## Why TPM usually trips first For most LLM workloads, TPM is the binding constraint. A single retrieval-augmented request stuffing 30k tokens of context into the prompt consumes as much token budget as hundreds of short chat turns. So a service can sit at 5% of its RPM and be hard-throttled on TPM, which is baffling if you are only watching request counts on a dashboard. The second surprise is *when* tokens are counted. The limiter cannot know how long the answer will be before generating it, so it estimates the request's token cost up front from the input tokens **plus the completion cap you asked for**. If every call sets `max_completion_tokens` to a huge value "just in case", every call reserves that much of your TPM even when the model replies in forty tokens. Setting a realistic cap is therefore a rate-limit optimisation as well as a cost one. (Note `max_completion_tokens` is the current parameter for Chat Completions; the older `max_tokens` is deprecated and is not accepted by reasoning models.) ## Continuous replenishment, not a minute reset Budgets refill smoothly rather than resetting at the top of each minute. A 60 RPM allowance behaves approximately like one request per second of refill, with a small burst allowance — it does not mean "any 60 requests within a calendar minute are fine". Firing all 60 simultaneously exhausts the bucket and the surplus is rejected with 429, even though your total for the minute would have been legal. This is why smoothing traffic through a client-side limiter beats a naive per-minute counter, and why the arrival *shape* of your traffic matters as much as its volume. ## Reading your position Every response carries the current state of both budgets: - `x-ratelimit-limit-requests`, `x-ratelimit-remaining-requests`, `x-ratelimit-reset-requests` - `x-ratelimit-limit-tokens`, `x-ratelimit-remaining-tokens`, `x-ratelimit-reset-tokens` Exporting `remaining-tokens` as a gauge is the cheapest early-warning system available: you see headroom shrinking minutes before the first rejection, instead of learning about it from a 429 spike. The `reset` values are also the correct input to a first backoff interval. ## Usage tiers Allowances are not fixed per account — they scale with a **usage tier**. Free-tier keys get small allowances with tight daily caps; Tier 1 unlocks once the account has paid (historically a small amount, on the order of five dollars of credit), and higher tiers are granted automatically as cumulative spend accumulates and a minimum number of days have passed since first payment. Each tier raises RPM and TPM per model, sometimes by an order of magnitude. Two consequences matter operationally: a brand-new project key does **not** inherit the limits your mature project enjoys, which surprises teams during migrations; and published tier thresholds change, so treat the limits page in the dashboard as the source of truth rather than a number you memorised. ## Designing within them Once you see limits as a multi-dimensional, per-model, continuously-refilling budget, the design moves follow: - Shape traffic with a shared limiter sized to the tightest dimension, not the friendliest one. - Keep completion caps honest so reservations match reality. - Push bulk work to the Batch API, which draws on its own enqueued-token budget. - Split traffic classes across projects so a backfill cannot starve user-facing calls. - Spread across models where the workload allows, since each model has its own bucket. - Alert on remaining-token headroom, not only on 429 counts.
- How does setting max_completion_tokens to a large value affect rate limiting?The limiter estimates a request's token cost before generation from your input tokens plus the completion cap you requested, so an inflated cap reserves TPM headroom the call may never use. On a busy service that inflates effective token consumption and produces 429s well below your real usage. Set the cap to what the task plausibly needs; it protects both the token budget and the bill.
- You create a new project and its key immediately hits 429s. Why?Rate limits are scoped per project and per model, and allowances depend on the organisation's usage tier and how the project's own limits were configured. A new project does not automatically inherit the throughput your established project ran at, and owners can also set per-project caps below the organisation ceiling. Check the project's limits page for the specific model before assuming a platform incident.
- Does sending work through the Batch API consume your synchronous TPM?No. Batch jobs draw on a separate enqueued-token allowance rather than the per-minute budget your interactive traffic uses, which is why moving backfills and offline scoring to Batch relieves pressure on live paths. It also carries a discounted price and a completion window measured in hours, so it suits work that tolerates delay, not anything a user is waiting on.
saying these in an interview costs you the question
- Assuming one number, requests per minute, is the only limit
- Thinking limits reset at the top of each minute
- Setting a huge completion cap with no rate-limit consequence
- Believing limits are per API key rather than per project and model
- Expecting a new project to inherit an established project's throughput