Your live feature keeps hitting OpenAI 429s at peak — how do you fix it?
answer
- Retries redistribute pain, do not remove it
- Control admission, not just failure
- One budget shared across all replicas
- Offline work belongs off the hot path
- Separate projects, separate buckets
basics
~20 sStop relying on retries and add admission control: a shared token-bucket limiter sized to the tightest limit, so work waits in your queue instead of bouncing off the API. Then reduce demand — realistic completion caps, smaller models for cheap traffic, bulk work moved to the Batch API.
solid answer
~50 sChronic 429s mean demand shape exceeds capacity, and retries only redistribute the pain. First put a limiter *in front* of the call: a shared token bucket, backed by something like Redis so all replicas draw on one budget, sized to the tightest dimension (usually TPM) and fed an estimated token cost per request. Then cut demand: trim prompts and history, set completion caps to what tasks really need so each call reserves less token budget, and route cheap traffic to a smaller model with its own bucket. Move anything offline — backfills, embeddings, bulk scoring — to the Batch API, which draws on a separate enqueued-token allowance and costs less. Isolate traffic classes into separate projects so a backfill cannot starve interactive users. Finally, raise the usage tier if the business genuinely needs the throughput, and monitor `x-ratelimit-remaining-tokens` so you see headroom shrinking before users do.
go deeper
Know that repeated 429s mean you are asking for more than your allowance and that the answer involves slowing down and sending smaller requests, not just retrying harder.
Describe a token-bucket limiter sized to TPM, realistic completion caps, and moving bulk work to the Batch API, and explain why each reduces rejections.
Demonstrate lived judgment: diagnosing which dimension is binding from headers, sharing the limiter across replicas, bounding queues with backpressure, and defining explicit degradation when capacity runs out.
Own the capacity architecture: isolating traffic classes into separate projects, deciding which workloads justify flagship throughput, and planning tier headroom against forecast growth rather than reacting to incidents.
## Diagnose before you tune The first question is *which* budget you are exhausting, because the fixes differ. Split your 429 metric by model and by error code, and export `x-ratelimit-remaining-requests` and `x-ratelimit-remaining-tokens` from response headers as gauges. Typical findings: - **TPM exhausted while RPM is idle** — you are sending large prompts. The lever is prompt size and completion caps, not concurrency. - **RPM exhausted with small prompts** — you are chatty. The lever is batching work into fewer calls or smoothing bursts. - **429s clustered in the first seconds of each minute** — you have a scheduler firing everything on a cron boundary. The lever is jitter and pacing. - **`insufficient_quota` mixed in** — this is not a rate-limit problem at all; billing is empty. ## Admission control beats retrying Retry is what you do *after* being rejected; a limiter is what stops you being rejected. Implement a token bucket that refills at your allowance and requires each request to acquire an estimated cost before it goes out. Two details make or break it: 1. **It must be shared.** If every one of ten pods runs a local limiter sized to the full TPM, you have provisioned ten times your budget. Centralise the counter (Redis, or a dedicated gateway service) or divide the allowance by replica count and accept the waste. 2. **It must be fed a token estimate.** Count the request locally before admission, using the tokenizer plus the completion cap, so large requests consume proportionally more of the bucket than small ones. Behind the limiter, queue with **backpressure**, not an unbounded buffer. A bounded queue that sheds or degrades when full keeps latency honest; an unbounded one converts a capacity shortfall into a slow, invisible failure and eventually an out-of-memory incident. ## Reduce the demand itself - **Right-size completion caps.** The pre-flight token estimate counts your requested cap, so an inflated cap consumes budget you never use. This is the highest-leverage one-line change on most services. - **Shrink the prompt.** Retrieve fewer, better chunks; trim conversation history; move rarely-needed instructions out of the system prompt. Every token removed is TPM back. - **Exploit the prompt cache.** A stable prefix repeated across requests is billed at a discounted input rate; putting the volatile portion last maximises the hit rate. - **Tier your models.** Classification, routing and extraction rarely need the flagship. Smaller models are cheaper *and* draw on their own separate bucket, so moving traffic across models moves it across capacity pools. - **Batch semantically.** Where the task allows, ask one call to handle several items instead of issuing one call per item, converting RPM pressure into a modest TPM increase. ## Move offline work off the hot path The Batch API accepts a file of requests, returns results within a completion window measured in hours, is billed at a discount, and — the operationally decisive part — consumes a **separate enqueued-token allowance** rather than your per-minute budget. Nightly embedding refreshes, historical backfills, bulk evaluation and offline classification belong there. Anything a user is waiting on does not. ## Isolate traffic classes Rate limits are scoped per project as well as per model. Putting interactive traffic and background pipelines in separate projects gives each its own bucket, so a runaway backfill degrades itself rather than your product. This is the same bulkhead reasoning as separate connection pools for a database, and it is much easier to arrange before an incident than during one. ## Raise the ceiling — deliberately Allowances scale with usage tier, which advances automatically as cumulative spend and account age grow. If the workload is legitimate and growing, moving up a tier is the honest answer. But raise the ceiling *after* the demand-side work, not instead of it: a service that wastes half its TPM on inflated completion caps will waste half of a larger allowance too, at a higher bill. ## Degrade rather than fail Decide in advance what happens when capacity truly runs out: serve a cached or precomputed answer, fall back to a smaller model, queue the request and notify the user, or return a clear message. An interactive path should have a deadline measured in what the user will tolerate, and when that deadline passes, retries should stop. The teams that handle peak well are not the ones that never hit a limit — they are the ones whose behaviour at the limit was designed rather than discovered.
- Why is a per-pod local rate limiter dangerous when you autoscale?Each pod enforces the full allowance locally, so N pods collectively admit N times the budget and the API rejects the surplus — and the problem worsens exactly when you scale up to handle load. Either centralise the counter so all replicas draw from one bucket, or divide the allowance by the replica count and accept the unused headroom. Autoscaling must feed into whichever choice you make.
- Which workloads should move to the Batch API, and what do you give up?Anything offline and delay-tolerant: embedding refreshes, historical backfills, bulk classification, evaluation runs. You give up latency — results arrive within a completion window measured in hours, not seconds — and interactive streaming. In exchange you get a discounted rate and, operationally most valuable, a separate enqueued-token allowance that leaves your per-minute budget for live traffic.
- How do you spot rate-limit trouble before users do?Export the x-ratelimit-remaining-requests and remaining-tokens headers as gauges and alert on sustained low headroom rather than on 429 counts, which only fire after the fact. Pair that with tokens-per-request as a tracked metric: a prompt template or tool catalogue that quietly grows shows up there days before it exhausts the budget at peak.
saying these in an interview costs you the question
- Answering only with more retries and longer backoff
- Running a full-allowance limiter independently in every replica
- Unbounded queues that hide the capacity shortfall
- Sending backfills through the same project as live traffic
- Raising the tier before fixing inflated completion caps