skip to content

Batching & Concurrency Control

You learn to push throughput without tripping provider limits: batch what can wait, bound how many calls are in flight, and back off politely on 429s. Under burst, a queue with timeouts and load shedding beats an unbounded fan-out every time.

on this pageshow

questions

4

What does a 429 from an LLM provider mean, and why is an immediate retry wrong?

level: juniorimportance: must knowfreq 62%

answer

  1. a quota signal, not a bug
  2. two ceilings, not one
  3. the response tells you how long
  4. double the wait, then randomise it
  5. synchronised retries rebuild the spike

basics

~20 s

A 429 means you crossed the provider's rate limit, usually a per-minute cap on requests or on tokens. Retrying instantly adds load to an already-throttled account; wait, honour any Retry-After, and back off exponentially with random jitter.

solid answer

~50 s

A 429 is a throttling signal, not a server bug: the account has exceeded a per-minute ceiling. LLM providers typically enforce **two** ceilings — requests per minute and tokens per minute, sometimes counting input and output tokens separately — so you can be far under the request cap and still be throttled on tokens. The correct response is to slow down. If the response carries a `Retry-After` value, honour it; otherwise sleep for an exponentially growing interval (0.5s, 1s, 2s, …) with **random jitter**, so that every client throttled in the same second does not retry in the same later second and rebuild the spike. Cap the number of attempts and the total wait against the caller's deadline — past that point the honest answer is to shed the request, not to keep retrying. Adding more workers or more connections while being throttled makes the situation strictly worse.

code

python · 9 lines
python
import random

def backoff_seconds(attempt, retry_after=None, base=0.5, cap=30.0):
    if retry_after is not None:
        return max(float(retry_after), base)
    ceiling = min(cap, base * (2 ** attempt))
    return random.uniform(0.0, ceiling)

print([round(backoff_seconds(n), 2) for n in range(5)])

go deeper

for a junior

Be able to say that 429 means rate-limited rather than broken, that the request usually needs no change, and that the client should wait and retry with a growing, randomised delay rather than looping immediately.

for a middle

Explain the mechanics: separate request and token ceilings, honouring Retry-After, exponential growth with jitter, and an attempt cap. Be ready to say why a synchronised retry after a fixed delay recreates the burst.

for a senior

Show you treat 429 rate as a monitored signal. Talk about tying retry envelopes to the caller's deadline, distinguishing transient overage from structural overage, and fixing the offered load upstream rather than retrying harder.

for a principal

Own the policy: one shared retry helper across services, quota consumption metered as a first-class metric, and an explicit decision about which workloads may retry long and which must fail fast when the account is saturated.

## What a 429 actually says HTTP 429 ("Too Many Requests") from a model provider means your account sent more work than its allowance for the current window. It is not a failure of the request itself: the same request, sent a few seconds later, will usually succeed unchanged. That distinction matters, because it tells you the fix is *pacing*, not *repair*. Nothing about the prompt needs to change; the rate at which you are offering work does. ## The two ceilings The most common surprise for people new to LLM APIs is that staying under the documented requests-per-minute number is not enough. Providers meter **tokens** as well as **requests**, typically as a tokens-per-minute allowance, and often with input and output tokens tracked separately. A workload that sends 20 requests per minute but attaches a 60,000-token document to each is a heavy consumer even though the request count is trivial. Two further details bite in practice. First, some providers reserve budget against the request's *declared maximum output length* rather than the tokens actually produced, so a habit of setting a very large maximum output on every call burns quota you never use. Second, when extended-thinking or reasoning modes are enabled, the tokens spent thinking count as output, and they can be several times the visible answer — a workload that fit comfortably yesterday can start throttling the moment reasoning effort is raised, with no change to request volume. ## Read the response before guessing A well-behaved client inspects the 429 before deciding what to do. The standard `Retry-After` header, when present, is the provider telling you exactly how long to wait; guessing a shorter interval simply earns another 429. Providers also commonly return headers describing remaining request and token quota and when the window resets, which lets a client throttle *before* it trips the limit rather than discovering the wall by hitting it. ## Exponential backoff and why jitter matters The standard client-side algorithm is exponential backoff: wait a base interval, then double it on each successive failure, up to a cap. This gives the provider's window time to reset and stops a struggling client from hammering. Backoff alone is not enough at scale, because throttling is a *correlated* event. If fifty in-flight calls are all rejected in the same second and every one of them waits exactly two seconds, they all return in the same later second and the spike reproduces itself — a retry storm that can keep an account throttled long after the original burst ended. Randomising the wait breaks the synchronisation. The common form, often called full jitter, picks a uniformly random delay between zero and the current exponential ceiling; the effect is to smear a synchronised herd across the whole interval. ## What a retry cannot fix Retrying is the right response to a *transient* overage. It is the wrong response to a structural one. If a single request is itself larger than the per-minute token allowance, no amount of backoff helps — that request will fail forever and needs to be split or shortened. If steady-state demand simply exceeds the account's quota, retries convert an obvious throttle into a slow, expensive, invisible one: latency climbs, work piles up in memory, and the real signal ("we need more quota, fewer calls, or smaller prompts") never reaches anyone. ## Where retries stop and backpressure starts Every retry policy needs two bounds: a maximum number of attempts and a deadline. A user-facing call with a two-second budget should not be retried three times over eight seconds — by then the answer is worthless, and the polite thing is to fail fast and let the caller decide. Bulk or offline work can afford a much longer retry envelope precisely because nobody is waiting. Deciding which class a call belongs to is part of designing the client, not an afterthought. Retries are also a symptom, not a strategy. A client that is regularly retrying is a client that is offering work faster than its allowance permits, and the durable fix lives upstream: bound how many calls are in flight, pace token consumption against the known ceiling, and shed or defer work that cannot be served. Backoff is the shock absorber; it is not the suspension. ## A practical shape Wrap every provider call in one retry helper rather than scattering ad-hoc loops. The helper should: classify the error (throttling versus a bad request versus a transient network fault — only the first two categories behave differently), honour `Retry-After` when present, otherwise sleep a jittered exponential interval, cap attempts, respect the caller's deadline, and emit a metric for every throttled attempt. That last point is easy to skip and the most useful in an incident: the count of 429s per minute is the earliest signal that a workload has outgrown its quota.

  • Your client backs off correctly but every attempt still returns 429. What would you look at?
    That pattern points at a structural overage rather than a burst. Check whether you are hitting the token ceiling rather than the request ceiling, whether a single request exceeds the per-minute token allowance on its own, whether an oversized declared maximum output is reserving quota you never use, and whether reasoning or thinking tokens have inflated output. If steady demand simply exceeds quota, no backoff policy will help — reduce concurrency, shrink prompts, or raise the limit.
  • Why add random jitter instead of just doubling the delay?
    Because throttling hits many callers at once. If all of them wait the same doubled interval, they retry in the same later second and recreate the burst that caused the throttle — a self-sustaining retry storm. Randomising each wait within the current backoff ceiling spreads the herd across the interval, so the offered load decays smoothly instead of oscillating.
  • Should a client retry every 429 until it eventually succeeds?
    No. Every retry policy needs an attempt cap and a deadline tied to the caller's tolerance. A user-facing call with a two-second budget should fail fast rather than retry into irrelevance; offline work can afford a long envelope. Unbounded retrying hides an underlying capacity problem, inflates latency, and turns a visible throttle into invisible queueing.

saying these in an interview costs you the question

  • Treats 429 as a server bug and retries in a tight loop
  • Assumes only request count matters, ignoring the token-per-minute ceiling
  • Adds more workers or connections to push through the throttling
  • Retries every failed call after the same fixed delay
  • Ignores Retry-After and invents a shorter wait

context

open as a page

When does an LLM provider's batch API beat synchronous calls for bulk work?

level: middleimportance: should knowfreq 50%

basics

~20 s

Use it when the job has a deadline but no single item needs a fast answer. Batch endpoints trade a completion window of up to a day for roughly half price and much higher throughput off the provider's spare capacity.

open as a page

Why does a per-process semaphore fail to bound LLM API concurrency fleet-wide?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Because the provider meters the whole account while the semaphore counts one process. A cap of 40 in-flight calls becomes 400 the moment ten replicas run, and the fleet trips the account limit even though every instance is behaving.

open as a page

How do you split one LLM token budget across workloads with different deadlines?

level: principalimportance: should knowfreq 36%

basics

~20 s

Give each workload a deadline class, reserve headroom for the interactive one, and let the rest queue. Under burst, shed or defer the deadline-tolerant classes explicitly rather than letting one shared budget starve the path users are waiting on.

open as a page