skip to content

Why does a per-process semaphore fail to bound LLM API concurrency fleet-wide?

level: seniorimportance: should knowfreq 45%

answer

  1. the meter is the account, not the process
  2. multiply the cap by replica count
  3. autoscaling scales into the wall
  4. permits are not tokens
  5. discover the limit instead of declaring it

basics

~20 s

Because the provider meters the whole account while the semaphore counts one process. A cap of 40 in-flight calls becomes 400 the moment ten replicas run, and the fleet trips the account limit even though every instance is behaving.

solid answer

~50 s

A semaphore bounds the process that holds it, not the account the provider bills. Set it to 40 in-flight calls, deploy ten pods, and the provider sees roughly 400 — the fleet blows through the account-wide ceiling while every single instance is inside its configured limit, so nothing local looks wrong. Autoscaling makes it worse: throttling raises latency, latency raises queue depth, and the autoscaler adds pods, which adds concurrency. Three fixes exist. Divide the budget by replica count (simple, but wastes capacity under uneven load and breaks silently on a scaling change). Meter centrally, through a shared limiter or an internal gateway that all calls pass through, so the budget is account-shaped. Or adapt: drive concurrency from observed throttling, increasing it slowly and cutting it sharply on a 429. Concurrency is also a poor proxy for the token ceiling, since per-request token cost varies enormously.

code

python · 11 lines
python
import asyncio

MAX_IN_FLIGHT = 40
_sem = asyncio.Semaphore(MAX_IN_FLIGHT)

async def call_model(prompt, send):
    async with _sem:
        return await send(prompt)

async def run(prompts, send):
    return await asyncio.gather(*(call_model(p, send) for p in prompts))

go deeper

for a junior

Know that a semaphore limits only the process it lives in, and that running several copies of a service multiplies the concurrency the provider actually sees, even though no configuration changed.

for a middle

Explain the arithmetic and the fix options: divide the budget by replica count, meter centrally through a shared limiter, or adapt from observed throttling. Be ready to say why concurrency is only a rough proxy for a token ceiling.

for a senior

Show the diagnosis. Describe the autoscaling feedback loop, the per-pod-looks-fine signature, sizing with Little's Law, and what you monitor at account level to catch it before an incident.

for a principal

Own where the bound lives across the organisation: a shared gateway or limiter, priority classes with reserved shares, an explicit fail-open or fail-closed choice when the limiter is unreachable, and quota changes that do not require redeploying every caller.

## The semaphore that looked like a limit The idiom is everywhere: acquire a semaphore before each model call, release it after, size it to something that felt safe in load testing. It works beautifully on one machine. The failure appears at the deployment boundary, because the semaphore is a property of the *process* and the rate limit is a property of the *account*. A team sets a cap of 40 in-flight calls, tests it against a single instance, sees clean headroom, and ships. Later, someone adds a second worker pod for redundancy, then autoscaling takes the deployment to ten. Nothing in the application changed and no configuration was touched, but the provider now sees roughly 400 concurrent calls. Throttling starts, and the confusing part is that every instance's own metrics look healthy: each is at or below its configured 40. The limit that was violated is not visible from inside any one process. ## The autoscaling feedback loop This failure has a nasty amplifier. Throttling raises per-request latency (calls now wait through backoff). Higher latency means requests spend longer in flight and in local queues. Queue depth and latency are exactly the signals an autoscaler is usually configured to react to, so it adds replicas — and each new replica brings another full semaphore's worth of concurrency to an account that is already over its ceiling. The system scales *into* the wall. Any fleet whose scaling signal is downstream latency needs its provider concurrency governed outside the replica, or it will do this. ## Concurrency is not the quota anyway Even a correctly shared concurrency cap only approximates the thing being metered. Providers meter tokens per minute as well as requests, and per-request token cost varies by orders of magnitude across a real workload: a 200-token classification and a 60,000-token document summary occupy one semaphore permit each. Forty in-flight summaries can exceed a token ceiling that forty in-flight classifications would not approach. Extended-thinking modes widen the gap further, because reasoning tokens count as output and can multiply an answer's token cost several times over with no change in request count. A concurrency cap tuned when a workload ran at low reasoning effort silently becomes far too generous when the effort dial is raised. If tokens are what is metered, tokens are what you should pace against; concurrency is a convenient, coarse proxy that needs a margin. ## Three ways to make the bound real **Static division.** Give each replica `total_budget / replica_count`. It is trivial and needs no shared state, and it is genuinely fine for a fixed-size deployment. Its weaknesses are that capacity is stranded whenever load is uneven (an idle pod's share is unusable by a busy one) and that the divisor is a hidden coupling: a scaling change that nobody thinks of as a rate-limit change quietly invalidates it. **Central metering.** Route model calls through one place — a shared limiter backed by a common store, or an internal gateway every service calls — so the budget is enforced at account shape. This is the version that actually holds under autoscaling, and it brings a second benefit: one place to observe consumption, apply per-workload priority, and change limits without redeploying callers. The costs are a network hop, a dependency that must itself be highly available, and a decision about what happens when the limiter is unreachable (fail open and risk throttling, or fail closed and stop all model traffic). **Adaptive concurrency.** Rather than deriving the cap from a published number, discover it: increase the in-flight allowance gradually while calls succeed, and cut it sharply when a 429 arrives — the classic additive-increase / multiplicative-decrease shape. This handles quota changes, noisy neighbours inside the same account, and heterogeneous request sizes without anyone maintaining a constant. It is more code, it needs a floor so the system does not collapse to one, and it must be per-account rather than per-pod for the same reason the static cap was. These compose. A common production shape is central metering for the account budget, plus a local adaptive cap so an instance backs off promptly on throttling instead of relying on a round trip. ## Sizing, when you must pick a number Little's Law gives the starting point: concurrency ≈ throughput × average latency. To sustain 20 requests per second against calls averaging 4 seconds you need about 80 in flight. Do that arithmetic per workload, then sanity-check it against the token ceiling using the workload's mean tokens per request, and take the lower of the two. Remember that latency here includes provider-side thinking time, which is the term most likely to move. ## What to watch Three signals separate this failure from every other latency problem: throttled responses per minute at the *account* level, concurrency observed in aggregate rather than per pod, and the ratio between the two. A fleet whose per-pod concurrency looks healthy while account-level throttling climbs is showing you exactly this bug, and the fix is to move the bound to where the meter is.

  • What is wrong with simply dividing the account budget by the replica count?
    Nothing, for a fixed deployment — it is cheap and needs no shared state. It fails in two ways: capacity is stranded when load is uneven, because an idle replica's share cannot be used by a busy one, and the divisor is an invisible coupling, so any change to replica count silently invalidates the limit without anyone treating it as a rate-limit change. Autoscaled fleets make both problems permanent.
  • Why is a fixed in-flight cap a poor proxy for a tokens-per-minute ceiling?
    Because one permit can represent wildly different token cost. A short classification and a 60,000-token summarisation each occupy one slot, so the same concurrency produces very different token rates depending on workload mix. Extended-thinking modes widen the gap, since reasoning tokens count as output. If tokens are metered, pace on estimated tokens and keep concurrency as a coarse secondary guard with margin.
  • How would an adaptive concurrency controller decide the limit at runtime?
    Additive increase, multiplicative decrease: raise the in-flight allowance by a small step while calls succeed, and cut it by a large factor as soon as throttling appears, with a floor so the system never collapses to zero throughput. It discovers the real limit rather than trusting a constant, absorbs quota changes and noisy neighbours in the same account, and must be scoped per account rather than per process.
  • Which metrics would show you this failure before customers do?
    Account-level throttled responses per minute, aggregate in-flight concurrency across all replicas, and the observed token rate against the ceiling. The signature is per-pod concurrency sitting comfortably under its configured cap while account-level throttling climbs — a combination that only makes sense when the bound is in the wrong place. Alarm on the aggregate, never on the per-instance view.

saying these in an interview costs you the question

  • Assumes a process-local semaphore bounds the account's concurrency
  • Raises the per-pod cap when throttling appears
  • Thinks extra API keys multiply the account's quota
  • Treats in-flight request count as equivalent to token rate
  • Lets an autoscaler add replicas in response to throttling-induced latency

context