A synchronous AWS Lambda function behind an API starts returning throttling errors during a short traffic spike, even though account-wide concurrency stays far below the quota. What is happening, and how do you fix it?
answer
- two limits, not one
- level versus rate
- headroom is not the same as speed
- throttles on the leading edge
- warm floor, buffer, or backoff
basics
~20 sLambda limits how fast a function adds environments, not just how many it may have. A spike that outruns that ramp is throttled while headroom still exists. Fix it with pre-warmed provisioned concurrency, a queue to absorb the burst, or client backoff.
solid answer
~50 sConcurrency has two separate limits and this failure is the second one. The quota caps the *level* of concurrency; a separate scaling rate caps how fast a single function can add execution environments. As of the November 2023 scaling update, each function can add up to 1,000 concurrent executions every 10 seconds, independently of other functions. A spike that jumps from near zero to several thousand simultaneous requests in a second outruns that ramp, and the excess is throttled with 429s even though `ConcurrentExecutions` never approaches the account quota — which is exactly why the graph looks innocent. Diagnose it by overlaying `Throttles` against `ConcurrentExecutions` at one-minute resolution and looking at the shape: rate-limited scaling gives you throttles on the leading edge of a spike that clear within seconds. The fixes are to have capacity already warm (provisioned concurrency, scheduled ahead of known spikes), to buffer the burst behind a queue so it becomes latency instead of errors, or to make the caller retry with backoff and jitter.
go deeper
Know that Lambda can throttle even when the account limit is not reached, because new execution environments are added at a limited rate rather than instantly.
Explain the difference between a level limit and a rate limit, and state that scaling headroom is per function under the current model rather than a shared Region-wide burst pool.
Drive the diagnosis: one-minute Maximum metrics, the leading-edge throttle signature, ruling out reserved caps and duration regressions, then choose between pre-warming, buffering and client backoff with reasons.
Own the architectural call — decide which paths must be synchronous at all, where queues belong between tiers, and what the platform guarantees about spike absorption so teams are not each rediscovering this in production.
## Two different limits with the same word on them Most people learn one Lambda concurrency limit — the per-Region account quota — and assume throttling always means they hit it. There is a second, independent constraint: **how quickly a function may grow**. Standing up an execution environment is real work: allocate a microVM, fetch the code, start the runtime, run Init. AWS bounds the rate at which it will do that for any one function. As of the November 2023 change to Lambda scaling, each function scales by **up to 1,000 additional concurrent executions every 10 seconds**, and this ramp is **per function** rather than shared across the account. (Before that change, a Region-wide burst quota was shared by all functions, so one function's spike could throttle another's; the newer model removed that coupling.) So a function sitting at 50 concurrent executions cannot instantly serve 4,000 simultaneous requests, even in an account with a 10,000 quota and nothing else running. It climbs: ~1,050 in the first 10 seconds, ~2,050 in the next, and so on. Requests arriving above the current level during the climb are rejected with `TooManyRequestsException`. ## Why the dashboards mislead This failure mode is diagnostically nasty for three reasons. 1. **`ConcurrentExecutions` looks fine.** It might peak at 1,200 against a quota of 10,000, so the obvious "are we near the limit?" check says no. 2. **The function's own telemetry is silent.** Throttled invocations never reach a handler, so there are no logs, no `Errors`, no traces. The evidence exists only in the `Throttles` metric and in the caller's 429s. 3. **It is over in seconds.** By the time someone opens a dashboard the ramp has caught up and everything looks healthy. At five-minute metric resolution the spike may be invisible entirely — look at one-minute resolution, Maximum statistic. The distinguishing signature is *shape*: throttles clustered on the **leading edge** of a traffic step, decaying as concurrency climbs, with `ConcurrentExecutions` rising in a staircase rather than a vertical line. Sustained throttling at a flat concurrency ceiling means something else — you hit the account quota or the function's own reserved concurrency, which are level limits, not rate limits. ## Three real fixes **1. Have the capacity already warm.** Provisioned concurrency on the alias means those environments exist before the spike, so the ramp starts from a high floor instead of zero. For *predictable* spikes — a market open, a scheduled campaign, a batch kickoff — schedule the increase with Application Auto Scaling before the event rather than reacting to it; reactive target-tracking on `LambdaProvisionedConcurrencyUtilization` is always behind a step function by definition. **2. Buffer the burst.** If the work does not have to be synchronous, put a queue between the client and the function. The client's request is accepted immediately and the function drains the backlog at whatever rate it can scale to, converting a wall of 429s into a few seconds of extra end-to-end latency. This is the structurally correct answer whenever the caller does not need the result in the same round trip, and it also protects downstream systems from the same spike. **3. Retry properly.** A 429 is a *retryable* signal, and Lambda's ramp means a retry a second later has a good chance of landing. Exponential backoff **with jitter** is essential; synchronised retries from thousands of clients recreate the spike exactly and turn a brief throttle into a sustained one. Note that AWS SDK clients already retry throttling errors by default, so check whether your effective request rate is higher than you think. What is *not* a fix is requesting a quota increase. The quota governs the ceiling; you never reached it. Raising it changes nothing about how fast a single function can add environments. ## The neighbouring diagnoses to rule out Before concluding "scaling rate", check the alternatives, because the remedy differs: - **Reserved concurrency set too low** on the function — throttles at a flat, suspiciously round concurrency number. - **Account quota reached**, often because a *different* function in the same Region consumed the shared unreserved pool — check `UnreservedConcurrentExecutions`. - **Duration regression** — a slower dependency raises concurrency for identical traffic; overlay `Duration` with `ConcurrentExecutions` and the correlation is obvious. - **A downstream limit** being reported as a Lambda problem, where the function itself is fine and the errors come from what it calls. Getting this diagnosis right is most of the interview answer: the fix follows immediately once you can say *which* limit was hit and show the metric shape that proves it.
- How would you distinguish scaling-rate throttling from hitting a reserved concurrency cap on the metric graph?Look at the shape. Rate-limited scaling shows throttles concentrated on the leading edge of a traffic step, with ConcurrentExecutions climbing in a staircase and throttles clearing within seconds. A reserved concurrency cap shows sustained throttling while ConcurrentExecutions sits pinned flat at exactly the reserved number for as long as the load lasts.
- Why is requesting a concurrency quota increase the wrong response here?The account quota is a ceiling on how many concurrent executions may exist, and in this scenario you never approached it. The constraint is the rate at which one function may add execution environments, which the quota does not affect. Raising it costs a support round trip and changes nothing; warm capacity, buffering or backoff address the actual limit.
- For a spike you know is coming at a specific time, what is better than target-tracking auto scaling of provisioned concurrency?A scheduled scaling action that raises provisioned concurrency before the event. Target tracking reacts to observed utilisation, so against a step change it is always behind and the first wave still pays throttles or cold starts. Scheduling puts the environments in place ahead of the traffic and you scale back down afterwards to stop paying for idle capacity.
saying these in an interview costs you the question
- Assumes throttling always means the account quota was hit
- Requests a quota increase without checking actual concurrency
- Retries immediately without backoff or jitter
- Concludes the function is healthy because its logs show no errors
- Looks only at five-minute average metrics for a seconds-long spike