How do you work out how many concurrent executions an AWS Lambda function needs, and what does Lambda do to a synchronous invocation once the account's concurrency limit is reached?
answer
- capacity is in-flight, not per second
- Little's Law
- duration multiplies capacity
- rejected, not queued
- 429 for the synchronous caller
basics
~20 sConcurrency equals invocation rate multiplied by average duration in seconds: 500 requests per second at 200 ms needs about 100 concurrent executions. Beyond the account's per-Region limit Lambda throttles, and a synchronous caller gets HTTP 429 with TooManyRequestsException.
solid answer
~50 sLambda's unit of capacity is the concurrent execution — one environment handling one request. The steady-state number you need is Little's Law: concurrency equals requests per second times average duration in seconds. Five hundred requests a second at 200 ms is roughly 100; the same rate at two seconds is 1,000. That is why cutting duration is a capacity lever, not just a latency one. Every account has a per-Region quota on total concurrent executions across all functions — 1,000 by default, raisable via a Service Quotas request. When you hit it, Lambda rejects the invocation rather than queueing it: a synchronous caller such as `Invoke` with `RequestResponse` gets an HTTP 429 with `TooManyRequestsException`, and the `Throttles` CloudWatch metric increments. Asynchronous and poller-driven invocations aren't surfaced to a caller at all — Lambda retries them on its own schedule. Watch `ConcurrentExecutions` against the quota, not invocation count.
code
bash · 12 lines# Current per-Region account limit and how much is in use right now
aws lambda get-account-settings \
--query 'AccountLimit.ConcurrentExecutions'
# Peak concurrency over the last hour, one-minute resolution
aws cloudwatch get-metric-statistics \
--namespace AWS/Lambda \
--metric-name ConcurrentExecutions \
--start-time "$(date -u -v-1H +%Y-%m-%dT%H:%M:%SZ)" \
--end-time "$(date -u +%Y-%m-%dT%H:%M:%SZ)" \
--period 60 \
--statistics Maximumgo deeper
Be able to state the formula — concurrency equals requests per second times average duration in seconds — and to work a simple example out loud without hesitating.
Explain that the quota is per Region per account and shared across functions, that a synchronous caller sees a 429 TooManyRequestsException, and that the Throttles metric is where you see it.
Show you use the formula operationally: alarm on Throttles and on ConcurrentExecutions Maximum, treat a Duration regression as a capacity event, and decide when a queue in front is the right buffer.
Own the account-level allocation: who gets reserved slices of the shared pool, how much headroom the platform keeps, and whether latency-critical workloads should be isolated into their own account entirely.
## Concurrency, not requests per second The question "how much Lambda do I need?" is not answered in requests per second, because Lambda does not meter capacity that way. One execution environment serves exactly one invocation at a time, so the quantity that matters is **how many are in flight simultaneously**. That is Little's Law: ``` concurrency = invocations per second x average duration in seconds ``` Work the examples until they're reflexive: - 100 rps, 50 ms average -> 100 x 0.05 = **5** concurrent executions. - 500 rps, 200 ms -> 500 x 0.2 = **100**. - 500 rps, 2 s -> 500 x 2 = **1,000** — same traffic, twenty times the capacity. - 10 rps, 60 s (a slow report generator) -> **600**, from traffic that looks trivial. The last two are the interviewer's point: **duration is a capacity multiplier**. A function that waits three seconds on a downstream call consumes concurrency for the whole wait, doing nothing. Removing that wait (or moving it off the request path) frees capacity in exact proportion. This formula gives the *average*. Real traffic is bursty, so provision against your peak and remember that duration is a distribution: if p99 is ten times the mean, your concurrency peaks are far above the average calculation. ## The account quota AWS enforces a per-Region, per-account limit on total concurrent executions summed across every function in that Region. The default is 1,000 and it is raisable through Service Quotas. You can read the current value from the API: ```bash aws lambda get-account-settings --query 'AccountLimit.ConcurrentExecutions' ``` Two things about this quota trip people up. It is **shared across all functions in the Region**, so a runaway batch function can consume the pool your latency-critical API depends on — the noisy-neighbour problem that reserved concurrency exists to solve. And it is **per Region**, so a multi-Region deployment has separate pools that must be raised separately. ## What throttling actually looks like When a new invocation would push concurrency past the applicable limit, Lambda does not queue it. It rejects it, and *how* you observe that depends entirely on the invocation path: - **Synchronous** (`RequestResponse`, an API Gateway or ALB integration, a direct `Invoke` call): the caller receives an HTTP **429** with the error type `TooManyRequestsException`. The client must decide whether to retry; the AWS SDKs retry throttling errors with exponential backoff by default, which can itself amplify the storm if the backoff is too aggressive. - **Asynchronous and event-source-driven** invocations are not surfaced to a caller — Lambda holds and retries them on its own schedule, so throttling shows up as *delay* rather than an error, which is why it often goes unnoticed until latency alarms fire. Either way the **`Throttles`** CloudWatch metric increments. It is the single most important Lambda metric to alarm on that most teams forget, because a throttled invocation produces **no** `Errors` datapoint and no function logs — from the function's own telemetry, the request simply never happened. ## Reading the right metrics - **`ConcurrentExecutions`** — the number in flight, sampled per minute. Compare its **Maximum** statistic against your quota; the Average will lie to you about spikes. - **`UnreservedConcurrentExecutions`** — how much of the account pool is left after every reserved allocation. - **`Throttles`** — rejections. Any non-zero value on a user-facing function deserves an alarm. - **`Duration`** — the other half of the formula; a regression here silently raises your concurrency needs. ## Designing with the number Once you can compute concurrency, several design decisions become arithmetic rather than intuition. Sizing the headroom you need before a launch. Deciding whether a synchronous API can absorb a spike or needs a queue in front to convert a burst of requests into a bounded stream of work. Predicting the load your function will put on a downstream database, since peak connections track peak concurrency, not request rate. And, in the other direction, understanding that a quota increase is not a fix for a function whose duration has quietly tripled — that is a capacity leak, and the formula tells you where it went.
- Your function's average duration doubles after a dependency slows down. What happens to concurrency?It doubles for the same request rate, because concurrency is rate times duration. If that pushes you into the account quota you start throttling even though traffic never grew. This is why a downstream slowdown often presents as Lambda throttling: alarm on Duration as a leading indicator, and keep the slow dependency off the synchronous path where you can.
- A throttled invocation produces no entry in the function's logs. Why, and what do you alarm on instead?Throttling happens before the invocation reaches an execution environment, so no handler runs and nothing is written to that function's log group — and the Errors metric stays at zero. Alarm on the Throttles metric, and watch ConcurrentExecutions using the Maximum statistic against your quota so you see the pressure building before rejections start.
- How does putting an SQS queue between the client and the function change the throttling picture?It converts a synchronous rejection into buffering. Requests are accepted into the queue and the function drains them at whatever concurrency it can get, so a burst becomes latency rather than 429s. The tradeoffs are that the caller now gets an acknowledgement rather than a result, and you need to watch queue age so backlog is visible.
saying these in an interview costs you the question
- Sizes Lambda in requests per second instead of concurrency
- Thinks Lambda queues synchronous requests when the limit is hit
- Believes the concurrency quota is per function by default
- Assumes throttles show up as errors in the function's logs
- Treats a quota increase as the fix for a duration regression