What is a 'concurrency limit' in a serverless compute platform, and what happens to incoming requests when a function's executions hit that limit at the same moment?
answer
- concurrency != requests/sec
- duration x rate ~= concurrent slots
- often a shared account-wide pool
- sync throttles immediately, async retries/queues
- downstream connection storms
basics
~10 sIt's a cap on how many copies of your function can run at once. If more requests arrive than the cap allows, the extra ones get rejected or queued/retried instead of running immediately.
solid answer
~50 sConcurrency is the number of function instances executing at the same moment, not requests per second — roughly, concurrency equals requests-per-second times average duration in seconds. Serverless platforms enforce a concurrency ceiling, often shared across an entire account or region rather than per function. When simultaneous executions would exceed that ceiling, synchronous callers (like HTTP clients) typically get an immediate throttling error such as HTTP 429 and must retry themselves, while asynchronous invocations (from a queue or event source) are usually buffered and retried automatically by the platform with backoff, eventually going to a failure destination if retries are exhausted. Because the limit is often shared account-wide, one function's traffic spike can throttle unrelated functions with no direct connection to the spike, and can also overwhelm downstream systems like databases that weren't designed for that many simultaneous connections.
go deeper
Can state that there's a cap on simultaneous executions and that going over it causes errors or delays.
Can explain the relationship between duration, request rate, and concurrency, and knows sync throttling looks different from async queuing/retry.
Proactively designs for downstream protection (connection pooling, concurrency caps, backpressure) and anticipates noisy-neighbor effects within a shared account limit.
Sets concurrency policy across an org's account/region to balance burst capacity against blast radius, ensuring one team's spike can't silently starve another team's critical path.
## Concurrency versus throughput Concurrency in a serverless platform is the count of function instances that are **actively executing at a single instant in time** — it is fundamentally different from throughput (requests per second). - A function invoked 1,000 times per second that finishes each invocation in 50ms only needs about 50 concurrent execution slots at any given moment (roughly requests-per-second times average duration, a relationship similar to **Little's Law** from queueing theory). - The same 1,000-requests-per-second rate with a 2-second average duration needs roughly 2,000 concurrent slots. Every serverless platform enforces some ceiling on this concurrent-execution count, and critically, that ceiling is very often shared across an entire account or region rather than allocated per function by default — every function without special configuration draws from the same finite pool. ## Why the ceiling exists This limit exists for two protective reasons. 1. First, it protects the platform's own **physical and orchestration capacity**: unbounded, instantaneous scale-out for every customer simultaneously isn't something any provider's underlying infrastructure can guarantee without some ceiling. 2. Second, and just as important in practice, it protects **whatever your function talks to downstream**. A serverless platform can scale out concurrent executions far faster than a traditional fixed-size application server fleet ever could, so without a ceiling, a burst of traffic could spin up thousands of concurrent function instances almost instantly, each potentially opening its own connection to a database, calling a third-party API, or writing to a shared resource — capacity that those downstream systems were never sized to absorb. ## What happens when the limit is hit What happens when the ceiling is hit differs by invocation type. | Invocation type | Platform behaviour | |---|---| | **Synchronous** — a client making an HTTP request and waiting for a response | Once the concurrency ceiling is reached, new requests are typically rejected immediately with a throttling error (commonly modeled as an HTTP 429 Too Many Requests), and it's the caller's responsibility to retry, ideally with backoff and jitter to avoid making the situation worse. | | **Asynchronous** — a function triggered by a queue message, a storage event, or a scheduled rule | The platform usually doesn't drop the work; instead it buffers the event and retries the invocation automatically after a delay, following a retry policy, and after enough failed retries routes the event to a dead-letter queue or configured failure destination so it isn't silently lost. | ## Failure modes The most important production trade-off is the **shared-pool effect**: because the concurrency ceiling is often account- or region-wide, one function's misbehavior or legitimate success can throttle every other function in that scope, even ones with no logical relationship to the spike. - **A bug that re-invokes itself.** A classic failure mode is a function with a bug that causes it to recursively re-invoke itself (for example, an error handler that republishes to the very queue that triggers it); within minutes it can consume the entire shared concurrency pool, and unrelated production functions — a payment processor, an authentication check — start failing with throttling errors even though nothing about their own code or traffic changed. - **The downstream-connection-storm.** A second common failure mode is the downstream-connection-storm: a burst of concurrent function executions each opens its own connection to a relational database sized for a handful of long-lived connections from a traditional app server; the database's connection pool exhausts almost instantly, and the outage looks like a database problem when its root cause is unmanaged serverless fan-out. ## Where it shows up A concrete real-world scenario is an event-driven pipeline where a serverless function is triggered by messages arriving on a queue after a bulk file upload. If nothing constrains that function's concurrency, a single large upload can drive it to scale out to thousands of simultaneous executions almost immediately, each one hitting the same downstream relational database to record results — a pattern platform teams commonly mitigate by deliberately capping that function's concurrency (via a **reserved concurrency** setting) to a number the database and its connection pool can actually sustain, deliberately trading maximum theoretical throughput for downstream stability, and accepting that some messages will simply take longer to process during a burst rather than taking the database down entirely.
- How is concurrent execution count related to request rate and average function duration?Roughly, concurrency equals requests-per-second multiplied by average duration in seconds, a relationship similar to Little's Law: longer-running functions consume more concurrency slots per unit of throughput than fast ones handling the same request rate.
- Why can a spike in a serverless function's concurrency cause a relational database to fall over even though the database itself never saw a request-rate spike from a single client before?Each concurrent execution environment often opens its own database connection, so N concurrent function instances can mean N simultaneous connections hitting a database whose connection pool was sized for a handful of long-lived app-server connections, exhausting it almost instantly; this is typically mitigated with a connection-pooling proxy layer or by capping the function's concurrency.
- What's the difference between throttling on a synchronous invocation versus an asynchronous one when the concurrency limit is hit?Sync callers get an immediate error response they must handle and retry themselves, while async invocations are automatically retried by the platform with backoff and eventually routed to a dead-letter queue or failure destination if retries are exhausted, so the work isn't necessarily lost, just delayed.
Like a restaurant with a fixed number of tables: it's not about how many customers arrive per hour, but how many are seated (occupying a table) at the same instant; once every table is full, new customers wait at the door or are turned away.
saying these in an interview costs you the question
- Equates the concurrency limit with a requests-per-second limit
- Believes raising memory or timeout settings increases the concurrency limit
- Doesn't know concurrency limits are often shared account-wide across all functions
- Assumes serverless scales infinitely with no need to protect downstream systems
- Unaware that concurrent executions can exhaust a database's connection pool