In a serverless platform's concurrency model, what is the difference between 'reserved concurrency' (a dedicated slice of an account's shared concurrency pool set aside for one function) and 'unreserved concurrency' (the shared pool available to all other functions), and what problem does each solve?
answer
- reserved = floor + ceiling for one function
- unreserved = shared leftover pool
- reservation is not the same as pre-warming
- protects against noisy neighbors
- self-imposed throttling risk if under-sized
basics
~10 sReserved concurrency carves out a guaranteed, capped slice of the shared pool for one function only. Unreserved concurrency is the leftover shared pool every other function draws from.
solid answer
~50 sIn a shared concurrency pool, reserved concurrency dedicates a fixed slice exclusively to one function — it guarantees that function a minimum available capacity (protecting it from being starved by other functions) while simultaneously capping its maximum (protecting the rest of the account, and often a fragile downstream dependency, from that function scaling unbounded). Unreserved concurrency is what's left in the shared pool after all reservations are subtracted; any function without an explicit reservation draws from it, and can be throttled if other functions or its own traffic exhausts that shared remainder. Reserving too little for a genuinely bursty function causes self-inflicted throttling even when the account has spare capacity elsewhere; reserving too much across many functions starves the unreserved pool for everything else. It's distinct from provisioned concurrency, which pre-warms instances to avoid cold starts and costs extra even when idle.
go deeper
Knows reserved concurrency guarantees capacity for one function; doesn't need to reason about account-wide pool math.
Can explain both the guarantee and the ceiling effect of reserved concurrency, and distinguishes it from provisioned concurrency.
Designs reservation policy across multiple functions sharing an account, sizing reservations against peak traffic and downstream capacity, and recognizes self-throttling failure modes.
Sets org-wide concurrency governance — default reservations, alerting on shared-pool exhaustion, escalation paths for limit increases — so no single team's function can silently starve another team's critical path.
## The two pools | Pool | What it holds | |---|---| | **Unreserved** | A serverless platform's total concurrency limit at the account or region level is a single shared pool that, by default, every deployed function draws from without distinction — this is the unreserved pool. The remaining, un-reserved portion of the account's total limit stays in the shared pool for every function that has no explicit reservation. | | **Reserved** | Setting a reserved concurrency value of `N` on a specific function removes N slots from that shared pool and dedicates them exclusively to that function: it can now run up to N concurrent executions at any time, and even if every other function in the account is being throttled to zero, this function still has its guaranteed N available. | ## The two risks one knob addresses This mechanism exists because a purely shared pool creates two distinct risks that a single knob can address simultaneously. 1. **Noisy-neighbor starvation.** The first is noisy-neighbor starvation: in a shared pool, one high-traffic or buggy function — for instance, one caught in a retry loop — can consume the account's entire concurrency budget, throttling every unrelated function in that account, including business-critical ones that had nothing to do with the incident. Reserved concurrency solves this for critical functions by guaranteeing them a **floor** that cannot be taken away by anyone else's behavior. 2. **Uncontrolled fan-out.** The second risk is a function driven by a bursty event source (a bulk file upload, a backlogged queue) that can scale out to thousands of concurrent executions and overwhelm a downstream dependency, like a connection-limited database or a rate-limited external API, long before the account's overall limit is even reached. Reserved concurrency solves this by acting as a self-imposed **ceiling** — capping how far that specific function can scale regardless of how much traffic arrives, deliberately bounding its worst case. ## The trade-off The trade-off cuts both ways and is easy to get wrong in each direction. - **Setting a reservation too low.** Setting a reservation too low on a function that legitimately experiences bursty, unpredictable traffic causes self-inflicted throttling: the function appears broken during a real spike even though the account has spare, unused capacity sitting in the unreserved pool the whole time. - **Setting reservations too generously.** Setting reservations too generously across many functions, on the other hand, can starve the unreserved pool: since a platform typically won't allow the sum of all reservations to exceed the account's total limit, over-reserving leaves little or nothing for functions deployed without an explicit reservation, causing surprise throttling on anything new or unconfigured. - **What it does not do.** It's also important to note that a reservation doesn't pre-warm any execution environments, so a reserved function can still pay a cold-start penalty on its first invocation into a fresh slot — that's the job of a separate, additional, and separately billed feature (**provisioned concurrency**), which keeps a specified number of environments initialized and warm at all times, at an ongoing cost regardless of whether they're currently handling traffic. ## Failure modes - **A reservation based on average daily traffic.** A realistic production failure mode: a team sizes a checkout function's reserved concurrency based on its typical average daily traffic, not its peak; during a flash sale, the function throttles legitimate paying customers the moment it hits its own reservation ceiling, even though the account overall has plenty of unused headroom sitting idle in other functions' unreserved allocations — a self-inflicted outage caused by under-provisioning a guarantee that was also acting as a cap. - **A function left entirely unreserved.** The mirror-image failure is a debug or rarely used function left entirely unreserved with an uncontrolled recursive-invocation bug; it silently exhausts the shared unreserved pool and takes down unrelated production functions in the same account, with the on-call team initially having no idea why services with no connection to the buggy function suddenly started failing. ## Where it shows up A concrete, deliberate real-world pattern is platform teams explicitly setting a function's reserved concurrency to **zero** to disable it without deleting its code (blocking all invocations instantly for an incident or deprecation), and setting a modest, headroom-padded reservation on functions that write to a small, connection-limited database — using the reservation as a load-tested **backpressure valve**, not merely a leftover configuration detail.
- If a function has no reserved concurrency set at all, where does its concurrency come from, and what's the risk?It draws entirely from the shared unreserved pool along with every other unreserved function in the account or region; the risk is that any other function, or a spike in its own traffic, can exhaust that shared pool, throttling it with no guaranteed floor to fall back on.
- Does setting reserved concurrency reduce cold starts?No — it only guarantees and caps the number of concurrent execution slots available to the function; each new slot invoked for the first time still needs to cold-start unless the function is also configured with provisioned concurrency, a separate, paid feature that pre-warms environments in advance.
- What happens if the sum of every function's reserved concurrency in an account approaches the account's total concurrency limit?The unreserved pool shrinks toward zero, so any function without an explicit reservation, including newly deployed ones, gets throttled almost immediately under any load — a common self-inflicted outage when teams over-reserve without tracking the account-wide total.
Like reserving a block of parking spots for one department in a shared office garage: that department always has spots waiting (a guarantee), but they can never park more cars than the block size even if the rest of the garage is empty (a cap); everyone else fights over what's left.
saying these in an interview costs you the question
- Confuses reserved concurrency with provisioned concurrency (pre-warmed instances)
- Thinks a high reservation costs money even at zero traffic
- Doesn't realize reservations are subtracted from one shared account-wide total
- Assumes unreserved functions are unaffected by other functions' concurrency use
- Omits that a reservation is also a hard ceiling, not just a guarantee