You're designing cost and capacity governance for an organization running dozens of serverless functions that share a single account-level concurrency limit and several downstream dependencies with their own capacity ceilings (e.g., a relational database, a rate-limited third-party API). What's your approach to preventing one function's traffic pattern from silently degrading cost or reliability for the rest of the account?
answer
- shared pool = finite governance resource
- reserve floor for critical, cap ceiling for risky
- protect downstream, not just the platform limit
- connection-pooling proxy for DB fan-out
- review reservations against real peak traffic
basics
~10 sGive critical functions a guaranteed slice of shared capacity, cap risky/bursty ones so they can't eat everything, protect fragile downstream systems like databases with pooling or their own limits, and monitor the shared pool.
solid answer
~50 sTreat the account's total concurrency limit as a finite, shared budget requiring active governance, not a platform detail to ignore. Give critical, unpredictable-traffic functions a reserved concurrency floor sized against observed peak (not average) traffic, so they can't be starved by noisy neighbors. Give bursty or fan-out-prone functions (queue consumers, event-storm handlers) an explicit reserved concurrency ceiling so their worst case is bounded rather than able to consume the whole shared pool. Separately and explicitly protect downstream dependencies — a connection-pooling proxy in front of a database, rate-limit-aware clients for third-party APIs — since the platform's own concurrency limit protects the platform, not dependencies that can be overwhelmed well before that limit is reached. Add per-function cost and throttle-rate observability, and revisit reservations as traffic patterns evolve rather than treating the initial configuration as permanent.
go deeper
Not expected to design this; should recognize that many functions sharing one account can affect each other.
Can describe reserved concurrency as a way to protect one function from others, without necessarily designing full account-wide governance.
Can design reservation policy for their own service and knows to protect a specific downstream dependency, such as via connection pooling, but may not yet think at whole-account or org governance scale.
Designs and owns account or org-wide concurrency and cost governance across many teams' functions, balancing reliability, cost, and downstream dependency capacity as a continuously reviewed system, not a static configuration.
## The shared pool is a finite budget This is a governance problem sitting at the intersection of the shared-resource nature of serverless concurrency and the real cost and reliability **blast radius** that follows from it, so the approach has to operate at both the platform-configuration layer and the organizational-process layer simultaneously. Start from the fact that the account or region's total concurrency ceiling is a fixed, finite budget — treat it exactly like a shared physical resource in capacity planning, not an implementation detail each team is free to ignore. ## Classify every function, then size the reservations Classify every function along two axes: - **Criticality** — does throttling this function cause real user or business impact? - **Traffic predictability** — steady versus bursty versus driven by an upstream event source that can spike unpredictably, like a bucket receiving a bulk upload or a queue backlog draining all at once. From those two axes: - **Critical, low-predictability functions** get an explicit reserved concurrency **floor** sized against observed peak traffic with real headroom, not average traffic, so they're never starved by a noisy neighbor elsewhere in the account. - **Non-critical or clearly bursty, fan-out-prone functions** — batch processors, event-storm consumers, anything triggered by user-uploaded content — get an explicit reserved concurrency **ceiling** that deliberately caps their maximum draw on the shared pool, even if that means they queue or backpressure under extreme load; the goal is making their worst case bounded and predictable rather than 'as much as the account will give me.' - **The remaining unreserved pool** is kept with real headroom, not driven toward zero by over-reservation elsewhere, so newly deployed or lower-priority functions aren't throttled by default the moment they ship. ## Why this governance is necessary Why this governance is necessary at all: without it, a shared concurrency pool creates a classic multi-tenant blast-radius problem — one team's bug, such as an infinite retry loop, or one team's unplanned success, such as a feature going viral, can throttle every other function in the account, including ones entirely unrelated to the incident, simply because they all silently draw from the same finite pool. The identical logic applies downstream: a serverless function can scale out far faster than a traditional connection-pooled application server ever could, so an un-throttled burst of concurrent executions each opening a database connection can exhaust a connection-limited database in seconds, or blow through a third-party API's rate limit and get an account's API key temporarily banned — failures that concurrency limits inside the serverless platform alone don't prevent, because the platform's limit protects the platform's own capacity, not your specific downstream dependency's capacity. ## The trade-off, in both directions The trade-offs run in both directions, with cost named on each side. | Approach | What you get | The catch | |---|---|---| | **Aggressive reservation and capping** | protects reliability and bounds worst-case cost | but it costs engineering overhead — someone has to classify functions, size reservations against real observed traffic data, and revisit those numbers as traffic evolves — and it risks over-conservative caps causing self-inflicted throttling during a legitimate spike nobody anticipated | | **Looser, unmanaged concurrency** | avoids that governance overhead entirely | but it leaves both cost and reliability exposed to whatever any single function's worst day happens to look like, with no floor and no ceiling protecting anyone | ## Failure modes - **A marketing-triggered burst.** A concrete production failure mode: a marketing-triggered burst — a push notification, a viral post — drives a spike into a webhook-processing function with no reserved ceiling; it scales out to thousands of concurrent executions, each opening a connection to a shared relational database sized for a handful of long-lived application-server connections, and the database falls over, taking down every other service that depends on it, not just the function that spiked. - **A runaway recursive-invocation bug.** A separate failure mode, where a function republishes to the same queue that triggers it on error, consumes the entire account's unreserved concurrency pool within minutes, silently throttling unrelated production functions with no direct connection to the bug — and the on-call team initially has no idea why an unrelated payment function suddenly started failing, because the root cause is nowhere near the symptom. ## Where it shows up A concrete real-world mitigation pattern that addresses both governance layers at once: pairing serverless functions that talk to a relational database with a **connection-pooling proxy** layer, so that N concurrent function executions share a small, bounded pool of actual database connections instead of each opening its own, combined with a hard reserved-concurrency ceiling on that function sized to what the proxy and database can actually sustain under load testing — turning an unbounded fan-out risk into a deliberately bounded, tested one. That ceiling is then reviewed on a recurring cadence, whenever the database is resized or the function's traffic pattern materially changes, rather than being set once at launch and forgotten — governance, in this domain, is a **living configuration** tied to observed reality, not a one-time setup step.
- Why isn't a serverless platform's own concurrency limit sufficient to protect a downstream database from an unbounded fan-out spike?The platform's limit protects the platform's own shared capacity, not the specific capacity of a downstream dependency; a concurrency limit set higher than what the database can actually sustain still lets a spike overwhelm the database well before the platform-level limit is even reached, so the database needs its own explicit protection such as connection pooling or a tightly sized reserved concurrency ceiling on the calling function.
- How would you detect that reserved concurrency reservations across an account have drifted into an unhealthy state?Track and alert on the aggregate sum of reserved concurrency across all functions versus the account's total limit, alongside per-function throttle-rate metrics; a rising throttle rate on functions without reservations, or an unreserved pool shrinking toward zero as new reservations are added, is the concrete signal that reservations need rebalancing.
- When would you deliberately choose NOT to give a critical function reserved concurrency?When the function's downstream dependency has effectively unbounded capacity of its own and its traffic is genuinely unpredictable in a way that makes any fixed reservation likely to be wrong in one direction or the other — in that case, a well-provisioned unreserved pool with strong monitoring can be more robust than a static number that will inevitably drift wrong as traffic evolves.
Like a shared office building's total electrical circuit capacity: you don't let any one floor draw unlimited power, because a runaway space heater on one floor can trip the breaker for the whole building; instead you allocate guaranteed circuits to critical systems, cap discretionary ones, and size the wiring for what the building — and the transformer feeding it — can actually handle.
saying these in an interview costs you the question
- Proposes raising the account-wide concurrency limit as the only fix
- Doesn't distinguish protecting the platform from protecting a specific downstream dependency
- Treats concurrency reservation sizing as a one-time setup rather than continuously reviewed
- Has no answer for how one function's bug can throttle unrelated functions in the same account
- Ignores connection-pooling/proxy patterns when discussing database protection