In AWS Lambda, what is the difference between reserved concurrency and provisioned concurrency, and when would you configure each?
answer
- one allocates, one pre-warms
- cap and floor in the same number
- zero is the off switch
- version or alias, never $LATEST
- billed while idle
basics
~20 sReserved concurrency partitions the account's concurrency pool: it both guarantees a function that many slots and caps it there, at no extra cost. Provisioned concurrency pre-initialises a number of environments so they answer without a cold start, and you pay for them hourly.
solid answer
~50 sThey solve different problems and are often confused because both take a number. Reserved concurrency is about **allocation**: setting it to 50 carves 50 slots out of the account's per-Region pool for that function, guaranteeing nobody else can take them and simultaneously preventing that function from ever using a 51st. It costs nothing and is how you stop a batch job from starving a customer-facing API, or how you cap a function whose downstream database has a connection ceiling. Setting it to zero throttles the function completely, which is the standard emergency stop. Provisioned concurrency is about **latency**: it keeps N environments initialised and waiting on a specific version or alias, so those requests skip the Init phase. You pay for that warm capacity by the hour whether it is used or not, and traffic beyond N still falls back to on-demand with normal cold starts. They compose: you can reserve capacity and provision a portion of it.
code
bash · 15 lines# Allocation: carve 200 slots out of the account pool for this function
aws lambda put-function-concurrency \
--function-name checkout-api \
--reserved-concurrent-executions 200
# Latency: keep 40 environments initialised on the 'live' alias
aws lambda put-provisioned-concurrency-config \
--function-name checkout-api \
--qualifier live \
--provisioned-concurrent-executions 40
# Emergency stop: throttle every invocation without deleting anything
aws lambda put-function-concurrency \
--function-name runaway-processor \
--reserved-concurrent-executions 0go deeper
Be able to state the one-line difference: reserved concurrency limits and guarantees how many slots a function gets, provisioned concurrency keeps environments pre-initialised so requests skip the cold start.
Explain that reserved concurrency is simultaneously a floor and a ceiling drawn from the shared account pool, that it costs nothing, and that provisioned concurrency attaches to a version or alias and is billed while idle.
Show the operational judgement: size a reservation from measured peak concurrency, protect a connection-limited database with a cap, tune provisioned capacity against utilisation and spillover metrics, and know that setting reserved to zero is the emergency stop.
Own the economics and the policy: when warm capacity is cheaper than re-architecting, how the account concurrency pool is allocated across teams, and whether latency-critical workloads deserve isolation in their own account.
## Two knobs, two problems Both settings take an integer and both mention "concurrency", which is why candidates blur them. Keep the two questions apart: - **Reserved concurrency** answers *how much of the account's pool may this function use, and how much is guaranteed to it?* - **Provisioned concurrency** answers *how many environments should already be warm before requests arrive?* One is about sharing a finite resource; the other is about paying to eliminate the Init phase. ## Reserved concurrency: a cap and a floor at once By default all functions in a Region draw from one shared account pool. Reserving concurrency for a function partitions that pool: the reserved amount is subtracted from the unreserved pool available to everyone else, and the function can never exceed its reservation. That dual nature is the key insight, and interviewers probe it: - **As a floor**, it protects a critical function from noisy neighbours. If a nightly batch job can burst to thousands of environments, your checkout API can still get its guaranteed 200. - **As a cap**, it protects everything *downstream*. A function that opens a database connection per environment must not scale to 1,000 environments against a database that accepts 200 connections; the reservation is how you enforce that. It equally protects a rate-limited third-party API from your own scale-out. Reserved concurrency costs nothing — it is bookkeeping, not capacity. Two practical constraints: you cannot reserve so much that the unreserved pool available to all other functions drops below 100, and a function *with* a reservation can no longer borrow from the unreserved pool at all, so under-reserving turns into throttling. Setting it to **zero** is a special and very useful case: every invocation is throttled, which effectively switches the function off without deleting it or unwiring its triggers. That is the standard containment move for a function misbehaving in production. ```bash aws lambda put-function-concurrency \ --function-name checkout-api \ --reserved-concurrent-executions 200 # emergency stop aws lambda put-function-concurrency \ --function-name runaway-processor \ --reserved-concurrent-executions 0 ``` ## Provisioned concurrency: buying away the Init phase Provisioned concurrency tells Lambda to create and initialise N execution environments *ahead of time* and hold them ready. Requests routed to them skip Init entirely — the handler runs immediately. Important properties: - It is configured on a **published version or an alias**, never on `$LATEST`, because Lambda must snapshot a fixed code/config combination to initialise. That forces a version-and-alias deployment discipline. - You are **billed for the provisioned amount for as long as it is configured**, whether or not any request arrives, plus the usual request and duration charges. Idle warm capacity is pure cost. - It is **not a cap**. Traffic beyond the provisioned amount still runs on on-demand environments, with ordinary cold starts. The `ProvisionedConcurrencySpilloverInvocations` metric counts exactly those, and `ProvisionedConcurrencyUtilization` tells you how much of what you bought is being used — the two metrics you tune against. - It can be **scheduled or auto-scaled** with Application Auto Scaling, either on a schedule for predictable daily shapes or with target tracking on the `LambdaProvisionedConcurrencyUtilization` predefined metric. - During a deploy, shifting an alias to a new version means the new version's provisioned environments must initialise before they are useful, so a naive cutover can reintroduce the cold starts you paid to remove. ```bash aws lambda put-provisioned-concurrency-config \ --function-name checkout-api \ --qualifier live \ --provisioned-concurrent-executions 40 ``` ## How they interact They are complementary, not alternatives. Provisioned concurrency is drawn from the function's reserved concurrency if one is set, and from the unreserved account pool otherwise; you cannot provision more than you have reserved. A realistic configuration for a latency-sensitive API is: reserve 200 so the function is isolated and bounded, and provision 40 of those to absorb the normal baseline warm, letting genuine spikes spill to on-demand. ## Choosing Reach for **reserved** when the problem is a *shared-resource* problem: noisy neighbours, a downstream connection ceiling, a third-party rate limit, or a blast radius you want bounded. Reach for **provisioned** when the problem is a *tail-latency* problem on a user-facing synchronous path, and you have measured that `Init Duration` is what is hurting p99 — and when the arithmetic works: warm capacity costs money continuously, so at very low traffic it can cost more than the requests themselves, and at very high steady traffic cold starts are already a small fraction of invocations. The uncomfortable middle — spiky, latency-sensitive, moderate volume — is where it earns its keep.
- You set reserved concurrency to 20 on a function that previously scaled freely. What can now go wrong?The reservation is also a hard ceiling, so any traffic needing a 21st concurrent execution is throttled — synchronous callers get 429s where they previously succeeded. The reserved slots also leave the shared unreserved pool, so other functions have less headroom. Size the reservation from measured peak concurrency, not average, and alarm on Throttles for that function afterwards.
- Why can provisioned concurrency not be configured on $LATEST?Lambda has to initialise environments against a fixed code and configuration combination, and $LATEST changes the moment you publish new code — the pre-warmed environments would instantly be stale. Provisioned concurrency is therefore set on a published version or an alias pointing at one, which is why using it pushes you toward a version-and-alias deployment workflow.
- Which metrics tell you whether the provisioned concurrency you bought is correctly sized?ProvisionedConcurrencyUtilization shows what fraction of the warm pool is actually serving requests — persistently low means you are paying for idle environments. ProvisionedConcurrencySpilloverInvocations counts requests that overflowed to on-demand and therefore may have paid a cold start; a steady non-zero value means the pool is too small for your baseline.
- How would you disable a runaway Lambda function immediately without deleting it?Set its reserved concurrency to zero with put-function-concurrency. Every invocation is throttled at once, triggers stay wired, code and configuration are untouched, and you restore service by removing or raising the reservation. It is faster and far more reversible than deleting event source mappings or the function itself during an incident.
saying these in an interview costs you the question
- Thinks reserved concurrency keeps environments warm
- Believes provisioned concurrency also caps the function's scaling
- Forgets that a reservation is a hard maximum, not just a guarantee
- Tries to set provisioned concurrency on $LATEST
- Assumes provisioned concurrency is free when idle