How does AWS Lambda's provisioned concurrency mitigate cold starts, and what are the cost and operational trade-offs of relying on it versus letting the platform auto-scale on demand?
answer
- pre-warms N environments continuously
- billed even when idle — reserved capacity cost
- overflow beyond N still cold-starts
- tied to a version/alias, not $LATEST
- pair with scheduled/target-tracking auto scaling
basics
~20 sProvisioned concurrency pays to keep a set number of function instances pre-warmed and ready at all times, so requests hit already-initialized code instead of waiting for a fresh one to boot — but you pay for that idle readiness even when there's no traffic.
solid answer
~40 sProvisioned concurrency reserves and pre-initializes a specified number of execution environments ahead of time — running the runtime bootstrap and function init code before any request arrives — so invocations up to that count are served warm with no cold-start penalty. You're billed for that reserved capacity continuously (a separate rate from standard per-invocation pricing) whether or not it's used. Traffic beyond the provisioned count still cold-starts on standard on-demand Lambda. The main trade-off is cost predictability vs. serverless's core value proposition of paying only for actual usage — it reintroduces a capacity-planning problem (often paired with scheduled or target-tracking auto-scaling policies) and needs re-provisioning after each deployment before traffic cuts over, or you eat cold starts on the new version.
go deeper
Knows provisioned concurrency 'keeps functions warm' so they respond faster.
Understands it pre-initializes a fixed number of environments and that you pay for that continuously, separate from per-invocation pricing.
Can reason about sizing the pool against real traffic patterns, the version/alias binding requirement, and overflow behavior beyond the provisioned count.
Weighs provisioned concurrency against alternative architectures (always-on containers/EKS, different runtime choice) at the cost/SLA/operational-complexity level, and designs deployment/traffic-shifting strategy so releases don't reintroduce cold starts.
## What the feature actually does **Provisioned concurrency** is AWS Lambda's mechanism for eliminating cold starts on a bounded slice of traffic by paying to keep environments pre-warmed continuously, rather than waiting for real invocations to trigger environment creation. Mechanically, when you configure provisioned concurrency for `N` on a specific function version or alias, Lambda immediately and continuously initializes `N` execution environments ahead of any traffic: - it provisions the microVMs; - it boots the runtime; - it runs your function's init-phase code (imports, static blocks, DI wiring, connection pool setup); so all `N` environments sit in a *warm and idle* state, ready to be invoked with zero setup latency the moment a request routes to them. As long as concurrent demand stays at or below `N`, every invocation is served by one of these already-initialized environments — the caller experiences only the handler's actual execution time, with the cold-start tax already paid up front and amortized across all subsequent invocations. ## The problem it buys off **Why this exists.** Some workloads simply cannot tolerate the occasional multi-second cold start that a standard on-demand Lambda might incur, particularly synchronous, user-facing, or latency-SLA-bound endpoints (a mobile app's login call, a payment authorization API) where an occasional 2-3 second stall is a customer-visible failure, not just a metric blip. Standard Lambda pricing lets you pay nothing when idle, but that flexibility is precisely what causes cold starts — there is no free way to guarantee *always instantly ready* without paying for readiness continuously, and provisioned concurrency buys exactly that guarantee, function by function. ## The trade-off The trade-off is the core of what makes this an interesting architectural decision rather than a free win. **First, cost model.** Provisioned concurrency is billed for the duration it's configured, per GB-second at a separate rate, regardless of whether it's actually invoked — an over-provisioned pool wastes money in exactly the way serverless is supposed to avoid, and this can eat into or exceed the cost savings that justified going serverless. Sizing it correctly requires understanding real traffic patterns (peak concurrency, not just average throughput), which reintroduces capacity-planning burden — teams commonly pair it with **AWS Application Auto Scaling** target-tracking or scheduled scaling policies to raise the provisioned pool ahead of known traffic patterns (e.g., before a 9am ramp), which itself is an operational surface to build, test, and monitor. ## Coverage is bounded **Second, coverage is bounded, not absolute.** Provisioned concurrency only guarantees warm environments up to `N` concurrent executions. If real traffic bursts past `N`, the overflow falls back to standard on-demand scaling and cold-starts normally — so provisioned concurrency reduces the cold-start rate, it does not categorically eliminate cold starts unless `N` is provisioned generously above realistic peak (which costs more). Teams need monitoring on metrics like `ProvisionedConcurrencyUtilization` and `ProvisionedConcurrencySpilloverInvocations` to know whether their pool is actually sized correctly, not just assume the feature solved the problem once configured. ## Deployment interaction **Third, deployment interaction.** Provisioned concurrency is tied to a specific published function version or an alias pointing to one — not to `$LATEST`. Every new deployment requires either 1. pre-warming a new version's provisioned pool before shifting traffic to it (e.g., via a weighted alias and gradual traffic shift, similar to a blue/green rollout), or 2. accepting that the cutover moment reintroduces cold starts until the new pool finishes initializing. Teams that naively update an alias to point at a freshly deployed version without waiting for its provisioned concurrency to finish initializing effectively defeat the feature during every release window — a subtle and common failure mode. ## In the field A concrete real-world scenario: a ride-hailing company's driver-matching API on Lambda had a strict p99 latency SLA because a slow match directly affected user-perceived reliability. They set provisioned concurrency sized to their historical p99 peak concurrent request count during rush hours, paired with a scheduled Application Auto Scaling policy that raised the pool ahead of morning and evening rush and lowered it overnight to control cost, and used a canary/linear deployment strategy via an alias with traffic shifting so the new version's provisioned pool was fully warmed before receiving production traffic. This cut cold-start-driven SLA violations to near zero during known peak windows, while still letting cost scale down during predictably quiet hours — an explicit acknowledgment that pure on-demand elasticity and zero-idle-cost weren't compatible with their latency requirements, so they bought back predictability where it mattered and let the rest of the system stay elastic.
- If a function has provisioned concurrency set to 10, and it suddenly receives 30 concurrent requests, what happens to the extra 20?The first 10 are served by the pre-warmed, provisioned environments with no cold start. The remaining 20 concurrent requests are handled by standard on-demand Lambda scaling, which creates fresh environments and pays the normal cold-start cost for each, so a third of the burst still experiences full cold-start latency.
- Why can't you just set provisioned concurrency on $LATEST and call it done after every deploy?AWS Lambda ties provisioned concurrency to a published, immutable version or an alias pointing at one — $LATEST is mutable and not supported as a provisioned-concurrency target. This forces teams to publish a version, attach or update provisioned concurrency on it, wait for it to finish warming, and only then shift traffic (often via a weighted alias), otherwise the deployment moment itself reintroduces cold starts.
- Is provisioned concurrency always cheaper than accepting occasional cold starts?No — it's a cost-for-latency trade, not a strict win. For low-traffic or latency-tolerant workloads, the continuous reserved-capacity billing of provisioned concurrency can cost more than the business impact of occasional cold starts, so it's typically reserved for genuinely latency-sensitive, synchronous, user-facing paths rather than applied blanket across every function.
Like a restaurant keeping a fixed number of chefs on shift and pans pre-heated even during slow hours, so the first orders of a rush are served instantly — you pay those chefs' wages the whole time, and if a bigger rush arrives than staffed for, the overflow orders still wait for a chef to be called in.
saying these in an interview costs you the question
- Thinks provisioned concurrency eliminates all cold starts unconditionally
- Doesn't know it's billed continuously regardless of usage
- Assumes it applies automatically to $LATEST after every deploy with no extra steps
- Unaware that overflow traffic beyond the provisioned count still cold-starts
- Treats it as a free feature with no cost or capacity-planning trade-off