skip to content

What does 'scaling to zero' mean in a serverless platform, and what latency and cost trade-off does it introduce via cold starts?

level: seniorimportance: should knowfreq 60%

answer

  1. idle -> zero instances -> $0
  2. cold start = first-request setup tax
  3. JVM cold starts slower than scripting runtimes
  4. provisioned/min instances trade cost back for latency
  5. cold-start storms hit during bursts, not just idle gaps

basics

~20 s

When idle, the platform fully shuts down your function's instances, so you pay nothing. The catch: the next request after idle has to wait for a fresh instance to start up — that delay is a cold start.

solid answer

~50 s

Scaling to zero means the platform fully deallocates every running instance of a function once it's been idle for some period, so idle time is billed at zero — this is what makes the 'pay only for what you use' promise literal. The cost is latency: the next request after that idle gap (or any request needing a new instance during a burst) triggers a cold start, where the platform must provision a fresh sandbox, load the runtime and code, and run initialization logic before handling the request, adding anywhere from tens of milliseconds to several seconds depending on the runtime. JVM-based runtimes typically cold-start slower than scripting languages. The mitigation is provisioned or minimum-instance concurrency, which keeps some instances permanently warm to avoid this — but that reintroduces a continuous, always-on-style cost for exactly the capacity being kept warm.

go deeper

for a junior

Knows scaling to zero means no cost when idle, and that the first request after idle can be slower.

for a middle

Can explain what happens during a cold start at a high level (new environment plus init code) and why it costs latency.

for a senior

Can reason about runtime-specific cold-start differences, cold-start storms under burst traffic, and the cost trade-off of provisioned/minimum-instance concurrency as a targeted mitigation.

for a principal

Sets platform-wide policy on which paths get scale-to-zero versus guaranteed-warm capacity based on SLA criticality and traffic shape, and factors cold-start risk into runtime and language choice for latency-sensitive serverless services.

## What a cold start actually does A serverless platform runs your code inside **ephemeral, sandboxed execution environments** — commonly lightweight microVMs or containers. When a request arrives and no warm environment currently exists to serve it, the platform must, in order: 1. provision a fresh sandbox, 2. load the language runtime and your deployed code or container image, 3. and execute any module-level or initialization code your function defines (opening a database connection, loading configuration, initializing SDK clients), before it can finally run the actual handler logic for that specific request. All of that upfront work is what's called a **cold start**, and it adds latency ranging from tens of milliseconds to several seconds on top of the function's normal execution time, depending heavily on the runtime — interpreted languages like `Python` or `Node.js` are typically much faster to cold-start than JVM-based runtimes, which pay a heavier class-loading, dependency-injection, and JIT-warmup penalty on top of runtime startup. - Once an environment is warm, subsequent requests reuse it directly and skip all of that overhead, running as a **warm start** with only single- or low-double-digit milliseconds of framework overhead. - After some period of no traffic to a given environment — implementation-specific, often on the order of minutes — the platform tears that environment down entirely, freeing the underlying resources and, critically, stopping any billing tied to keeping it around. That teardown-on-idle behavior is exactly what **scaling to zero** refers to. ## Why the platform tears environments down This behavior exists because it's the mechanism that makes serverless's core economic promise literally true. If the platform kept some minimum capacity running by default for every deployed function, every low-traffic function across every customer would silently accrue a continuous baseline cost, undermining the entire 'pay only for what you use, nothing while idle' pitch — especially at the scale of millions of functions running across a major provider's customer base, where that idle baseline would add up enormously. Scaling to zero is what lets a rarely used function cost genuinely nothing between invocations. ## The trade-off The trade-off, with cost named explicitly on both sides: - **Scaling to zero** costs nothing during silence — a true zero-dollar idle state — but it costs latency on the way back up, and that first-request cold start can be severe enough to violate a service-level objective for a user-facing, latency-sensitive endpoint that's expected to respond in well under a second. - **The available mitigation** — provisioned or minimum-instance concurrency, which keeps a specified number of environments permanently initialized and warm — trades that latency risk back for a real, continuous dollar cost: you now pay for those N warm environments around the clock whether or not they're actively handling requests, which is functionally re-introducing an always-on cost profile, just scoped down to a tunable minimum rather than a full fixed fleet. ## Failure modes - **A low-traffic internal service.** A very common production failure mode is a low-traffic internal service — a few requests per hour — that scales to zero between calls and, if built on a JVM-based runtime with a heavy dependency-injection framework, pays a multi-second cold start on essentially every call, making the service feel randomly and confusingly slow to anyone unfamiliar with the pattern. It's a classic on-call puzzle: p50 latency looks perfectly fine, but tail latency spikes correlate tightly with gaps in traffic rather than with load, which is the opposite of what most engineers instinctively check first. - **A cold-start storm.** A second, distinct failure mode is a sudden burst of concurrent traffic that requires more simultaneously warm instances than currently exist, so a large fraction of that burst's requests each independently trigger their own cold start at the same time — producing a visible latency spike at precisely the moment the system is under the most load and least able to absorb it. - **Teams that over-provision warm capacity.** A third, subtler failure is the inverse of the intended savings: teams that over-provision minimum-instance or provisioned concurrency 'just in case' quietly convert a function that was supposed to be a true scale-to-zero cost center into an always-on-priced one, and are sometimes surprised months later to discover they've been paying continuously for warm capacity they believed was billed strictly on use. ## Where it shows up A concrete real-world scenario illustrates the right way to split the difference: - **A webhook receiver** for a third-party payment provider that fires only a handful of times a day is a great fit for true scale-to-zero, since the occasional few-hundred-millisecond cold start is invisible inside an asynchronous callback flow that nothing user-facing is waiting on. - **A customer-facing checkout API** running on the same platform, however — expected to respond in under a second at the 99th percentile even during a bursty flash-sale spike — is a poor fit for pure scale-to-zero, and teams typically size provisioned or minimum-instance concurrency to the expected burst floor for that specific critical path, deliberately accepting the associated always-on-like cost there while still letting less critical, lower-traffic paths scale to zero freely.

  • Why do JVM-based runtimes typically suffer worse cold starts than something like Python or Node.js on the same serverless platform?
    JVM startup involves loading and verifying classes, initializing the class-loader hierarchy, and JIT-compiling hot paths essentially from scratch, which is inherently heavier than an interpreted language beginning to execute source or bytecode almost immediately; frameworks that rely on heavy reflection-based dependency injection compound this further, though ahead-of-time compilation approaches can substantially reduce the gap.
  • What is a 'cold-start storm' and when does it happen?
    It's when a sudden burst of traffic requires many new concurrent execution environments at once, and since none of them are warm yet, a large fraction of the burst's requests each pay the full cold-start penalty simultaneously — happening exactly when the system is under the most load and least able to absorb extra latency.
  • If a team sets provisioned or minimum-instance concurrency to eliminate cold starts entirely, what have they effectively given up from the original scale-to-zero cost model?
    They've reintroduced a continuous, always-on-style cost for that reserved warm capacity — they still avoid managing servers directly, but they no longer get true zero-cost idle time for that function, since they're paying to keep instances initialized whether or not they're handling requests.

Like a pop-up coffee stand that fully packs up when there's no line: it costs nothing to keep it running while closed, but the very next customer has to wait for the stand to be set back up before getting served — unless you pay staff to keep it assembled and ready even when no one's in line.

saying these in an interview costs you the question

  • Claims scaling to zero has no meaningful downside
  • Can't describe what actually happens during a cold start
  • Assumes all language runtimes cold-start at roughly the same speed
  • Unaware that provisioned/minimum-instance concurrency reintroduces continuous cost
  • Can't explain why sudden bursts, not just idle periods, worsen cold-start impact

context