skip to content

Cold Starts

The first invocation on a fresh instance pays for container and runtime initialization, which shows up squarely in your p99. You will cover what drives the delay and the mitigations — provisioned concurrency, SnapStart, lighter runtimes, keep-alive pings — with their costs.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In a serverless platform like AWS Lambda, what is a 'cold start' and why does it add latency to the first request?

level: juniorimportance: must knowfreq 75%

answer

  1. idle -> reclaimed environment
  2. provisioning + runtime boot + init
  3. warm reuse skips setup
  4. elasticity's unavoidable tax
  5. concurrency scale-up = new cold instances

basics

~20 s

A cold start is the extra delay when a serverless function runs for the first time (or after being idle) because the cloud provider has to set up a fresh container and boot the code before it can process the request, instead of reusing one that's already running.

solid answer

~30 s

When a serverless request arrives and no warm execution environment exists, the platform must provision a new sandbox (container or microVM), initialize the language runtime, and load/initialize application code (static initializers, DB connection pools, SDK clients) before invoking the handler. This 'cold start' adds latency on top of normal execution time. Subsequent requests within the idle-timeout window reuse that same warm environment ('warm start'), skipping provisioning and init. Cold starts recur whenever concurrency scales up, a function has been idle long enough to be reclaimed, or a new code deployment happens.

go deeper

for a junior

Can state that a cold start is the extra delay from spinning up a fresh environment versus reusing a warm one, and knows it happens on 'first' requests.

for a middle

Can name the concrete phases (provisioning, runtime boot, init code) and explain why concurrency scale-up and redeploys both trigger fresh cold starts, not just idle timeout.

for a senior

Connects cold start behavior to language/runtime choice and package size, and can reason about where in a system (sync vs async paths) cold starts actually matter for user-facing SLAs.

for a principal

Reasons about cold starts as an inherent trade-off of the elastic-to-zero cost model, and can make organization-level calls on runtime standardization, provisioned concurrency budget, and SLA design given that trade-off.

## What a cold start is A **cold start** is the latency penalty incurred the first time a serverless function instance handles a request, as opposed to a *warm start* where an already-initialized instance handles it. To understand why this penalty exists, you need to understand how serverless platforms like **AWS Lambda**, **Azure Functions**, or **Google Cloud Functions** actually execute code: they don't keep a fleet of instances running for your function at all times. Instead, they run functions inside ephemeral, isolated execution environments — historically containers, and in AWS Lambda's case since 2018 backed by **Firecracker microVMs** — that are created on demand and torn down when unused. ## How an environment gets built The mechanism, step by step: when an invocation arrives and the platform's routing layer finds no idle warm environment for that function, it must - **(1) provision compute** — find capacity on a host, create or restore a microVM/container, allocate CPU/memory; - **(2) bootstrap the runtime** — start the language runtime process, e.g. the JVM, the Node.js V8 engine, the Python interpreter, the .NET CLR; - **(3) initialize the function** — download/mount the deployment package or image, run module-level/static initialization code (imports, singleton construction, DB connection pools, SDK client instantiation, DI wiring); - and only then **(4) invoke the actual handler** with the event payload. Steps 1-3 are the cold-start overhead — pure tax a warm invocation skips because the environment from a previous invocation is frozen and reused, with only step 4 running. ## Why this exists The entire economic and operational proposition of serverless is that you don't pay for or manage idle capacity — the platform multiplexes a shared fleet of hosts across enormous numbers of tenants, spinning environments up only when there's traffic and reclaiming them after a period of inactivity (roughly minutes, undocumented as a hard SLA). This is what lets a rarely-used endpoint cost nothing when idle and still scale to thousands of concurrent invocations during a spike, with no capacity planning. Cold starts are the direct, unavoidable cost of that elasticity: you cannot get *zero idle cost* and *always warm* simultaneously without keeping something running, which is exactly what **provisioned concurrency** (a paid feature) does. ## Upside and downside **Trade-offs.** The upside: - you never provision capacity; - you pay per-invocation/per-ms; - you scale automatically. The downside is unpredictable added latency on some fraction of requests — worse for low-traffic functions and worse the moment traffic scales up faster than existing warm instances can absorb it, since each new concurrent execution needs its own environment. Language and package size matter enormously: | What you ship | Cold-start cost | Why | |---|---|---| | A small **Go** or **Rust** binary | might cold-start in tens of milliseconds | because there's no separate interpreter/VM startup | | A **JVM-based** function with a large dependency graph (Spring, especially) | can take multiple seconds | because interpreted startup, classloading, and reflection-heavy DI wiring are all serial costs paid before application code runs | **Deployment/image size** adds latency too because code must be fetched/mounted before init. ## Where the pain shows up **Failure modes in production.** The most common way this bites teams is a latency SLA violation that only shows up in the tail — p50 looks fine because it's dominated by warm requests, but p99/p999 spikes because a fraction of requests land on a cold environment, especially right after a deploy (every deploy invalidates warm environments) or during a traffic burst (new concurrency always cold-starts). Synchronous request paths behind an **API Gateway** or **ALB** are the most painful because the caller is blocked; async/event-driven consumers are far more tolerant since a few extra hundred milliseconds rarely matters. Teams often *fix* this by - shrinking dependencies; - switching runtime for latency-sensitive paths; - enabling provisioned concurrency; - lazy-initializing expensive resources only when actually needed rather than at module load. ## In the field A concrete scenario: a fintech's payment-authorization API on Lambda, backed by Java/Spring Boot, met p50 SLAs comfortably but periodically breached p99 SLAs during morning traffic ramp-up, when auto-scaling created a burst of new cold environments simultaneously; JVM init plus Spring context startup added roughly 2-4 seconds per cold instance. The team combined provisioned concurrency sized to the morning ramp-up rate with lazy-initializing the heaviest DI-based startup work, cutting the residual cold-start penalty by more than half.

  • Does every single invocation of a serverless function on a busy, high-traffic endpoint experience a cold start?
    No. Once an execution environment is warm, the platform reuses it for subsequent invocations as long as it stays busy or idle within the retention window, so a steady stream of traffic to a single-concurrency workload mostly hits warm instances. Cold starts happen specifically when concurrency needs to increase or when an instance has been idle long enough to be reclaimed, so bursty or spiky traffic sees proportionally far more cold starts than steady, high-throughput traffic.
  • Why does a Java/Spring function typically cold-start much slower than a Go function on the same platform?
    The JVM has nontrivial startup cost (class loading, bytecode verification, no JIT warm-up yet), and Spring's DI container adds heavy reflection-based bean instantiation and wiring on top, all executing before the handler runs. Go compiles to a static native binary with no separate runtime to boot and typically has a much smaller init graph, so its cold start is often tens of milliseconds versus low seconds for Spring.
  • Does a code deployment affect cold start rates even if traffic volume doesn't change?
    Yes. A new deployment creates a new function version, and existing warm environments running the old code aren't reused for the new version — every environment for the new version starts cold on its first invocations. This is why teams sometimes see a latency spike immediately after a rollout even though request volume is flat.

Like a restaurant kitchen that's shut down between rushes — the first order after a lull has to wait for the stove to heat up and staff to arrive, while orders right after are served immediately because everything's already running.

saying these in an interview costs you the question

  • Says cold start only happens on the very first invocation ever, not on every scale-up event
  • Claims warm instances are guaranteed to stay warm indefinitely
  • Ignores that new deployments reset warm pools
  • Thinks cold start latency is the same across all languages/runtimes
  • Confuses cold start with network latency or database latency

context

open as a page

Break down where the time actually goes during an AWS Lambda cold start — what are the distinct phases, and which one typically dominates for an interpreted language like Python versus a JVM-based language like Java?

level: middleimportance: must knowfreq 70%

basics

~20 s

A cold start has three parts: finding a machine and starting an isolated sandbox, booting the language engine (like the JVM or Python interpreter), and running your own startup code (imports, DB connections). For Python it's usually your own code and the interpreter that dominate; for Java it's usually the JVM and framework startup.

open as a page

How does AWS Lambda's provisioned concurrency mitigate cold starts, and what are the cost and operational trade-offs of relying on it versus letting the platform auto-scale on demand?

level: seniorimportance: must knowfreq 80%

basics

~20 s

Provisioned concurrency pays to keep a set number of function instances pre-warmed and ready at all times, so requests hit already-initialized code instead of waiting for a fresh one to boot — but you pay for that idle readiness even when there's no traffic.

open as a page

Why is pinging a serverless function on a schedule (a 'keep-alive' or 'warming' ping) considered a weak, sometimes actively misleading, mitigation for cold starts?

level: middleimportance: should knowfreq 55%

basics

~20 s

A scheduled ping keeps one instance warm, but real traffic often needs many instances at once, and pings can't predict or prevent the cold starts that happen when demand suddenly scales up beyond what's already warm.

open as a page

How does AWS Lambda SnapStart reduce cold starts for Java functions, and what correctness pitfalls does its snapshot-and-restore approach introduce that provisioned concurrency doesn't have?

level: seniorimportance: should knowfreq 45%

basics

~20 s

SnapStart boots and initializes your function once, then saves a memory snapshot of that fully-warmed state; every future cold start just restores from that snapshot instead of re-running startup, making it much faster. The catch: cached state from before the snapshot (like random numbers or timestamps) can leak into later invocations unless you handle it explicitly.

open as a page

For a serverless API with a strict p99 latency SLA, why can average or even p50 latency numbers hide a cold-start problem entirely, and what architectural options exist beyond provisioned concurrency to keep tail latency in check?

level: principalimportance: should knowfreq 50%

basics

~20 s

Most requests hit already-warm functions and look fast, so the average looks great, but the rare slow ones (cold starts) show up only in the tail (p99, the slowest 1%). Fixing that tail needs more than just averages — options include keeping capacity pre-warmed, timing out and retrying to a warm instance, or picking faster-starting runtimes.

open as a page