skip to content

What is a cold start in AWS Lambda, and why does the same function usually respond faster on the next invocation a moment later?

level: juniorimportance: must knowfreq 85%

answer

  1. new sandbox before your code runs
  2. two phases, only one repeats
  3. frozen, not destroyed, between calls
  4. Init Duration in the REPORT line
  5. scale-out pays it again

basics

~20 s

A cold start is the extra latency Lambda spends creating a new execution environment for an invocation: downloading the code, starting the runtime, and running your initialization code before the handler runs. A later invocation reuses that environment and skips all of it.

solid answer

~50 s

Lambda runs your code inside an execution environment — a small isolated sandbox. When a request arrives and no idle environment exists, Lambda creates one: it fetches the deployment package, starts the language runtime, and runs everything in your module's global scope, which AWS calls the Init phase. Only then does it call your handler. That whole preamble is the cold start, and you see it in the CloudWatch Logs `REPORT` line as `Init Duration`. After the handler returns, Lambda does not destroy the environment; it freezes it and keeps it around for a while, so the next invocation is routed to the warm environment and only pays the Invoke phase. Cold starts therefore show up on the first request after a deploy, whenever traffic scales out to more concurrent environments, and after an idle period — they hit a small share of requests but land in your tail latency.

go deeper

for a junior

Be able to say plainly that a cold start is the setup time for a new execution environment and that the next call reuses it. Knowing that init code lives outside the handler is enough at this level.

for a middle

Explain the Init/Invoke split and the freeze-thaw behaviour between invocations, and name the concrete triggers: deploys, scale-out, and reclaimed idle environments.

for a senior

Show that you treat cold starts as a tail-latency measurement problem — pull Init Duration out of the logs, quantify the p99 impact, and only then decide whether provisioned concurrency or SnapStart is worth the money.

for a principal

Own the framing question: is this workload latency-sensitive enough that serverless cold starts are a real product risk, or is the right answer a different compute model? Be ready to price warm capacity against re-platforming.

## The unit Lambda actually manages Lambda does not run "a function"; it runs your code inside an **execution environment** — an isolated sandbox (a Firecracker microVM) with your deployment package, the language runtime, a slice of memory, and its own writable `/tmp`. One environment handles **exactly one invocation at a time**. That single sentence explains nearly everything about Lambda scaling and cold starts: to serve two simultaneous requests, Lambda needs two environments; if it doesn't have a spare one, it has to build it, and building it takes time. ## The lifecycle: Init, Invoke, Shutdown An environment moves through three phases. **Init** happens once per environment. Lambda downloads and unpacks your code, starts any extensions, starts the language runtime, and then executes your function's *initialization code* — everything at module/static scope, outside the handler. In Node.js that is the top level of your module; in Python the module body; in Java static initializers and the constructor of your handler class. **Invoke** is the part you normally think of as "the function running": Lambda passes the event to your handler and waits for the response. This repeats for every request the environment serves. **Shutdown** happens when Lambda decides to reclaim the environment; the runtime and extensions get a chance to stop. A **cold start** is an invocation that had to pay for an Init phase first. A **warm start** is one routed to an environment that already exists. ## Freeze and thaw Between invocations Lambda **freezes** the environment rather than destroying it. Execution is suspended: background threads, timers, and pending I/O stop making progress, and they resume (thaw) only when the next event arrives. Two consequences follow. First, work you kicked off but did not await before returning may simply not finish — and may resume, confusingly, during a *later* invocation. Second, everything you built during Init is still in memory when the environment thaws, which is exactly why the second call is fast. ```javascript // runs once per environment (Init) — reused by every later invocation const client = new SomeClient(); export const handler = async (event) => { // runs once per request (Invoke) return client.doWork(event); }; ``` ## When cold starts actually happen - **First request after a deploy.** Publishing new code or changing configuration invalidates existing environments; the next requests all initialize fresh. - **Scale-out.** Traffic rising from 10 to 60 simultaneous in-flight requests means Lambda must stand up roughly 50 more environments, and each of those requests eats an Init. - **After idleness.** Lambda eventually reclaims unused environments. AWS does not publish the idle timeout, so never design as if a warm environment is guaranteed. - **Periodic recycling.** Environments are replaced over time even under steady traffic, for patching and health reasons. Steady, high-volume traffic therefore has a *low percentage* of cold starts, but a spiky or low-volume function can see them constantly. ## How you observe it Every invocation writes a `REPORT` line to CloudWatch Logs. On a cold start that line carries an extra `Init Duration` field: ``` REPORT RequestId: ... Duration: 12.34 ms Billed Duration: 13 ms Memory Size: 512 MB Max Memory Used: 90 MB Init Duration: 430.12 ms ``` Because only a subset of requests are cold, cold starts are a **tail-latency** problem: they barely move your p50 and can dominate p99. Judge them there, not on averages. ## What influences init time, and what to do Init time is dominated by how much work happens before your handler is reachable: runtime startup plus loading dependencies and constructing clients. Runtimes differ substantially — an interpreted or ahead-of-time-compiled runtime typically initializes faster than a JVM that must load and JIT a large classpath. Doing less at Init (lazy-loading rarely used dependencies, importing narrowly rather than pulling in an entire SDK) is the cheapest lever. When a workload cannot tolerate the tail, AWS sells two answers: **provisioned concurrency**, which keeps a pool of environments already initialized and waiting, and **SnapStart**, which restores an environment from a snapshot taken after Init instead of re-running it. Both remove the Init phase from the request path rather than making Init faster. What does **not** work is the folklore fix: a scheduled "ping" that invokes the function every few minutes keeps *one* environment warm and does nothing for the scale-out case, which is where cold starts actually hurt.

  • Does keeping a function warm with a scheduled ping every five minutes solve cold starts?
    It only keeps roughly one environment alive. Cold starts hurt most during scale-out, when Lambda must create many environments at once for concurrent requests, and a single ping does nothing there. It also costs invocations forever. If the latency tail genuinely matters, provisioned concurrency or SnapStart address it directly; a warming ping is a folk remedy that hides the problem at low traffic.
  • Where would you look to tell whether a latency spike was caused by cold starts?
    Compare the `Duration` distribution against the `Init Duration` values in the CloudWatch Logs `REPORT` lines — a Logs Insights query filtering for records that contain `Init Duration` isolates the cold invocations. If your p99 spikes line up with a burst of initialized environments right after a deploy or a traffic ramp, cold starts are the cause; if slow invocations have no `Init Duration`, look at the handler itself.
  • Can two requests ever be handled by the same execution environment at the same time?
    No. An execution environment processes one invocation at a time, which is why concurrency is measured in environments. Two simultaneous requests always need two environments. This is also why in-memory state in a single environment is safe from data races between invocations, but is not shared with any other environment.

saying these in an interview costs you the question

  • Says every invocation starts a brand-new container
  • Thinks cold starts only happen on the very first request ever
  • Claims a scheduled warming ping eliminates cold starts
  • Believes init code runs again on every invocation
  • Judges cold-start impact from average latency instead of p99

context