skip to content

What specifically causes a FaaS cold start's latency (walk through the init phase), and what mechanisms do AWS Lambda, Google Cloud Functions, and Azure Functions each offer to reduce or avoid it - and what do those mechanisms cost you?

level: seniorimportance: must knowfreq 75%

answer

  1. cold start = sandbox boot + package pull + runtime boot + init code
  2. Lambda: Provisioned Concurrency (always warm) vs SnapStart (snapshot/restore)
  3. Cloud Functions/Cloud Run: min instances
  4. Azure Functions Premium: Always Ready instances
  5. all pre-warm options cost idle capacity except snapshot-based ones

basics

~20 s

A cold start is slow because the platform has to boot a brand-new environment: download your code, start the language runtime, and run any setup code before your function can even begin. All major clouds let you pay extra to keep a pool of environments pre-warmed and ready, so real users never hit that slow first call.

solid answer

~50 s

A cold start's latency comes from a sequence of steps the platform must do before your handler runs: provisioning the sandbox (container or microVM), pulling/mounting your deployment package or image, starting the language runtime process, and executing your init-phase code (imports, static initializers, client construction) - each step adds real time, and heavier runtimes (JVM, .NET) or large dependency trees make it worse. All three major providers offer a way to pre-warm capacity: AWS Lambda has Provisioned Concurrency, which keeps a fixed number of environments permanently initialized and ready (and SnapStart, which snapshots a fully-initialized JVM environment and resumes from that snapshot instead of booting fresh); Google Cloud Functions/Cloud Run let you set a minimum instance count so some instances stay warm; Azure Functions offers 'Always Ready' instances on its Premium plan. All of these cost real money for idle capacity you're paying to keep warm even when there's no traffic, so they're a targeted fix for latency-sensitive paths, not a default you flip on everywhere.

go deeper

for a junior

Should know that cold starts can be mitigated somehow and that cloud providers offer settings for this, without needing the mechanism names.

for a middle

Should be able to name at least one concrete mitigation (e.g., provisioned concurrency or minimum instances) and know it costs extra.

for a senior

Should walk through the init-phase pipeline concretely, name provider-specific mechanisms accurately, and reason about when the cost of pre-warming is justified versus not.

for a principal

Should evaluate mitigation strategy as a cost/latency portfolio decision across an entire system - which paths deserve provisioned/pre-warmed capacity, how to detect snapshot-related correctness traps in a review process, and when the right call is to move a workload off FaaS entirely rather than keep paying to patch around its cold-start characteristics.

## The init-phase pipeline A cold start is not one atomic event but a **pipeline of discrete steps** the platform must execute before your handler code runs for the first time in a given execution environment, and understanding that pipeline is what lets you reason about where the latency actually comes from and which mitigation targets which step. 1. **First, the platform's control plane has to provision the sandbox itself** - for AWS Lambda this means launching a Firecracker microVM (a lightweight, KVM-based virtualization technology built specifically for fast-booting, secure multi-tenant sandboxes), for other providers it's typically a gVisor-sandboxed or similarly isolated container. 2. **Second, the platform has to get your code into that sandbox** - pulling a container image or mounting a deployment package, which is slower the larger your package/image is (bloated dependency trees are a common, avoidable cause of slow cold starts). 3. **Third, the language runtime itself has to start** - a JVM has to boot the HotSpot virtual machine and JIT-compile hot code paths from scratch, a .NET function has to start the CLR, whereas Go compiles to a static binary with almost no runtime startup cost and Node.js/Python interpreters start comparatively fast too. 4. **Fourth, and often overlooked, your own init-phase code runs** - module-level imports, static initializer blocks, framework bootstrapping (a Spring Boot application context, for instance, can easily take 2-4 seconds to build on its own) - before your handler is finally invoked with the event. ## What each provider offers All three major providers offer mechanisms to short-circuit some or all of this pipeline by keeping capacity pre-warmed, though the specifics differ. - **AWS Lambda's Provisioned Concurrency** lets you specify a number of execution environments that AWS keeps permanently initialized - already booted, already through your init-phase code - so invocations routed to them skip the entire cold-start pipeline; you can also attach Application Auto Scaling to adjust the provisioned count on a schedule or based on utilization. - **Lambda additionally offers SnapStart** (originally for Java, later extended to other runtimes) which takes a different approach: instead of keeping environments alive continuously, it snapshots a fully-initialized environment's memory and disk state after your init code has run once, caches that snapshot, and on a cold start restores execution from the snapshot rather than re-running the JVM boot and your init code from scratch - this can cut multi-second Java cold starts down to a few hundred milliseconds without paying for continuously-idle capacity. - **Google Cloud Functions** (and its close relative Cloud Run, which the newer Cloud Functions gen2 is built on) lets you configure a minimum number of instances that stay resident and warm even at zero traffic, directly analogous to provisioned concurrency. - **Azure Functions** offers 'Always Ready' instances on its Premium plan (and dedicated App Service plans inherently avoid the cold-start problem since the underlying compute is never deallocated), while its default Consumption plan has no such option and is the tier most exposed to cold starts. ## What pre-warming costs you The cost side of all of these mechanisms is real and needs to be weighed deliberately. Every pre-warming mechanism means you are **paying for idle capacity** - a provisioned-concurrency Lambda environment or an Azure 'Always Ready' instance is billed continuously whether or not it handles a request, which directly undermines the pay-per-invocation economics that make serverless attractive in the first place. This means these mechanisms make sense as a **targeted fix** for specific latency-sensitive paths - a payment API with a hard SLA, a user-facing endpoint where a multi-second tail spike would breach a UX budget - rather than something you apply blanket-style to every function in a system. ## The snapshot caveat SnapStart is a partial exception since it doesn't require paying for continuously-idle capacity, but it comes with its own caveats: any state captured in the snapshot (like a random seed, a unique ID generator, or a network connection) can be problematically 'frozen' and reused across restores unless you explicitly hook into the snapshot/restore lifecycle to regenerate it, which is a subtle correctness trap teams have been bitten by. ## A concrete real-world example A concrete real-world example: some teams running Java payment-processing Lambdas found Provisioned Concurrency's continuous cost unacceptable for functions with spiky, unpredictable traffic, so they migrated those functions to SnapStart instead once it became available for their runtime - eliminating multi-second cold starts on genuinely cold invocations without paying an idle-capacity bill, at the cost of auditing their init code for any non-deterministic state (like cryptographic nonces generated at startup) that needed to be explicitly refreshed via SnapStart's before-checkpoint/after-restore hooks rather than trusted to run fresh.

  • What's a subtle correctness risk introduced by AWS Lambda SnapStart that provisioned concurrency doesn't have?
    Because SnapStart resumes execution from a memory/disk snapshot taken after your init code ran once, any state that should be unique or fresh per environment - like a random seed, a generated UUID, or an open network connection - can get silently reused across many restores unless you explicitly regenerate it using SnapStart's before-checkpoint/after-restore hooks. Provisioned concurrency doesn't have this issue because each environment genuinely runs its own init code once, live, rather than resuming from a shared snapshot.
  • Why wouldn't a team just turn on provisioned concurrency for every function in their system to eliminate cold starts everywhere?
    Provisioned concurrency bills for the reserved capacity continuously regardless of traffic, so applying it everywhere effectively converts a large chunk of your serverless bill back into paying for idle compute, undermining the cost efficiency that made FaaS attractive in the first place. It's typically reserved for a small set of latency-sensitive, high-value paths rather than applied blanket-style.

Like the difference between a restaurant that keeps a chef standing by fully prepped at all times (provisioned concurrency - you pay for their idle time) versus one that keeps a perfectly reheatable pre-made meal in the fridge ready to microwave (SnapStart-style snapshot) instead of cooking from raw ingredients every time a customer walks in cold.

saying these in an interview costs you the question

  • Thinks cold starts are unavoidable with no mitigation options
  • Can't name any concrete mitigation mechanism beyond 'just wait for it to warm up'
  • Doesn't realize pre-warming mechanisms cost money for idle capacity
  • Confuses SnapStart-style snapshotting with simply 'caching the response'
  • Assumes mitigation should be applied to every function regardless of latency sensitivity

context