In a Functions-as-a-Service (FaaS) platform like AWS Lambda, what is the difference between a 'cold start' and a 'warm start' when a function is invoked, and why does it happen?
answer
- cold = new container/microVM boot + init code
- warm = reuse of live environment, skip init
- idle window ~5-45 min before reclaim
- heavier runtime (JVM/.NET) = slower cold start
- cold starts hit tail latency (p99), not median
basics
~20 sA cold start is when the cloud has to build a fresh environment to run your function, which takes extra time. A warm start reuses an environment left over from an earlier call, so it responds almost instantly.
solid answer
~50 sFaaS platforms don't keep your code running continuously; they create an execution environment - a lightweight container or microVM - only when a request needs one. A cold start is that first invocation in a new environment: the platform has to fetch your deployment package, boot the language runtime, and run any module-level initialization code before your handler function even starts, adding anywhere from tens of milliseconds to a few seconds depending on runtime and package size. A warm start happens when a later request lands on an environment that's still alive from a previous invocation - all that setup is skipped and only the handler body runs, so latency drops to single-digit milliseconds. Providers keep idle environments around for a limited window (roughly 5-45 minutes) hoping to reuse them, then reclaim the resources if no traffic arrives.
go deeper
Should be able to state the basic definition - cold start builds a new environment, warm start reuses one - and give a rough sense that cold starts are slower. Doesn't need to know mitigation techniques or provider-specific mechanics.
Should additionally know that cold starts show up as tail-latency spikes, that heavier runtimes (JVM, .NET) cold-start slower than lightweight ones (Go, Node, Python), and that idle environments get reclaimed after a time window.
Should be able to explain the init-phase mechanics (runtime boot, static initializers), name at least one concrete mitigation (provisioned/min concurrency, memory bump), and reason about when a cold start is actually a business problem worth paying to fix versus an acceptable trade-off.
Should reason about cold starts as a systemic architectural concern - e.g., choosing runtimes/languages for latency-sensitive FaaS workloads, designing SLAs around p99 rather than average latency, deciding when FaaS is the wrong tool entirely versus a long-running container service, and weighing the cost of provisioned concurrency against traffic patterns at scale.
## What the platform does when a request arrives To understand cold starts you first need to understand that a FaaS platform does not keep a server sitting around waiting for your code to run. Instead, when a request or event arrives, the platform's control plane finds (or creates) an **execution environment** - typically a lightweight, sandboxed container or a microVM such as AWS's Firecracker - and loads your function's code into it. If no existing environment is available (the very first call, or a burst of concurrent calls that exceeds the number of already-running environments), the platform must perform what's called a **cold start**: - it downloads or mounts your deployment package (zip file or container image) - it starts the language runtime process (the JVM, the Node.js V8 engine, the Python interpreter, the .NET CLR, and so on) - it then executes any code that lives outside your handler function - imports, static initializers, SDK client construction, configuration loading Only after all of that 'init phase' completes does the platform invoke your actual handler with the event payload. All of this adds latency on top of your function's own execution time, and that added latency is what people mean by 'the cold start penalty' or 'cold start tax.' ## The warm start A **warm start** is the opposite case: a later invocation is routed to an environment that is still alive from handling an earlier request. Because the runtime is already booted and your init-phase code has already run, the platform can go straight to invoking the handler, so the extra latency disappears and you're left with just your business logic's execution time, often in the single-digit-to-low-double-digit milliseconds. This is why the same function can log 20ms on one call and 900ms on another with no code difference - the difference is purely whether the platform had a warm environment sitting ready. ## Why the model exists This model exists because it is what makes serverless economically and operationally attractive: providers only want to pay for (and bill you for) compute that is actively doing work, and they want to scale environments up and down **elastically** without you provisioning capacity. If every function kept a server running permanently, you'd be back to paying for idle EC2 instances and managing autoscaling groups yourself - the whole point of FaaS is that the platform handles that elasticity for you, and the cost of that elasticity is the occasional cold start when demand grows or when an environment has been idle long enough to be reclaimed. ## The trade-off The trade-off is **latency predictability** versus **cost and operational simplicity**. Cold starts are worse the 'heavier' your runtime is: - a Go or Rust or Node.js function with a small deployment package might cold-start in tens of milliseconds - a JVM-based Java or .NET function with a large dependency tree and a big deployment artifact can take one to several seconds, because the JVM/CLR itself is slow to boot and JIT-warm Higher configured memory usually also grants more CPU, which speeds up cold starts, so bumping memory is a common cheap mitigation. VPC-attached functions historically had much worse cold starts too, because the platform had to attach an elastic network interface before the function could reach VPC resources - most providers have since optimized this away, but it's a real historical failure mode worth knowing. ## Where it shows up in production In production, cold starts show up as latency spikes concentrated in your p95/p99/p99.9 percentiles rather than your median or average - a dashboard that looks fine on p50 latency can be hiding a nasty tail. They bite hardest in three scenarios: - **(1) low-traffic or spiky functions**, where environments idle out between calls and every burst starts with a wave of cold environments (a 'cold start storm' during a traffic spike, since the platform must provision N new environments simultaneously to handle N concurrent new requests) - **(2) synchronous, latency-sensitive paths** like an API Gateway-fronted Lambda backing a user-facing web request, where a multi-second cold start can breach a client timeout or an SLA - **(3) after a deployment**, where every environment for the new code version starts cold because the platform can't reuse containers running the old code ## A concrete example A concrete real-world example: a team runs a Java Spring Boot function on AWS Lambda behind API Gateway for a checkout API. Under steady low traffic, most requests land on cold environments because idle containers get recycled between sporadic calls, and p99 latency is dominated by JVM boot plus Spring context initialization (often 2-5 seconds) rather than the actual checkout logic. The team mitigates this with AWS Lambda Provisioned Concurrency, which keeps a fixed number of environments pre-initialized and warm at all times, trading a flat hourly cost for eliminating the cold-start tail on the critical path - a decision that only makes sense once cold starts are shown to be violating a real latency SLA.
- Why does increasing a Lambda function's configured memory often reduce its cold start time, even if the function doesn't need that much memory to run?AWS Lambda allocates CPU power proportionally to configured memory, so a higher memory setting also grants more vCPU and network bandwidth during the init phase, which speeds up runtime boot and any heavy static initialization (like loading a big Spring context). It's a common, cheap lever teams pull before reaching for provisioned concurrency, though it does raise the per-invocation cost.
- What is a 'cold start storm' and when does it happen?It's when a sudden spike in concurrent traffic forces the platform to provision many new execution environments simultaneously because the existing warm pool can't cover the concurrency, so a large fraction of requests in that burst all pay the cold-start penalty at once. It's common right after a deploy or during a flash-traffic event, and it can cascade into timeouts if downstream systems (like a database connection pool) also get hit with a burst of new connections at once.
Like a pop-up food stall that has to be assembled from a folded cart before it can serve the first customer of the day (cold start), but once it's set up it can serve the next dozen customers in line almost instantly (warm start) - until it's been idle so long the stall gets folded back up.
saying these in an interview costs you the question
- Claims FaaS functions 'run continuously' waiting for requests
- Doesn't know cold starts are about environment setup, thinks it's just 'slow code'
- Assumes every invocation is a cold start
- No awareness that runtime choice (JVM vs Go/Node) changes cold start magnitude
- Confuses cold start with network latency or database latency