skip to content

What does AWS Lambda SnapStart do to reduce cold-start latency, and what must your function code account for when it is enabled?

level: seniorimportance: nice to knowfreq 30%

answer

  1. initialize once, restore many
  2. a frozen moment replayed
  3. uniqueness stops being unique
  4. sockets do not survive restore
  5. hooks before and after

basics

~20 s

SnapStart runs your initialization once when you publish a version, snapshots the initialized execution environment, and restores from that snapshot instead of re-running init. Code must therefore assume anything captured at init — connections, cached credentials, random seeds — may be stale or duplicated after restore.

solid answer

~50 s

SnapStart attacks cold starts from the other direction to provisioned concurrency. Rather than keeping environments running, it runs the Init phase once at publish time, takes an encrypted snapshot of the initialized environment's memory and disk, caches it, and restores new environments from that image when traffic arrives. The expensive part of initialization — runtime startup, class loading, framework wiring — is paid once instead of per environment, which is why it helps most for runtimes with heavy startup such as the JVM. The catch is that the snapshot is a *frozen moment* replayed many times. Network connections opened during init are dead on restore. Credentials or config cached at init may have expired. Anything unique generated at init — a random seed, a UUID, an instance identifier — is now identical across every restored environment. Java exposes CRaC runtime hooks (`beforeCheckpoint` and `afterRestore` on `org.crac.Resource`) so you can close resources before the snapshot and re-establish them after restore.

code

java · 28 lines
java
import java.sql.Connection;
import org.crac.Context;
import org.crac.Core;
import org.crac.Resource;

public class OrderHandler implements Resource {

    private Connection connection;

    public OrderHandler() {
        Core.getGlobalContext().register(this);
    }

    @Override
    public void beforeCheckpoint(Context<? extends Resource> context) throws Exception {
        // A live socket must not be captured in the snapshot.
        if (connection != null) {
            connection.close();
            connection = null;
        }
    }

    @Override
    public void afterRestore(Context<? extends Resource> context) throws Exception {
        // Every restored environment gets its own fresh connection.
        connection = DataSourceHolder.newConnection();
    }
}

go deeper

for a junior

Know the one-line idea: SnapStart snapshots an already-initialized environment and restores from it, so the initialization work is not repeated for every cold start.

for a middle

Explain that the snapshot is taken when a version is published, that it applies to versions rather than $LATEST, and that it is an alternative to provisioned concurrency rather than a complement.

for a senior

Show you can spot the correctness hazards — frozen randomness, dead sockets, expired cached credentials — and know that runtime hooks such as CRaC's beforeCheckpoint and afterRestore exist to handle them.

for a principal

Own the decision framing: whether the latency win justifies the code constraints and runtime lock-in, versus buying provisioned concurrency, versus reducing initialization work or choosing a lighter runtime outright.

## The idea A cold start is dominated by the Init phase, and for some runtimes Init is expensive in a very repetitive way: the JVM starts, loads hundreds of classes, wires a dependency-injection container, and arrives at a state that is *identical every single time*. SnapStart exploits that. When you publish a function version with SnapStart enabled, Lambda runs Init once, then takes a snapshot of the initialized execution environment — memory and disk — encrypts it, and caches it. When a request needs a new environment, Lambda **restores** from that snapshot rather than re-running Init. The result is a restore-and-resume path instead of an initialize-from-scratch path. For heavyweight startup this is a large reduction in cold-start latency, and unlike provisioned concurrency you are not paying to keep idle environments running. ## What it is not - It is **not** a warm pool. Environments are still created on demand; they just come up faster. Restore is fast, not free. - It does **not** speed up your handler, only the path to reaching it. - It applies to **published versions**, not `$LATEST`, because a snapshot must correspond to fixed code and configuration. Publishing a version triggers the snapshot, which makes deployments take longer. - It cannot be combined with provisioned concurrency on the same function version — the two are alternative answers to the same problem. As of 2025 SnapStart is available for the Java, Python and .NET managed runtimes; support has expanded over time, so confirm the current list rather than reciting one. ## The uniqueness problem This is the part interviewers actually care about, because it is where correctness breaks. A snapshot taken once and restored N times means **everything captured at init is duplicated N times**. The classic hazard is randomness. If your code seeds a pseudorandom generator during init, every restored environment starts from the identical seed and produces the identical sequence — catastrophic if you use it for tokens, nonces or identifiers. AWS mitigated the most common case by making the JVM's `java.security.SecureRandom` reseed after restore, but that only covers the cryptographic generator provided by the runtime; a library or your own code that captured entropy or generated a value at init is still frozen. The same applies to a "unique instance id" computed at init, to a timestamp cached at init, and to any counter you expected to start fresh per environment. ## The staleness problem A snapshot may be restored minutes or days after it was taken, so anything time-sensitive captured at init is suspect: - **Network connections and sockets** opened during init do not survive; the peer closed them long ago. Establish connections lazily on first use, or re-open them in a restore hook. - **Credentials and secrets** fetched at init may have expired or rotated. Cache with a TTL and treat an auth failure as a signal to re-fetch. - **DNS resolutions and cached endpoints** captured at init can point at addresses that no longer serve. - **Ephemeral data in `/tmp`** written at init is part of the snapshot and will reappear identically in every restored environment. ## Runtime hooks Java implements SnapStart on the CRaC (Coordinated Restore at Checkpoint) API, which gives you two callbacks: `beforeCheckpoint`, to release things that must not be frozen, and `afterRestore`, to rebuild them. ```java import org.crac.Context; import org.crac.Core; import org.crac.Resource; public class Handler implements Resource { private Connection conn; public Handler() { Core.getGlobalContext().register(this); } @Override public void beforeCheckpoint(Context<? extends Resource> ctx) throws Exception { if (conn != null) { conn.close(); conn = null; } // don't snapshot a live socket } @Override public void afterRestore(Context<? extends Resource> ctx) throws Exception { conn = openConnection(); // fresh per restored environment } } ``` Other supported runtimes expose equivalent before-snapshot and after-restore hooks; the discipline is identical regardless of language. ## Choosing between SnapStart and provisioned concurrency Both remove Init from the request path, with different economics and different constraints. Provisioned concurrency works with any runtime, gives the most predictable latency because environments are genuinely already running, has no snapshot-correctness caveats — and bills continuously for warm capacity. SnapStart has no idle capacity charge and scales with demand rather than a number you guessed, but is limited to supported runtimes, requires code that tolerates restore, and lengthens the publish step. The pragmatic reading: a JVM function with spiky traffic and no init-time uniqueness hazards is SnapStart's home ground. A function with a hard, predictable latency contract and budget for it is provisioned concurrency's. And a function whose Init phase is fast to begin with needs neither — measure `Init Duration` before buying either.

  • Why is a random seed generated during initialization dangerous under SnapStart?
    The snapshot captures that seed, so every restored environment starts from the identical state and produces the identical sequence of values. If those values are used for tokens, nonces or identifiers you get collisions across environments. AWS reseeds the runtime's SecureRandom after restore, but any generator your own code or a library seeded at init stays frozen — regenerate it in an after-restore hook.
  • How does SnapStart differ from provisioned concurrency in what you pay for?
    Provisioned concurrency bills hourly for environments held ready whether or not requests arrive, so you pay for idle warm capacity you sized in advance. SnapStart has no idle capacity charge: environments are still created on demand, they just restore from a cached snapshot instead of re-running init. The tradeoff is runtime support and code that tolerates being restored.
  • Why does SnapStart require a published version rather than $LATEST?
    A snapshot corresponds to one exact combination of code and configuration; $LATEST changes whenever you deploy, which would immediately invalidate it. Lambda takes the snapshot when you publish a version, so the artifact is immutable. The practical consequence is that publishing gets slower, and you route traffic through an alias pointing at the published version.

saying these in an interview costs you the question

  • Describes SnapStart as keeping environments warm like a pool
  • Assumes database connections opened at init survive a restore
  • Believes values generated during init stay unique per environment
  • Thinks SnapStart can be enabled on $LATEST
  • Expects SnapStart to speed up handler execution itself

context