skip to content

How does Spring Boot 3.2+ integrate with CRaC, and how do you trigger an automatic checkpoint?

level: middleimportance: should knowfreq 40%

answer

  1. org.crac façade + CRaC-enabled JDK
  2. Spring stops/starts Lifecycle beans on checkpoint/restore
  3. spring.context.checkpoint.restore=onRefresh = auto, pre-traffic
  4. jcmd <pid> JDK.checkpoint = on-demand, JIT-warmed
  5. -XX:CRaCCheckpointTo / -XX:CRaCRestoreFrom

basics

~10 s

Add the org.crac dependency and run on a CRaC-enabled JDK. Spring drives its Lifecycle beans on checkpoint/restore. You can auto-checkpoint at startup with -Dspring.context.checkpoint.restore=onRefresh, or trigger manually with jcmd JDK.checkpoint.

solid answer

~40 s

Spring Framework 6.1 / Boot 3.2 add first-class CRaC support. You run on a CRaC-enabled JDK and add the `org.crac:crac` API on the classpath. Spring registers a CRaC `Resource` that, on checkpoint, calls `stop()` on all `Lifecycle`/`SmartLifecycle` beans (pausing thread pools, closing pools where they opt in), and on restore calls `start()` again. There are two modes. **Automatic / on-refresh**: set `-Dspring.context.checkpoint.restore=onRefresh`; the app initializes the context, then Spring triggers a checkpoint automatically right after refresh (before the web server starts accepting traffic), producing an image of an initialized-but-not-serving app. **On-demand**: start the app normally, warm it up with real traffic, then run `jcmd <pid> JDK.checkpoint` externally. You launch checkpointing with `-XX:CRaCCheckpointTo=<dir>` and restore with `-XX:CRaCRestoreFrom=<dir>`.

code

java · 16 lines
java
// build.gradle dependency (compile against the portable facade):
//   implementation("org.crac:crac:1.4.0")

// Automatic checkpoint after context refresh, before serving traffic:
//   java -XX:CRaCCheckpointTo=/opt/crac-image \
//        -Dspring.context.checkpoint.restore=onRefresh \
//        -jar app.jar
// -> JVM checkpoints itself right after ApplicationContext refresh, then exits.

// On-demand, JIT-warmed checkpoint:
//   java -XX:CRaCCheckpointTo=/opt/crac-image -jar app.jar   # start normally
//   # ... push warm-up traffic ...
//   jcmd <pid> JDK.checkpoint                                 # snapshot + exit

// Restore (either image) — no classpath/main needed:
//   java -XX:CRaCRestoreFrom=/opt/crac-image

go deeper

for a junior

Know you add org.crac and need a CRaC JDK; know a checkpoint can be triggered.

for a middle

Explain both trigger modes (onRefresh vs jcmd) and the CRaCCheckpointTo/RestoreFrom flags.

for a senior

Explain the JIT-warmth trade-off between the two modes and how Spring reuses Lifecycle beans.

for a principal

Design the build/deploy pipeline: where warm-up traffic comes from, image security, and choosing onRefresh vs warmed per SLO.

## The moving parts There are three layers involved: 1. **The JDK** must be a **CRaC-enabled build** (e.g. Azul Zulu with CRaC, BellSoft Liberica with CRaC). A stock JDK has no checkpoint support. 2. **The `org.crac` library** (`org.crac:crac`) — a thin, portable façade over the `jdk.crac` API. Frameworks compile against `org.crac` so they build on any JDK; at runtime it delegates to the real `jdk.crac` if present, otherwise no-ops. Spring depends on this. 3. **Spring's integration** — Spring Framework 6.1+ registers a global CRaC `Resource` and wires checkpoint/restore into its `Lifecycle` machinery. ## What Spring does on checkpoint and restore Spring's `ApplicationContext` already manages **`Lifecycle`** and **`SmartLifecycle`** beans (things that can be `start()`ed and `stop()`ped: message listener containers, task schedulers, the web server, etc.). When a checkpoint is requested, Spring: - **Before checkpoint**: calls `stop()` on lifecycle beans (in reverse phase order). This pauses schedulers, stops listener containers, and lets components release resources they can't safely hold across a snapshot. - **After restore**: calls `start()` again, resuming everything. Beans that need finer control implement the `org.crac.Resource` interface directly (`beforeCheckpoint`/`afterRestore`) — covered in the resource-handling question. ## Two ways to take the checkpoint ### 1. Automatic checkpoint on refresh Set the property: ``` -Dspring.context.checkpoint.restore=onRefresh ``` With this, Spring initializes the whole `ApplicationContext` (all beans created, auto-config done) and then **immediately triggers a checkpoint right after context refresh, before the embedded web server begins accepting requests**. The resulting image is a fully-initialized application that has not served any traffic. Benefit: deterministic, easy to bake into a Docker build. Limitation: because no traffic has flowed, the **JIT is not warmed** — you get fast startup but not the peak-throughput-from-first-request benefit. > Note: this property historically was `spring.context.checkpoint.restore`; in newer versions the automatic-checkpoint value is `onRefresh`. The checkpoint itself is still performed by running the JVM with `-XX:CRaCCheckpointTo=<dir>`. ### 2. On-demand checkpoint (warmed) Start the app normally with `-XX:CRaCCheckpointTo=<dir>`, drive real or synthetic load to warm the JIT, then from outside trigger: ``` jcmd <pid> JDK.checkpoint ``` The JVM writes the image and exits. This image is **JIT-warmed**, so restore reaches peak throughput immediately — the fullest expression of CRaC's value. The cost is a more involved build pipeline (you must run a representative workload during the build). ## Restoring Regardless of how the checkpoint was made, restore with: ``` java -XX:CRaCRestoreFrom=<dir> ``` No main class or classpath is needed — everything is in the image. Spring runs its `afterRestore` handling and the app is serving within milliseconds. ## Gotchas - **Open resources must be closed before checkpoint.** An open socket or DB connection held across a checkpoint will be invalid on restore. CRaC even fails the checkpoint if it detects certain open file descriptors, forcing you to handle them. - **Secrets and randomness.** State captured at checkpoint (loaded config, in-memory secrets, RNG seeds) is duplicated across every restored instance — a security consideration. - **The image contains your heap.** Anything sensitive in memory is on disk; protect the image files. - **Time and identity.** System time jumps and the process starts fresh; anything caching 'now' or a startup timestamp may be stale.

  • What is the throughput difference between an onRefresh checkpoint and an on-demand warmed checkpoint?
    onRefresh checkpoints before any traffic, so the JIT is cold — you get fast startup but must still warm up under load. An on-demand checkpoint taken after warm-up traffic captures JIT-compiled code, so it hits peak throughput immediately on restore.
  • Why does Spring hook into Lifecycle/SmartLifecycle rather than inventing a new mechanism?
    Because those beans already model start/stop of stateful components (web server, listener containers, schedulers). Reusing them means existing components can pause and resume across a checkpoint without new APIs; only components needing resource close/reopen implement org.crac.Resource directly.

saying these in an interview costs you the question

  • Thinking any stock JDK can checkpoint (it needs a CRaC-enabled build)
  • Believing onRefresh checkpoints capture a JIT-warmed app
  • Confusing -XX:CRaCCheckpointTo (create) with -XX:CRaCRestoreFrom (restore)
  • Assuming Spring closes DB connections/sockets automatically for you in all cases

context