What are the ways to trigger a CRaC checkpoint with Spring, and how does spring.context.checkpoint=onRefresh differ from an on-demand checkpoint?
answer
- onRefresh = auto checkpoint at end of refresh, beans not started
- jcmd <pid> JDK.checkpoint = on-demand, app fully running
- -XX:CRaCCheckpointTo to arm, -XX:CRaCRestoreFrom to restore
- onRefresh = cold JIT; on-demand = warmed JIT
- restore doesn't name a main class
basics
~10 sTwo ways: set spring.context.checkpoint=onRefresh to auto-checkpoint during startup (context refreshed but lifecycle beans not yet started), or run jcmd <pid> JDK.checkpoint on a fully running app. Restore both with java -XX:CRaCRestoreFrom=<dir>.
solid answer
~40 sThere are two trigger modes. (1) **Automatic, at startup**: the Spring property `spring.context.checkpoint=onRefresh` makes Spring take a checkpoint at the end of context refresh — beans are instantiated but SmartLifecycle beans haven't been started, so no pools/listeners are open yet. That yields a clean 'wired but not connected' snapshot; on restore Spring runs start(). (2) **On demand, while running**: launch with `-XX:CRaCCheckpointTo=<dir>` and later issue `jcmd <pid> JDK.checkpoint`; here the app is live, so Spring calls stop() on lifecycle beans to release connections, then snapshot. Restore either with `java -XX:CRaCRestoreFrom=<dir>`. onRefresh is convenient for build-time snapshots but skips runtime warm-up; on-demand lets you snapshot after real traffic has warmed the JIT and caches.
code
java · 20 lines// Mode 1 — automatic checkpoint during startup:
// java -Dspring.context.checkpoint=onRefresh \
// -XX:CRaCCheckpointTo=./cr -jar app.jar
// (Spring snapshots at end of refresh; SmartLifecycle beans not yet started)
// Mode 2 — on-demand checkpoint of a warmed, running app:
// java -XX:CRaCCheckpointTo=./cr -jar app.jar # start & warm up
// jcmd <pid> JDK.checkpoint # Spring stop()s, then snapshot
// Restore (identical for both, no main class needed):
// java -XX:CRaCRestoreFrom=./cr
// Programmatic on-demand trigger from inside the app:
import org.crac.Core;
public final class Checkpointer {
public static void trigger() throws Exception {
Core.getGlobalContext().checkpointRestore();
}
}go deeper
Know that a checkpoint can be automatic at startup or triggered manually with jcmd.
Contrast onRefresh (before lifecycle start, cold) vs on-demand jcmd (fully running, warm) and the JVM flags involved.
Discuss warm-up strategy: script traffic then jcmd checkpoint for peak-perf snapshots.
Weigh build-time reproducibility (onRefresh) against warmed-image throughput (on-demand) in a deployment pipeline.
Spring exposes **two distinct checkpoint triggers**, and the difference matters because it changes *what state* ends up in the snapshot. ## The two triggers **1. Automatic checkpoint at startup — `spring.context.checkpoint=onRefresh`.** This is a Spring property (system property `-Dspring.context.checkpoint=onRefresh` or `SPRING_CONTEXT_CHECKPOINT=onRefresh`). When set, Spring initiates a checkpoint at the **end of the context refresh phase**. At that point: - all singleton beans have been created and wired, but - the `SmartLifecycle` auto-start beans have **not yet been started** — so the embedded web server hasn't bound its port and connection pools haven't opened. The snapshot is therefore "assembled but not connected," which sidesteps most open-resource problems automatically. You still must launch the JVM with `-XX:CRaCCheckpointTo=<dir>`. Typically you do this once (in a build/CI step or a throwaway boot) to produce an image. On **restore**, Spring proceeds past refresh and calls `start()`, so the server binds and pools open against live infrastructure. Downside: because the app never actually served traffic before the snapshot, the JIT and any lazy caches are not warmed. **2. On-demand checkpoint — `jcmd <pid> JDK.checkpoint`.** 1. You start the app normally (with `-XX:CRaCCheckpointTo=<dir>`), 2. let it run — ideally after sending warm-up traffic so the JIT compiles hot paths and caches fill — 3. then trigger a checkpoint from outside via `jcmd <pid> JDK.checkpoint` (or programmatically via `org.crac.Core.getGlobalContext().checkpointRestore()`). Now the app is **fully running with open resources**, so Spring's `beforeCheckpoint` runs `stop()` on lifecycle beans (in descending phase order) to close pools and sockets before the memory image is written. The process typically **exits** after the checkpoint is written. On **restore** (`java -XX:CRaCRestoreFrom=<dir>`), `afterRestore` runs `start()` to reopen everything. This mode captures a genuinely warmed JVM — best peak performance immediately after restore. ## Restore is identical for both `java -XX:CRaCRestoreFrom=<dir>` — you don't even name a main class, because the process is reconstructed from the image. ## Choosing - Use `onRefresh` when you want a simple, reproducible build-time snapshot and can tolerate cold JIT after restore. - Use on-demand when peak throughput immediately after restore matters and you can script a warm-up + checkpoint step. Both require a CRaC-enabled JDK. ## Gotcha With `onRefresh`, code in `ApplicationRunner`/`CommandLineRunner` and lifecycle `start()` runs *after* restore, not before the snapshot — so anything you do there executes on every restored instance, which is usually what you want for reconnecting.
- With onRefresh, why are there usually no open sockets to worry about at checkpoint time?Because the checkpoint fires at the end of context refresh, before SmartLifecycle auto-start beans run. The web server hasn't bound its port and pools haven't opened, so the snapshot is 'wired but not connected'.
- Which mode gives better throughput right after restore, and why?On-demand, because you can warm the app with real traffic first — the JIT compiles hot paths and caches fill — and that warmed state is captured in the snapshot. onRefresh snapshots a cold JVM.
saying these in an interview costs you the question
- Claiming onRefresh checkpoints after the app is serving traffic (it's before lifecycle start)
- Thinking you must pass the main class again on restore
- Assuming onRefresh gives a warmed JIT — it doesn't