skip to content

What are the ways to trigger a CRaC checkpoint with Spring, and how does spring.context.checkpoint=onRefresh differ from an on-demand checkpoint?

level: middleimportance: should knowfreq 22%

answer

  1. onRefresh = auto checkpoint at end of refresh, beans not started
  2. jcmd <pid> JDK.checkpoint = on-demand, app fully running
  3. -XX:CRaCCheckpointTo to arm, -XX:CRaCRestoreFrom to restore
  4. onRefresh = cold JIT; on-demand = warmed JIT
  5. restore doesn't name a main class

basics

~10 s

Two ways: set spring.context.checkpoint=onRefresh to auto-checkpoint during startup (context refreshed but lifecycle beans not yet started), or run jcmd <pid> JDK.checkpoint on a fully running app. Restore both with java -XX:CRaCRestoreFrom=<dir>.

solid answer

~40 s

There are two trigger modes. (1) **Automatic, at startup**: the Spring property `spring.context.checkpoint=onRefresh` makes Spring take a checkpoint at the end of context refresh — beans are instantiated but SmartLifecycle beans haven't been started, so no pools/listeners are open yet. That yields a clean 'wired but not connected' snapshot; on restore Spring runs start(). (2) **On demand, while running**: launch with `-XX:CRaCCheckpointTo=<dir>` and later issue `jcmd <pid> JDK.checkpoint`; here the app is live, so Spring calls stop() on lifecycle beans to release connections, then snapshot. Restore either with `java -XX:CRaCRestoreFrom=<dir>`. onRefresh is convenient for build-time snapshots but skips runtime warm-up; on-demand lets you snapshot after real traffic has warmed the JIT and caches.

code

java · 20 lines
java
// Mode 1 — automatic checkpoint during startup:
//   java -Dspring.context.checkpoint=onRefresh \
//        -XX:CRaCCheckpointTo=./cr -jar app.jar
//   (Spring snapshots at end of refresh; SmartLifecycle beans not yet started)

// Mode 2 — on-demand checkpoint of a warmed, running app:
//   java -XX:CRaCCheckpointTo=./cr -jar app.jar   # start & warm up
//   jcmd <pid> JDK.checkpoint                      # Spring stop()s, then snapshot

// Restore (identical for both, no main class needed):
//   java -XX:CRaCRestoreFrom=./cr

// Programmatic on-demand trigger from inside the app:
import org.crac.Core;

public final class Checkpointer {
    public static void trigger() throws Exception {
        Core.getGlobalContext().checkpointRestore();
    }
}

go deeper

for a junior

Know that a checkpoint can be automatic at startup or triggered manually with jcmd.

for a middle

Contrast onRefresh (before lifecycle start, cold) vs on-demand jcmd (fully running, warm) and the JVM flags involved.

for a senior

Discuss warm-up strategy: script traffic then jcmd checkpoint for peak-perf snapshots.

for a principal

Weigh build-time reproducibility (onRefresh) against warmed-image throughput (on-demand) in a deployment pipeline.

Spring exposes **two distinct checkpoint triggers**, and the difference matters because it changes *what state* ends up in the snapshot. ## The two triggers **1. Automatic checkpoint at startup — `spring.context.checkpoint=onRefresh`.** This is a Spring property (system property `-Dspring.context.checkpoint=onRefresh` or `SPRING_CONTEXT_CHECKPOINT=onRefresh`). When set, Spring initiates a checkpoint at the **end of the context refresh phase**. At that point: - all singleton beans have been created and wired, but - the `SmartLifecycle` auto-start beans have **not yet been started** — so the embedded web server hasn't bound its port and connection pools haven't opened. The snapshot is therefore "assembled but not connected," which sidesteps most open-resource problems automatically. You still must launch the JVM with `-XX:CRaCCheckpointTo=<dir>`. Typically you do this once (in a build/CI step or a throwaway boot) to produce an image. On **restore**, Spring proceeds past refresh and calls `start()`, so the server binds and pools open against live infrastructure. Downside: because the app never actually served traffic before the snapshot, the JIT and any lazy caches are not warmed. **2. On-demand checkpoint — `jcmd <pid> JDK.checkpoint`.** 1. You start the app normally (with `-XX:CRaCCheckpointTo=<dir>`), 2. let it run — ideally after sending warm-up traffic so the JIT compiles hot paths and caches fill — 3. then trigger a checkpoint from outside via `jcmd <pid> JDK.checkpoint` (or programmatically via `org.crac.Core.getGlobalContext().checkpointRestore()`). Now the app is **fully running with open resources**, so Spring's `beforeCheckpoint` runs `stop()` on lifecycle beans (in descending phase order) to close pools and sockets before the memory image is written. The process typically **exits** after the checkpoint is written. On **restore** (`java -XX:CRaCRestoreFrom=<dir>`), `afterRestore` runs `start()` to reopen everything. This mode captures a genuinely warmed JVM — best peak performance immediately after restore. ## Restore is identical for both `java -XX:CRaCRestoreFrom=<dir>` — you don't even name a main class, because the process is reconstructed from the image. ## Choosing - Use `onRefresh` when you want a simple, reproducible build-time snapshot and can tolerate cold JIT after restore. - Use on-demand when peak throughput immediately after restore matters and you can script a warm-up + checkpoint step. Both require a CRaC-enabled JDK. ## Gotcha With `onRefresh`, code in `ApplicationRunner`/`CommandLineRunner` and lifecycle `start()` runs *after* restore, not before the snapshot — so anything you do there executes on every restored instance, which is usually what you want for reconnecting.

  • With onRefresh, why are there usually no open sockets to worry about at checkpoint time?
    Because the checkpoint fires at the end of context refresh, before SmartLifecycle auto-start beans run. The web server hasn't bound its port and pools haven't opened, so the snapshot is 'wired but not connected'.
  • Which mode gives better throughput right after restore, and why?
    On-demand, because you can warm the app with real traffic first — the JIT compiles hot paths and caches fill — and that warmed state is captured in the snapshot. onRefresh snapshots a cold JVM.

saying these in an interview costs you the question

  • Claiming onRefresh checkpoints after the app is serving traffic (it's before lifecycle start)
  • Thinking you must pass the main class again on restore
  • Assuming onRefresh gives a warmed JIT — it doesn't

context