skip to content

Why must resources like connection pools be released before a CRaC checkpoint, and how does Spring coordinate that through the Lifecycle contract?

level: seniorimportance: should knowfreq 18%

answer

  1. open FD with no owner -> CheckpointException
  2. context registered once as org.crac.Resource
  3. stop() descending phase, start() ascending phase
  4. implement SmartLifecycle to participate
  5. acquire in start(), not constructor/@PostConstruct

basics

~20 s

A memory snapshot can't hold live OS handles; CRaC even fails the checkpoint if unmanaged file descriptors are open. Spring registers the context as a CRaC resource and, in beforeCheckpoint, calls stop() on Lifecycle beans (descending phase) to close pools/sockets, then start() on restore.

solid answer

~40 s

CRaC snapshots JVM memory, but live OS resources — TCP sockets, file descriptors, JDBC connections — cannot be meaningfully frozen and would be invalid on restore against different infrastructure. CRaC enforces this: any open file descriptor that no `org.crac.Resource` has accounted for makes the checkpoint fail with a `CheckpointException`. Spring's `DefaultLifecycleProcessor` registers the application context as an `org.crac.Resource`. In `beforeCheckpoint` it stops the context lifecycle — calling `stop()` on every running `Lifecycle`/`SmartLifecycle` bean in **descending phase order** (reverse of startup) — so pools drain, the embedded server unbinds, and sockets close. In `afterRestore` it calls `start()` in ascending phase order to reopen them. Framework integrations (Tomcat, HikariCP-backed data sources, messaging listeners) implement this so it's automatic; your own resource-holding beans should implement SmartLifecycle to participate.

code

java · 34 lines
java
import org.springframework.context.SmartLifecycle;
import org.springframework.stereotype.Component;

// A third-party client that holds a socket but doesn't know about CRaC.
// Wrapping it in SmartLifecycle makes it participate automatically.
@Component
public class MetricsPublisher implements SmartLifecycle {

    private volatile boolean running;
    private StatsdSocket socket; // owns a live UDP/TCP FD

    @Override
    public void start() {                 // afterRestore -> reopen
        socket = StatsdSocket.connect("metrics:8125");
        running = true;
    }

    @Override
    public void stop() {                  // beforeCheckpoint -> release
        if (socket != null) socket.close();
        running = false;
    }

    @Override
    public boolean isRunning() {          // must be accurate or stop/start is skipped
        return running;
    }

    // Higher phase = starts later, stops earlier. Tune ordering if needed.
    @Override
    public int getPhase() {
        return SmartLifecycle.DEFAULT_PHASE;
    }
}

go deeper

for a junior

Know that pools/sockets must be closed before the snapshot and that Spring uses stop()/start() to do it.

for a middle

Explain the CheckpointException-on-open-FD rule and that beans implement SmartLifecycle to participate.

for a senior

Detail phase ordering, where to acquire resources (start(), not constructor), and wrapping non-cooperative libraries.

for a principal

Reason about auditing an app's FDs before adopting CRaC and designing lifecycle boundaries so every resource owner is checkpoint-safe.

**The core constraint.** A CRaC checkpoint writes the JVM's memory image to disk. Memory can be reconstructed faithfully; **kernel-owned resources cannot**. An open TCP socket is a connection to a peer, tracked by the OS and the remote host — you cannot serialize "a live connection." A file descriptor is an index into the kernel's per-process table. After restore (often on a different host, minutes or hours later), any such handle would be stale, and the peer would have long since dropped the connection. So CRaC takes a hard line: at checkpoint time it inspects open file descriptors, and if it finds ones that no registered `org.crac.Resource` has claimed responsibility for, it aborts with a **`CheckpointException`** (jdk.crac). This is the mechanism that forces "coordination." **The org.crac contract.** The facade defines `org.crac.Resource` with `beforeCheckpoint(Context)` and `afterRestore(Context)`, and you register instances on `org.crac.Core.getGlobalContext()`. Resources are notified in **LIFO** order for `beforeCheckpoint` (last registered releases first) and reverse for `afterRestore`, giving a clean nesting of teardown/reacquire. **Spring's coordination.** Spring registers the `ApplicationContext` (via `DefaultLifecycleProcessor`) as a single `org.crac.Resource`. It then reuses the framework's existing **`Lifecycle`** machinery rather than inventing a new one: - `beforeCheckpoint` → `stopForRestart`-style processing: every running `SmartLifecycle`/`Lifecycle` bean gets `stop()`, ordered by **descending `getPhase()`** (the reverse of start), so higher-phase beans (which started last) stop first — the same discipline used at normal shutdown. Beans use `stop()` to close pools, unbind listeners, flush. - `afterRestore` → `start()` on those beans in **ascending phase**, reopening everything against the restored environment. **What this means for your code.** If a bean owns a socket, a pool, a scheduler thread, or a file handle, make it implement **`SmartLifecycle`** (or plain `Lifecycle`) and do teardown in `stop()`, setup in `start()`. Then it automatically participates in checkpoint/restore — no CRaC-specific code. Spring Boot's integrations already do this for the embedded Tomcat/Jetty/Netty server, JDBC `DataSource` pools, and many messaging listener containers, so a typical web app checkpoints cleanly out of the box. **Edge cases & gotchas.** - A bean that opens a resource in its constructor or `@PostConstruct` but *not* in `start()` won't reopen it after restore — move acquisition into `start()`. - Third-party libraries that hold sockets but don't integrate with Lifecycle/CRaC can cause `CheckpointException` (unmanaged FD). You then wrap them in your own `SmartLifecycle` or a raw `org.crac.Resource`. - Daemon threads blocked in native reads may need to be stopped in `stop()`. - `isRunning()` must accurately reflect state, or Spring may skip `stop()`/`start()`. **Bottom line:** the Lifecycle contract *is* Spring's CRaC coordination protocol — stop() = release before snapshot, start() = reacquire after restore.

  • What happens at checkpoint if a library holds an open socket and doesn't integrate with Lifecycle or CRaC?
    CRaC finds an unmanaged open file descriptor and aborts the checkpoint with a CheckpointException. You fix it by wrapping the library in a SmartLifecycle bean (or a raw org.crac.Resource) that closes the socket in stop()/beforeCheckpoint and reopens it in start()/afterRestore.
  • Why is descending phase order used for stop() at checkpoint?
    It's the same ordering as normal shutdown: beans that started last (higher phase, often depending on lower-phase beans) are stopped first, so a resource is torn down before the things it depends on. start() on restore uses ascending order to rebuild in dependency order.

saying these in an interview costs you the question

  • Acquiring resources in the constructor/@PostConstruct instead of start(), so they never reopen after restore
  • Believing CRaC silently tolerates open connections
  • Returning a hardcoded isRunning() that doesn't reflect actual state, causing stop()/start() to be skipped

context