Why must resources like connection pools be released before a CRaC checkpoint, and how does Spring coordinate that through the Lifecycle contract?
answer
- open FD with no owner -> CheckpointException
- context registered once as org.crac.Resource
- stop() descending phase, start() ascending phase
- implement SmartLifecycle to participate
- acquire in start(), not constructor/@PostConstruct
basics
~20 sA memory snapshot can't hold live OS handles; CRaC even fails the checkpoint if unmanaged file descriptors are open. Spring registers the context as a CRaC resource and, in beforeCheckpoint, calls stop() on Lifecycle beans (descending phase) to close pools/sockets, then start() on restore.
solid answer
~40 sCRaC snapshots JVM memory, but live OS resources — TCP sockets, file descriptors, JDBC connections — cannot be meaningfully frozen and would be invalid on restore against different infrastructure. CRaC enforces this: any open file descriptor that no `org.crac.Resource` has accounted for makes the checkpoint fail with a `CheckpointException`. Spring's `DefaultLifecycleProcessor` registers the application context as an `org.crac.Resource`. In `beforeCheckpoint` it stops the context lifecycle — calling `stop()` on every running `Lifecycle`/`SmartLifecycle` bean in **descending phase order** (reverse of startup) — so pools drain, the embedded server unbinds, and sockets close. In `afterRestore` it calls `start()` in ascending phase order to reopen them. Framework integrations (Tomcat, HikariCP-backed data sources, messaging listeners) implement this so it's automatic; your own resource-holding beans should implement SmartLifecycle to participate.
code
java · 34 linesimport org.springframework.context.SmartLifecycle;
import org.springframework.stereotype.Component;
// A third-party client that holds a socket but doesn't know about CRaC.
// Wrapping it in SmartLifecycle makes it participate automatically.
@Component
public class MetricsPublisher implements SmartLifecycle {
private volatile boolean running;
private StatsdSocket socket; // owns a live UDP/TCP FD
@Override
public void start() { // afterRestore -> reopen
socket = StatsdSocket.connect("metrics:8125");
running = true;
}
@Override
public void stop() { // beforeCheckpoint -> release
if (socket != null) socket.close();
running = false;
}
@Override
public boolean isRunning() { // must be accurate or stop/start is skipped
return running;
}
// Higher phase = starts later, stops earlier. Tune ordering if needed.
@Override
public int getPhase() {
return SmartLifecycle.DEFAULT_PHASE;
}
}go deeper
Know that pools/sockets must be closed before the snapshot and that Spring uses stop()/start() to do it.
Explain the CheckpointException-on-open-FD rule and that beans implement SmartLifecycle to participate.
Detail phase ordering, where to acquire resources (start(), not constructor), and wrapping non-cooperative libraries.
Reason about auditing an app's FDs before adopting CRaC and designing lifecycle boundaries so every resource owner is checkpoint-safe.
**The core constraint.** A CRaC checkpoint writes the JVM's memory image to disk. Memory can be reconstructed faithfully; **kernel-owned resources cannot**. An open TCP socket is a connection to a peer, tracked by the OS and the remote host — you cannot serialize "a live connection." A file descriptor is an index into the kernel's per-process table. After restore (often on a different host, minutes or hours later), any such handle would be stale, and the peer would have long since dropped the connection. So CRaC takes a hard line: at checkpoint time it inspects open file descriptors, and if it finds ones that no registered `org.crac.Resource` has claimed responsibility for, it aborts with a **`CheckpointException`** (jdk.crac). This is the mechanism that forces "coordination." **The org.crac contract.** The facade defines `org.crac.Resource` with `beforeCheckpoint(Context)` and `afterRestore(Context)`, and you register instances on `org.crac.Core.getGlobalContext()`. Resources are notified in **LIFO** order for `beforeCheckpoint` (last registered releases first) and reverse for `afterRestore`, giving a clean nesting of teardown/reacquire. **Spring's coordination.** Spring registers the `ApplicationContext` (via `DefaultLifecycleProcessor`) as a single `org.crac.Resource`. It then reuses the framework's existing **`Lifecycle`** machinery rather than inventing a new one: - `beforeCheckpoint` → `stopForRestart`-style processing: every running `SmartLifecycle`/`Lifecycle` bean gets `stop()`, ordered by **descending `getPhase()`** (the reverse of start), so higher-phase beans (which started last) stop first — the same discipline used at normal shutdown. Beans use `stop()` to close pools, unbind listeners, flush. - `afterRestore` → `start()` on those beans in **ascending phase**, reopening everything against the restored environment. **What this means for your code.** If a bean owns a socket, a pool, a scheduler thread, or a file handle, make it implement **`SmartLifecycle`** (or plain `Lifecycle`) and do teardown in `stop()`, setup in `start()`. Then it automatically participates in checkpoint/restore — no CRaC-specific code. Spring Boot's integrations already do this for the embedded Tomcat/Jetty/Netty server, JDBC `DataSource` pools, and many messaging listener containers, so a typical web app checkpoints cleanly out of the box. **Edge cases & gotchas.** - A bean that opens a resource in its constructor or `@PostConstruct` but *not* in `start()` won't reopen it after restore — move acquisition into `start()`. - Third-party libraries that hold sockets but don't integrate with Lifecycle/CRaC can cause `CheckpointException` (unmanaged FD). You then wrap them in your own `SmartLifecycle` or a raw `org.crac.Resource`. - Daemon threads blocked in native reads may need to be stopped in `stop()`. - `isRunning()` must accurately reflect state, or Spring may skip `stop()`/`start()`. **Bottom line:** the Lifecycle contract *is* Spring's CRaC coordination protocol — stop() = release before snapshot, start() = reacquire after restore.
- What happens at checkpoint if a library holds an open socket and doesn't integrate with Lifecycle or CRaC?CRaC finds an unmanaged open file descriptor and aborts the checkpoint with a CheckpointException. You fix it by wrapping the library in a SmartLifecycle bean (or a raw org.crac.Resource) that closes the socket in stop()/beforeCheckpoint and reopens it in start()/afterRestore.
- Why is descending phase order used for stop() at checkpoint?It's the same ordering as normal shutdown: beans that started last (higher phase, often depending on lower-phase beans) are stopped first, so a resource is torn down before the things it depends on. start() on restore uses ascending order to rebuild in dependency order.
saying these in an interview costs you the question
- Acquiring resources in the constructor/@PostConstruct instead of start(), so they never reopen after restore
- Believing CRaC silently tolerates open connections
- Returning a hardcoded isRunning() that doesn't reflect actual state, causing stop()/start() to be skipped