How do you handle open resources (DB connections, sockets) across a CRaC checkpoint/restore in a Spring app?
answer
- org.crac.Resource: beforeCheckpoint close, afterRestore reopen
- Core.getGlobalContext().register(this) — keep strong ref (weak refs!)
- CRaC may abort checkpoint on open file descriptors
- Hikari + Lettuce + Spring Lifecycle already integrate
- afterRestore fast + idempotent; refresh stale tokens/time
basics
~10 sImplement org.crac.Resource: in beforeCheckpoint() close/drain the resource (e.g. evict a connection pool); in afterRestore() re-open it. Register with Core.getGlobalContext().register(this). Spring already does this for its managed Lifecycle beans and many pools support it.
solid answer
~40 sCRaC will not (and often cannot) safely freeze live OS resources like sockets, files, and DB connections — a connection held across a checkpoint is dead on restore, and CRaC may abort the checkpoint if it detects open descriptors. The contract is the `org.crac.Resource` interface with two callbacks: `beforeCheckpoint(Context)` runs just before the snapshot — close sockets, drain/evict the connection pool, flush buffers; `afterRestore(Context)` runs just after restore — reopen connections, re-warm caches, resume. You register a resource via `Core.getGlobalContext().register(resource)`, keeping a strong reference (the context holds them weakly). In Spring, most of this is handled for you: Spring stops/starts `Lifecycle` beans, and CRaC-aware pools like **HikariCP** and clients like Lettuce integrate to close and reopen on the callbacks. You only implement `Resource` yourself for custom resources Spring doesn't manage.
code
java · 25 linesimport jdk.crac.*;
import org.springframework.stereotype.Component;
@Component
public class WarmCacheResource implements Resource {
private volatile ExpensiveClient client;
public WarmCacheResource() {
// Keep a strong ref to 'this' via the bean; the context holds it weakly.
Core.getGlobalContext().register(this);
}
@Override
public void beforeCheckpoint(Context<? extends Resource> ctx) throws Exception {
// Release OS resources so the checkpoint is clean.
if (client != null) client.close();
}
@Override
public void afterRestore(Context<? extends Resource> ctx) throws Exception {
// Re-acquire quickly and idempotently on the restore path.
client = ExpensiveClient.connect(); // fresh sockets, fresh token
}
}go deeper
Know that connections must be closed before checkpoint and reopened after — CRaC won't keep them alive.
Name the org.crac.Resource callbacks and that Spring/Hikari handle common cases.
Explain the weak-reference registration trap, checkpoint-abort-on-open-fd, and idempotent fast afterRestore.
Reason about ordering, secret/lease refresh on restore, and image-as-sensitive-artifact security posture across a fleet.
## Why resources are the hard part A checkpoint snapshots process memory, but many objects in that memory are just **handles to operating-system resources** — file descriptors, TCP sockets, DB connections. Those OS resources are NOT part of the snapshot in a resumable way: when you restore into a new process (possibly on a new machine, minutes or days later), a socket that was open at checkpoint time points to nothing. Worse, CRaC is deliberately strict: if it detects **open file descriptors** it doesn't know how to handle, it can **fail the checkpoint** rather than produce a broken image. This is the 'coordinated' in Coordinated Restore at Checkpoint — the app must cooperate to release and re-acquire resources. ## The `org.crac.Resource` contract ```java public interface Resource { void beforeCheckpoint(Context<? extends Resource> context) throws Exception; void afterRestore(Context<? extends Resource> context) throws Exception; } ``` - **`beforeCheckpoint`** — called (in reverse registration order) right before the snapshot. Do teardown here: close sockets, evict/close the connection pool, stop background threads, flush anything buffered. - **`afterRestore`** — called (in registration order) right after the process is restored. Do re-acquisition here: reopen connections, restart threads, refresh anything that went stale (leases, tokens, cached 'now'). If `beforeCheckpoint` throws, the checkpoint is aborted — a safety valve to prevent snapshotting a bad state. ## Registration and the reference gotcha You register with: ```java Core.getGlobalContext().register(myResource); ``` **Gotcha:** the global context holds resources via **weak references**. If you register an anonymous/temporary object and keep no strong reference, it can be garbage-collected and its callbacks silently never fire. Always keep the resource as a field/bean. ## What Spring handles automatically You rarely write `Resource` by hand in Spring because: - **Spring lifecycle**: on checkpoint Spring calls `stop()` on `Lifecycle`/`SmartLifecycle` beans and `start()` on restore. Web servers, `@Scheduled` task schedulers, JMS/Kafka listener containers, etc. pause and resume through this. - **CRaC-aware libraries**: **HikariCP** (the default Spring Boot connection pool) has CRaC support — it can suspend the pool and close idle connections `beforeCheckpoint`, and resume `afterRestore`. Redis (Lettuce), and others have similar integration. When these are on the classpath with a CRaC runtime, they register their own resources. So for a typical web app with Hikari + a scheduler, checkpoint/restore of resources 'just works'. You implement `Resource` only for **custom** long-lived resources — a hand-rolled socket client, an embedded broker, a native library handle, a warm in-memory cache you want to rebuild. ## Ordering and idempotency concerns - Order matters: register the pool before things that depend on it so teardown/reacquire nest correctly (Spring's phase ordering handles this for lifecycle beans). - `afterRestore` should be **fast and idempotent** — it runs on the critical restore path; heavy work here erodes the startup win. - Re-fetch anything time- or lease-sensitive: OAuth tokens, TTL-bound cache entries, and any 'startup timestamp' captured before checkpoint are now stale. ## Security note Because the heap is in the image, in-memory secrets and open-credential state are persisted to disk and duplicated across every restored instance. Treat checkpoint images as sensitive artifacts and prefer re-fetching short-lived secrets in `afterRestore`.
- Your custom Resource's callbacks never fire. What's the most likely cause?You registered a resource but kept no strong reference to it. Core.getGlobalContext() holds resources weakly, so it was garbage-collected before the checkpoint. Store it as a field/bean.
- Why might CRaC refuse to take a checkpoint at all?It detected open file descriptors (e.g. live sockets or files) it can't safely snapshot. You must close them in beforeCheckpoint — or the checkpoint aborts to avoid producing an image that would be broken on restore.
- What sensitive-state issue does restoring create?The heap is serialized into the image, so in-memory secrets/tokens are written to disk and copied into every restored instance. Prefer re-fetching short-lived credentials in afterRestore and protect the image files.
saying these in an interview costs you the question
- Thinking CRaC transparently keeps sockets/DB connections alive across restore
- Registering a Resource but keeping no strong reference (weak-ref GC trap)
- Doing heavy work in afterRestore and negating the startup win
- Forgetting cached tokens/timestamps go stale across a checkpoint