skip to content

How do you handle open resources (DB connections, sockets) across a CRaC checkpoint/restore in a Spring app?

level: seniorimportance: must knowfreq 38%

answer

  1. org.crac.Resource: beforeCheckpoint close, afterRestore reopen
  2. Core.getGlobalContext().register(this) — keep strong ref (weak refs!)
  3. CRaC may abort checkpoint on open file descriptors
  4. Hikari + Lettuce + Spring Lifecycle already integrate
  5. afterRestore fast + idempotent; refresh stale tokens/time

basics

~10 s

Implement org.crac.Resource: in beforeCheckpoint() close/drain the resource (e.g. evict a connection pool); in afterRestore() re-open it. Register with Core.getGlobalContext().register(this). Spring already does this for its managed Lifecycle beans and many pools support it.

solid answer

~40 s

CRaC will not (and often cannot) safely freeze live OS resources like sockets, files, and DB connections — a connection held across a checkpoint is dead on restore, and CRaC may abort the checkpoint if it detects open descriptors. The contract is the `org.crac.Resource` interface with two callbacks: `beforeCheckpoint(Context)` runs just before the snapshot — close sockets, drain/evict the connection pool, flush buffers; `afterRestore(Context)` runs just after restore — reopen connections, re-warm caches, resume. You register a resource via `Core.getGlobalContext().register(resource)`, keeping a strong reference (the context holds them weakly). In Spring, most of this is handled for you: Spring stops/starts `Lifecycle` beans, and CRaC-aware pools like **HikariCP** and clients like Lettuce integrate to close and reopen on the callbacks. You only implement `Resource` yourself for custom resources Spring doesn't manage.

code

java · 25 lines
java
import jdk.crac.*;
import org.springframework.stereotype.Component;

@Component
public class WarmCacheResource implements Resource {

    private volatile ExpensiveClient client;

    public WarmCacheResource() {
        // Keep a strong ref to 'this' via the bean; the context holds it weakly.
        Core.getGlobalContext().register(this);
    }

    @Override
    public void beforeCheckpoint(Context<? extends Resource> ctx) throws Exception {
        // Release OS resources so the checkpoint is clean.
        if (client != null) client.close();
    }

    @Override
    public void afterRestore(Context<? extends Resource> ctx) throws Exception {
        // Re-acquire quickly and idempotently on the restore path.
        client = ExpensiveClient.connect();   // fresh sockets, fresh token
    }
}

go deeper

for a junior

Know that connections must be closed before checkpoint and reopened after — CRaC won't keep them alive.

for a middle

Name the org.crac.Resource callbacks and that Spring/Hikari handle common cases.

for a senior

Explain the weak-reference registration trap, checkpoint-abort-on-open-fd, and idempotent fast afterRestore.

for a principal

Reason about ordering, secret/lease refresh on restore, and image-as-sensitive-artifact security posture across a fleet.

## Why resources are the hard part A checkpoint snapshots process memory, but many objects in that memory are just **handles to operating-system resources** — file descriptors, TCP sockets, DB connections. Those OS resources are NOT part of the snapshot in a resumable way: when you restore into a new process (possibly on a new machine, minutes or days later), a socket that was open at checkpoint time points to nothing. Worse, CRaC is deliberately strict: if it detects **open file descriptors** it doesn't know how to handle, it can **fail the checkpoint** rather than produce a broken image. This is the 'coordinated' in Coordinated Restore at Checkpoint — the app must cooperate to release and re-acquire resources. ## The `org.crac.Resource` contract ```java public interface Resource { void beforeCheckpoint(Context<? extends Resource> context) throws Exception; void afterRestore(Context<? extends Resource> context) throws Exception; } ``` - **`beforeCheckpoint`** — called (in reverse registration order) right before the snapshot. Do teardown here: close sockets, evict/close the connection pool, stop background threads, flush anything buffered. - **`afterRestore`** — called (in registration order) right after the process is restored. Do re-acquisition here: reopen connections, restart threads, refresh anything that went stale (leases, tokens, cached 'now'). If `beforeCheckpoint` throws, the checkpoint is aborted — a safety valve to prevent snapshotting a bad state. ## Registration and the reference gotcha You register with: ```java Core.getGlobalContext().register(myResource); ``` **Gotcha:** the global context holds resources via **weak references**. If you register an anonymous/temporary object and keep no strong reference, it can be garbage-collected and its callbacks silently never fire. Always keep the resource as a field/bean. ## What Spring handles automatically You rarely write `Resource` by hand in Spring because: - **Spring lifecycle**: on checkpoint Spring calls `stop()` on `Lifecycle`/`SmartLifecycle` beans and `start()` on restore. Web servers, `@Scheduled` task schedulers, JMS/Kafka listener containers, etc. pause and resume through this. - **CRaC-aware libraries**: **HikariCP** (the default Spring Boot connection pool) has CRaC support — it can suspend the pool and close idle connections `beforeCheckpoint`, and resume `afterRestore`. Redis (Lettuce), and others have similar integration. When these are on the classpath with a CRaC runtime, they register their own resources. So for a typical web app with Hikari + a scheduler, checkpoint/restore of resources 'just works'. You implement `Resource` only for **custom** long-lived resources — a hand-rolled socket client, an embedded broker, a native library handle, a warm in-memory cache you want to rebuild. ## Ordering and idempotency concerns - Order matters: register the pool before things that depend on it so teardown/reacquire nest correctly (Spring's phase ordering handles this for lifecycle beans). - `afterRestore` should be **fast and idempotent** — it runs on the critical restore path; heavy work here erodes the startup win. - Re-fetch anything time- or lease-sensitive: OAuth tokens, TTL-bound cache entries, and any 'startup timestamp' captured before checkpoint are now stale. ## Security note Because the heap is in the image, in-memory secrets and open-credential state are persisted to disk and duplicated across every restored instance. Treat checkpoint images as sensitive artifacts and prefer re-fetching short-lived secrets in `afterRestore`.

  • Your custom Resource's callbacks never fire. What's the most likely cause?
    You registered a resource but kept no strong reference to it. Core.getGlobalContext() holds resources weakly, so it was garbage-collected before the checkpoint. Store it as a field/bean.
  • Why might CRaC refuse to take a checkpoint at all?
    It detected open file descriptors (e.g. live sockets or files) it can't safely snapshot. You must close them in beforeCheckpoint — or the checkpoint aborts to avoid producing an image that would be broken on restore.
  • What sensitive-state issue does restoring create?
    The heap is serialized into the image, so in-memory secrets/tokens are written to disk and copied into every restored instance. Prefer re-fetching short-lived credentials in afterRestore and protect the image files.

saying these in an interview costs you the question

  • Thinking CRaC transparently keeps sockets/DB connections alive across restore
  • Registering a Resource but keeping no strong reference (weak-ref GC trap)
  • Doing heavy work in afterRestore and negating the startup win
  • Forgetting cached tokens/timestamps go stale across a checkpoint

context