skip to content

How do you make a custom resource participate in Spring's CRaC checkpoint/restore, and when would you use org.crac.Resource directly instead of SmartLifecycle?

level: seniorimportance: nice to knowfreq 12%

answer

  1. SmartLifecycle = default path, stop()/start()
  2. org.crac.Resource = beforeCheckpoint/afterRestore + register on Core.getGlobalContext()
  3. direct Resource for non-lifecycle concerns (SecureRandom, DNS, non-Spring code)
  4. hold strong reference to registered Resource
  5. acquire in start(), never constructor

basics

~20 s

Usually implement SmartLifecycle: close in stop(), reopen in start(), and Spring handles the CRaC callbacks. Use org.crac.Resource directly (register on Core.getGlobalContext()) for non-Spring components or when you need CRaC hooks decoupled from Spring's lifecycle phases.

solid answer

~40 s

The idiomatic path is a **SmartLifecycle** bean: acquire in `start()`, release in `stop()`, report state in `isRunning()`, and optionally order via `getPhase()`. Spring's context is a registered CRaC resource, so it drives your bean's stop()/start() around the snapshot automatically — no CRaC dependency in your code. Reach for **`org.crac.Resource`** directly when the component isn't a Spring-managed lifecycle bean, when a library exposes a resource outside the ApplicationContext, or when you need checkpoint hooks that run independently of Spring's lifecycle phase ordering (e.g. flushing an app-wide cache, reseeding SecureRandom, resetting a metrics client). You implement `beforeCheckpoint(Context)`/`afterRestore(Context)` and register on `org.crac.Core.getGlobalContext()`. Prefer SmartLifecycle inside Spring apps for testability and clean shutdown reuse; drop to org.crac.Resource only for the gaps.

code

java · 26 lines
java
import org.crac.Context;
import org.crac.Core;
import org.crac.Resource;
import org.springframework.stereotype.Component;

// Cross-cutting concern that isn't a 'start/stop' lifecycle: reseed randomness
// after restore so cloned instances don't share the pre-snapshot RNG state.
@Component
public class SecureRandomReseeder implements Resource {

    public SecureRandomReseeder() {
        // Keep a strong ref (the field/bean) so callbacks keep firing.
        Core.getGlobalContext().register(this);
    }

    @Override
    public void beforeCheckpoint(Context<? extends Resource> ctx) {
        // nothing to release here
    }

    @Override
    public void afterRestore(Context<? extends Resource> ctx) {
        // Force fresh entropy on the restored (possibly cloned) instance.
        java.security.SecureRandom.getInstanceStrong().nextLong();
    }
}

go deeper

for a junior

Know that implementing SmartLifecycle is enough to participate in checkpoint/restore.

for a middle

Map start()/stop() to afterRestore/beforeCheckpoint and know org.crac.Resource exists as the lower-level API.

for a senior

Choose between SmartLifecycle and org.crac.Resource by layer/semantics, and know the strong-reference and idempotency gotchas.

for a principal

Design a checkpoint-safe resource strategy across framework and non-framework components, including cross-cutting hooks like RNG/DNS reset.

**Two integration surfaces exist**, and choosing between them is about *what layer* owns the resource. **1. SmartLifecycle (preferred inside Spring).** Because Spring registers the `ApplicationContext` as an `org.crac.Resource` (through `DefaultLifecycleProcessor`), any bean implementing `Lifecycle`/`SmartLifecycle` is automatically driven around a checkpoint: `stop()` at `beforeCheckpoint`, `start()` at `afterRestore`. Contract details: - `start()` — acquire/open the resource. Runs on every restore, so it must be idempotent-safe and re-runnable. - `stop()` — release/close. Runs before the snapshot. - `isRunning()` — must accurately report state; Spring uses it to decide whether to call stop()/start(). A wrong value causes skipped teardown. - `getPhase()` — ordering. Higher phase starts later, stops earlier. Use it to sequence dependent resources. - `isAutoStartup()` — default true; keep it so the bean starts with the context and on restore. The big win: **zero CRaC-specific dependency** in your code, and the same stop()/start() also serves normal application shutdown, so it's easy to unit-test. **2. `org.crac.Resource` (direct).** The facade interface: ```java public interface Resource { void beforeCheckpoint(Context<? extends Resource> ctx) throws Exception; void afterRestore(Context<? extends Resource> ctx) throws Exception; } ``` Register with `org.crac.Core.getGlobalContext().register(this)`. CRaC notifies resources in LIFO for `beforeCheckpoint` and reverse for `afterRestore`. Use this when: - The component is **not a Spring lifecycle bean** — e.g. a static utility, a library object, or infrastructure created outside the context. - You need behavior that is **conceptually not 'stop the app'** — reseeding `SecureRandom`, refreshing cached DNS/host info, invalidating time-based tokens, resetting a random-seeded ID generator. These aren't really "start/stop," so modeling them as Lifecycle is awkward. - You need hooks that run **regardless of Spring's phase graph**. **Trade-offs.** SmartLifecycle keeps you framework-idiomatic, testable, and couples checkpoint behavior to shutdown behavior (usually desirable). `org.crac.Resource` couples you to the `org.crac` API and runs outside Spring's ordering, but is the right tool for cross-cutting, non-lifecycle concerns and for non-Spring code. **Gotchas.** - Hold a **strong reference** to any `org.crac.Resource` you register; the global context may keep weak references, and a GC'd resource silently stops receiving callbacks. - `beforeCheckpoint`/`afterRestore` can throw checked exceptions; an exception in `beforeCheckpoint` aborts the checkpoint. - Don't block indefinitely in these callbacks — checkpoint is a stop-the-world-ish operation. - Anything acquired in a constructor/`@PostConstruct` won't be reacquired on restore; move it into `start()`/`afterRestore`. **Rule of thumb:** model resource lifecycles as `SmartLifecycle`; use `org.crac.Resource` only for cross-cutting checkpoint concerns or non-Spring objects.

  • Why might registering an org.crac.Resource and forgetting to keep a reference to it be a bug?
    The global CRaC context may hold the registration weakly. If the only reference is dropped, the object can be garbage-collected and silently stops receiving beforeCheckpoint/afterRestore callbacks, so your teardown/reacquire never runs.
  • Give a case where org.crac.Resource fits better than SmartLifecycle.
    Reseeding SecureRandom or refreshing cached DNS after restore: these aren't 'start/stop the app' semantics, they only need to run on restore, and modeling them as Lifecycle phases is awkward. A direct Resource with a no-op beforeCheckpoint and real afterRestore is cleaner.

saying these in an interview costs you the question

  • Using org.crac.Resource for everything when SmartLifecycle would be idiomatic and testable
  • Registering a Resource as a throwaway local and losing the reference
  • Doing resource acquisition in the constructor so it never re-runs on restore

context