skip to content

You are designing checkpoint-and-rollback for a component whose work includes writing files, calling remote services, and holding open connections. Where does the approach of snapshotting and restoring an object's internal state stop working, and what do you do instead?

level: principalimportance: nice to knowfreq 20%

answer

  1. memento rewinds owned in-memory state only
  2. escaped effects need compensation, not restore
  3. saga: step + compensating step in reverse
  4. exclude handles/sockets; re-acquire on restore
  5. grace window: delay the effect, then undo is real

basics

~20 s

Snapshots only restore state the object owns in memory. Anything that escaped — files written, messages sent, remote calls made — is not rolled back by restoring a snapshot. Those need explicit compensating actions, and live resources such as connections should be rebuilt on restore rather than captured.

solid answer

~60 s

A memento restores the originator's internal state; it has no power over effects that have already crossed a boundary. Three failure classes: (1) externally visible side effects — files, sent messages, charged payments, rows other services already read — which require compensating actions (delete the file, publish a reversal, issue a refund) rather than restoration, and which are only approximately reversible; (2) live resources — sockets, file handles, GPU buffers, thread and transaction contexts — which must be excluded from the snapshot and re-acquired on restore, since a captured handle may be closed or stale by then; (3) shared and aliased state — other objects holding references into the originator, or observers already notified — where restoring one object leaves the graph inconsistent, so the snapshot boundary must be the whole consistency unit. Practically: snapshot only the pure, owned, in-memory state; defer irreversible effects until after the point of no return, or make them idempotent and reversible; use a saga or outbox pattern for cross-service work; and never present undo in the UI for actions the system cannot actually retract.

go deeper

for a junior

Say that restoring a saved state only undoes changes inside the object; things like sent emails or written files stay done.

for a middle

Add that live resources should be excluded and re-opened on restore, and that other objects holding references may need refreshing.

for a senior

Cover compensating actions, deferring effects until commit, idempotency keys, snapshot boundary as the consistency unit, and capturing under the mutation's lock.

for a principal

Turn it into policy: classify every action as undoable / compensatable / irreversible, drive UI affordances from that, use outbox and saga patterns for cross-service work, adopt grace windows to convert irreversible effects into rollbackable ones, and handle persisted-snapshot versioning, retention, and data-governance obligations.

## What a memento can and cannot reach A memento is a copy of the originator's *own in-memory state*. Restoring it rewinds exactly that and nothing else. Anything that already crossed a boundary — the process boundary, the disk, the network, another object's reference graph — is outside its reach. Interviewers use this to separate people who know the pattern from people who have shipped rollback. ## Failure class 1 — externally visible side effects If the operation wrote a file, sent an email, published an event, charged a card, or updated a row another service already read, restoring the snapshot leaves the world inconsistent: internal state says "never happened", the world says otherwise. The remedy is not restoration but **compensation** — an explicit action that approximately reverses the effect: | Effect | Compensation | Caveat | |---|---|---| | File written | Delete or restore prior version | Someone may have read it | | Event published | Publish a reversal/cancellation event | Consumers must handle it | | Payment charged | Refund | Fees, timing, ledger shows both entries | | Email sent | Send a correction | Cannot unsee it | This is the **saga** model: a long-running operation is a sequence of local steps, each with a compensating step, executed in reverse on failure. Compensation is *semantic* rather than exact — the system reaches a consistent state, not the identical prior state. Design moves that shrink the problem: - **Defer effects past the point of no return.** Accumulate intended effects, snapshot freely while everything is in memory, and flush the effects only when the operation commits. This is the outbox idea and the reason "pure core, effectful shell" architectures are easy to checkpoint. - **Make effects idempotent and keyed.** With an idempotency key, a retry after rollback does not double-charge. - **Two-phase style staging.** Write to a temp file and rename atomically on commit; hold messages in an outbox table committed with the state change and published after. ## Failure class 2 — live resources in the snapshot Sockets, file handles, database connections, GPU buffers, thread-local context, open transactions, timers, and subscriptions are *identities of live things*, not values. Capturing one is a bug in waiting: by restore time it may be closed, expired, or owned by someone else; restoring installs a stale handle that fails on first use, sometimes much later. The rule: **exclude non-value state from the memento and re-acquire it on restore.** Capture the descriptor (path, URL, offset, query parameters), not the handle. Have the originator expose an explicit `reattach()`/`rehydrate()` step, and document which fields are transient — this is also why serialization frameworks have transient/ignore markers. ## Failure class 3 — shared and aliased state If other objects hold references *into* the originator's structures, or the originator already notified observers, restoring it alone breaks the graph: caches and indexes elsewhere still describe the newer state, observers acted on an event that is now un-happened, and a parent may hold a child that the restored state no longer contains. Implications: - The snapshot boundary must be a **consistency unit** — the whole aggregate that must move together — not a convenient single object. - After restore, **invalidate or rebuild** dependents: clear derived caches, re-emit a coarse "state replaced" notification rather than fine-grained diffs, and re-establish parent/child links. - Prefer value semantics inside the snapshot boundary so no external aliases exist in the first place. ## Concurrency Capturing a snapshot of state that other threads are mutating yields a torn, never-actually-existed state. Capture under the same lock or consistency scope as the mutation, or make the state immutable so a reference *is* a consistent snapshot. Similarly, a restore that lands while another thread is mid-operation can clobber that work; restore usually needs to be a quiescent, exclusive operation. ## Persistence and lifetime Once mementos are written to disk or shipped elsewhere, they inherit every persistence concern: schema evolution (an old snapshot describes a state shape that no longer exists — version-stamp and migrate or reject), retention and cleanup, and data governance if the state holds personal data or secrets (encryption, deletion requests, audit). ## The honest product answer Decide, per action, which of three categories it falls into and make the UI reflect it: **undoable** (pure in-memory, snapshot works), **compensatable** (we can issue a reversal — offer "undo" with a delay window or an explicit "cancel" that emits compensation), and **irreversible** (require confirmation, never label it undo). The frequent real-world design is a **grace window**: hold the effect for a few seconds, expose undo during that window, then commit — which converts an irreversible effect into a snapshot-restorable one for as long as it has not escaped.

  • Why should an open database connection or file handle never be stored inside a memento?
    It is an identity of a live resource, not a value. By restore time it may be closed, timed out, or reassigned, so reinstating it installs a stale handle that fails unpredictably later. Capture the descriptor (path, URL, offset) and re-acquire on restore.
  • How does a saga differ from restoring a snapshot?
    A saga does not rewind state; it runs an explicit compensating action per completed step, in reverse order, to reach a semantically consistent state. It is approximate rather than exact — a refund leaves two ledger entries, not zero — and it is the only option once effects have been observed externally.
  • What is the 'grace window' technique and why is it popular for undo?
    Delay the irreversible effect (sending, deleting, publishing) for a few seconds while holding it locally. During that window nothing has escaped, so undo is a genuine in-memory rollback. After it expires the effect commits and undo is withdrawn from the UI.

You can rewind a chess engine's internal board position instantly, but if you already mailed your move to your opponent, rewinding your own board changes nothing about the game. You have to send a retraction and hope it arrives first.

saying these in an interview costs you the question

  • Believing restoring a snapshot rolls back files written or messages already sent.
  • Capturing open sockets, file handles, or transactions inside the snapshot.
  • Snapshotting one object while other objects hold live references into its internals, then wondering why the graph is inconsistent.
  • Taking a snapshot without the consistency scope/lock that guards the state, producing a torn state.
  • Offering an undo button for effects the system cannot actually retract.
  • Persisting snapshots with no version stamp, so an old one is restored into a changed state shape.

context