skip to content

Walk through, step by step, how an event-sourced aggregate is reconstructed when a snapshot exists: what does the loading code fetch, in what order, and how does it combine the snapshot with the event stream?

level: middleimportance: must knowfreq 70%

answer

  1. fetch snapshot, deserialize as starting state
  2. query events after snapshot version
  3. apply tail in order
  4. fallback: no snapshot -> replay all
  5. version mismatch = discard snapshot

basics

~20 s

The code fetches the newest saved snapshot, turns it into an in-memory object, then asks the event store for only the events recorded after that snapshot's position, and applies those events one by one to bring the object fully up to date.

solid answer

~40 s

Loading proceeds in three steps: (1) query the snapshot store for the latest snapshot for that aggregate ID, getting back its serialized state plus the version/sequence number it represents; (2) deserialize that into the aggregate's in-memory representation — this becomes the starting state instead of an empty/default aggregate; (3) query the event store for events on that aggregate's stream with version strictly greater than the snapshot's version, fetched in order, and apply each one through the aggregate's normal event-handling logic exactly as if it were being processed live. After the last event is applied, the aggregate is at the current version and ready to accept commands. If no snapshot exists (new aggregate or none taken yet), step 1/2 are skipped and step 3 replays from version zero.

go deeper

for a junior

Should describe the basic idea: load the saved state, then replay just the newer events on top.

for a middle

Should give the precise three-step algorithm including the strict version-boundary detail (greater-than, not greater-or-equal).

for a senior

Should identify the silent-corruption risk from a mistagged snapshot version and the schema-version-mismatch fallback pattern.

for a principal

Should design the defensive checks (gap detection, schema-version tagging) and reason about how a subtle apply-logic bug would make snapshot-based and full replay diverge.

## Reconstruction is a fold Reconstructing an event-sourced aggregate is, at its core, a **fold**: you start from some initial state and apply a sequence of events to it, each event transforming the state through the aggregate's own apply/when/on logic (whatever the framework calls the function that maps `(state, event) -> newState`). - **Without a snapshot**, that initial state is a blank aggregate — usually a default-constructed or 'empty' object of the aggregate's type — and the sequence is every event ever recorded for that aggregate's stream, fetched from version 1 (or 0, depending on numbering convention) onward. - **With a snapshot available**, the initial state changes: instead of starting from blank, you start from the state captured in the snapshot, and the event sequence shrinks to only the events recorded after the snapshot's version. ## The three-step algorithm Concretely, the algorithm has three steps. 1. **First**, the loading code asks the snapshot store — which may be a separate storage system, a special stream, or a table alongside the event store — for the most recent snapshot associated with the aggregate's identifier. That response includes the serialized state and, critically, the version or sequence number at which it was taken (e.g., 'this snapshot reflects the aggregate after event #4,820'). 2. **Second**, the code deserializes that payload into the aggregate's actual in-memory type, using whatever serialization format the system uses (JSON, Avro, protobuf, or a language-native format) — this deserialized object becomes the starting point of the fold instead of a blank aggregate. 3. **Third**, the code queries the event store for events belonging to that same aggregate's stream with a version strictly greater than the snapshot's version, fetched in ascending order, and applies each one in turn through the normal apply logic, exactly as if those events were arriving live. Once the last event in that tail has been applied, the aggregate object is now at the same version as the event store's head for that stream, and is ready to handle new commands or be inspected. If no snapshot exists yet — a brand-new aggregate, or a system where snapshotting was only just enabled — the first two steps are simply skipped, and the third step replays the entire stream from the beginning, which is the fallback path the whole mechanism exists to avoid needing routinely. ## Why the anchor is worth it The reason this exists is straightforward: the fold over the full event history is correct but its cost scales linearly with history length, and for long-lived, high-write aggregates that cost eventually violates whatever latency budget the read path has. Anchoring the fold's starting point at a recent snapshot converts an **O(total events)** operation into an **O(events-since-snapshot)** operation, which is the entire value proposition, and it costs nothing in correctness as long as the snapshot faithfully represents the state at its claimed version — which is the crux of most of the failure modes in this area. ## Hazard one — a snapshot that lies about its version The main correctness hazard in the restore path is a mismatch between what the snapshot claims (its version number) and what it actually contains. If a snapshot is tagged as version 4,820 but was actually captured after only 4,800 events were durably applied (a race between an async snapshot writer and concurrently-appended events, or a bug where the version counter is read before the state, or vice versa), then replaying events 4,821 onward on top of it silently produces wrong state — some events effectively get skipped or double counted relative to what they should be. This is worse than an unavailable snapshot, because an unavailable snapshot fails loudly (falls back to full replay) whereas a mistagged snapshot fails silently, producing plausible-looking but incorrect state. The standard defenses are: - derive the snapshot strictly from a durable read of the event store at a known version (never from 'whatever is currently in memory'); - and, defensively, verify on load that the event immediately after the snapshot's version actually exists and its version is exactly snapshot-version-plus-one — a gap there (e.g., due to stream truncation or a version-numbering bug) should trigger a fallback to full replay rather than blindly trusting the snapshot pointer. ## Hazard two — a shape change after a deploy A second common failure mode surfaces after a deploy: the aggregate's apply logic or field shape changes, and an old snapshot deserializes into a shape the new code doesn't expect (a missing field, a renamed field, a changed enum). Production systems typically tag each snapshot with a schema/aggregate-version number alongside the event-stream version, and on load, check that the snapshot's schema version matches what the current code expects; a mismatch causes the loader to discard the snapshot and fall back to full event replay, then optionally write a fresh, correctly-shaped snapshot afterward. **Akka Persistence's** snapshot recovery is a concrete, widely used implementation of exactly this restore algorithm: on actor restart, it loads the most recent snapshot (if any), then replays only the journal entries whose sequence number is greater than the snapshot's, applying each through the actor's `receiveRecover` handler before the actor becomes ready to process new commands.

  • What should happen if the event immediately after the snapshot's claimed version is missing from the event store?
    That's a strong signal the snapshot's version pointer is untrustworthy or the stream has a gap/truncation issue, so the safest response is to discard the snapshot and fall back to replaying the full stream from the beginning rather than silently continuing from a possibly-wrong starting point.
  • Does the tail-replay step apply events differently than they were applied when originally recorded?
    No — it must use the exact same apply/reduce function the aggregate uses for live events, so that replay produces identical results to the original sequential processing; any divergence (e.g., a bug fix that changes apply semantics) would make old snapshots and their tails produce different state than a full replay would, which is a serious correctness bug.
  • Why fetch events strictly greater than the snapshot version rather than greater-or-equal?
    The snapshot already reflects the state after that event was applied, so re-applying the same event again would double-apply it and corrupt the state; the query boundary has to be exclusive of the snapshot's own version.

Like resuming a saved document: you open the last autosave (snapshot) instead of a blank page, then reapply only the keystrokes (events) typed since that autosave to get back to exactly where you left off.

saying these in an interview costs you the question

  • thinks replay starts from event #1 even when a snapshot exists
  • doesn't mention checking/using the snapshot's version number
  • unaware that a version/schema mismatch should trigger fallback rather than blind trust
  • conflates snapshot restore with database row restore/backup

context