In an event-sourced system, why do teams take periodic snapshots of an aggregate's state, and what exactly gets stored in a snapshot?
answer
- snapshot = point-in-time state cache
- version/sequence number tag
- replay only the tail
- disposable, derivable from events
- schema version mismatch -> fallback to full replay
basics
~20 sEvent sourcing rebuilds an object's state by replaying every event that ever happened to it. If there are thousands of events, that replay gets slow. A snapshot is a saved copy of the state at some point, so you only replay events since then.
solid answer
~40 sA snapshot is a serialized copy of an aggregate's state at a specific event version/sequence number, stored so that rebuilding the aggregate doesn't require replaying its full event history from the beginning. Instead of loading and applying event #1 through #50,000, you load the snapshot taken at version 49,900 and replay only the last 100 events on top of it. It stores the aggregate's field values, the version/sequence number it corresponds to, and usually a schema version tag. It's purely a performance optimization — the event stream remains the source of truth, and the snapshot could be deleted and regenerated at any time without losing information.
go deeper
Should know a snapshot exists to avoid replaying everything and that it stores state at a point in time; doesn't need to know versioning edge cases.
Should be able to describe the load-snapshot-then-replay-tail algorithm precisely, including the version pointer.
Should discuss schema evolution of snapshots and safe-fallback-to-full-replay as a design principle.
Should reason about snapshot writing as a derived, idempotent side-process and its interaction with concurrent writes/versioning at scale.
## Why replay alone stops scaling Event sourcing stores every state change to an aggregate (an entity like an `Order` or a `BankAccount`) as an **immutable event**, and derives current state by replaying those events in order through the aggregate's apply/reduce logic. That gives a full audit trail and lets you rebuild state as of any point in time, but it has an obvious cost: the more events an aggregate accumulates, the longer it takes to load. - A **long-lived aggregate** — a bank account open for ten years, an IoT device streaming readings for months — can accumulate tens of thousands of events. - Replaying all of them on every load, especially every command handled, turns an **O(1)-feeling read** into an **O(n) operation** that grows without bound. ## What a snapshot is, and how a load uses it A snapshot breaks that growth. It is a **serialized capture** of the aggregate's state after applying events up to a known point, tagged with the version (or sequence number) it represents — e.g., 'Account-123 state after event #48,700'. To rebuild the aggregate, the loading code: 1. fetches the most recent snapshot at or before the version it needs; 2. deserializes it into the aggregate's in-memory shape; 3. then queries the event store for only the events with a version greater than the snapshot's version; 4. and replays that short tail on top of the snapshot. If the snapshot was taken at version 48,700 and the aggregate is now at version 48,900, you replay 200 events instead of 48,900. The event store is untouched and remains the **single source of truth**; the snapshot is a derived, disposable artifact — you could delete every snapshot in the system and reconstruct correct state purely from events, just slower. ## The trade-off The trade-off is what you'd expect from any cache-like optimization: you're trading storage and write-path complexity for read-path speed. - Every snapshot write is extra I/O, extra storage, and a second thing that can go stale or become inconsistent with the events if the two aren't written atomically or read together carefully. - You've also introduced a **versioning problem**: if the aggregate's internal shape changes (a field renamed, a new field added with a default, an enum value split into two), old snapshots serialized under the previous shape may not deserialize correctly under the new code. Systems handle this by tagging each snapshot with a schema/aggregate version number and either migrating old snapshots on read, or — the simpler and more common approach — treating an unreadable or stale-versioned snapshot as absent and falling back to full replay from event #1, then writing a fresh snapshot in the new format. Because snapshots are disposable, the safe default on any doubt is 'ignore it and replay everything'; correctness never depends on the snapshot being valid. ## Failure modes Failure modes tend to cluster around two axes: **staleness** and **corruption**. - **Staleness** happens when a snapshot lags far behind the event stream — for instance, a snapshot job that runs nightly on a system that also gets occasional huge event bursts; the day of the burst, loads temporarily replay thousands of events until the next snapshot catches up. That's a performance regression, not a correctness bug, and it self-heals. - **Corruption or inconsistency** is more dangerous: if a snapshot is written concurrently with new events being appended to the same aggregate without proper coordination, you can end up with a snapshot tagged at version 100 that actually reflects only 98 events' worth of state (a race between the snapshot job reading current state and a concurrent command appending event 99 and 100), or a snapshot whose version pointer doesn't match what's actually captured. The standard mitigation is to make snapshot writes derive strictly from a read of the event store at a specific, already-durable version — never from 'whatever the in-memory aggregate looks like right now' — so the snapshot's version tag is always trustworthy, and to treat snapshot writing as an idempotent, replayable side operation that can safely be retried or skipped without affecting correctness. ## Where you meet it in production A concrete example: Event Store DB (the open-source event-sourcing database) supports snapshots as an explicit pattern where client code stores a snapshot event back into the same stream (or a companion stream) periodically, and readers are told to start reading from the latest snapshot position rather than stream position zero. Similarly, **Akka Persistence** (JVM actor framework) has built-in snapshot support: an actor's persistent state is periodically saved via `saveSnapshot`, and on actor restart it loads the latest snapshot then replays only the journal entries recorded after that snapshot's sequence number — exactly the load-snapshot-then-replay-tail pattern described above, and it is the textbook implementation most engineers will meet in production.
- Is a snapshot ever the source of truth, or can you always regenerate it?It's never the source of truth — it's a derived cache. The event stream is authoritative, and any snapshot can be deleted and regenerated by replaying events from the beginning. This is why snapshot bugs are usually 'ignore and fall back to full replay' rather than data-loss incidents.
- What happens if the snapshot store is unavailable when loading an aggregate?The system should degrade gracefully to full event replay from the beginning of the stream. It's slower, but correctness is preserved because the event store remains available and authoritative.
- Why store the version number with the snapshot instead of just a timestamp?Events are ordered and replayed by sequence/version, not wall-clock time, so the snapshot needs to align with a specific position in that ordered stream. A timestamp doesn't tell you unambiguously which events still need replaying, especially with clock skew or events processed out of arrival order.
Like a video game checkpoint: instead of replaying the whole level from the start every time you die, you resume from the last checkpoint and only redo what happened since.
saying these in an interview costs you the question
- describes snapshot as replacing the event log
- thinks deleting a snapshot loses data
- can't explain what 'tail replay' means
- assumes snapshots are always safe to trust without a version check