When designing where snapshots live relative to the event stream (same store vs. a separate store), what trade-offs come into play, and how should the design handle the aggregate's serialization format changing over time?
answer
- colocated store = consistency simplicity
- separate store = tuned read perf, extra ops surface
- consistency gap risk when uncoordinated
- explicit schema/version tag distinct from stream version
- mismatch -> discard + full replay, or upcast if replay too slow
basics
~20 sYou can keep snapshots in the same database as events, or in a different, faster store. Same-store is simpler and stays consistent; a separate store can be optimized for fast reads but risks getting out of sync. Either way, you tag snapshots with a format version so old ones can be safely ignored if the code changes.
solid answer
~50 sStoring snapshots in the same store as events (e.g., a dedicated snapshot stream per aggregate, or a snapshot table in the same database) keeps writes transactionally close to the event append and simplifies operational reasoning — one system to back up, scale, and reason about consistency for. Storing snapshots in a separate, purpose-built store (e.g., a key-value cache like Redis, or blob storage) can be optimized for fast point lookups by aggregate ID and scaled independently of event-append throughput, at the cost of an extra system to operate and a potential consistency gap between the two stores if writes aren't coordinated. For schema evolution, every snapshot should carry an explicit schema/aggregate-version tag; on load, code checks that tag against what the current deserializer expects, and on any mismatch, discards the snapshot and falls back to full event replay rather than attempting a risky in-place migration of old snapshot bytes.
go deeper
Not expected to design this; awareness that snapshots are stored 'somewhere' alongside events is enough.
Should recognize that snapshots can live in the same or a different store and name one basic trade-off.
Should articulate the consistency-vs-performance trade-off precisely and design the schema-version-tag fallback mechanism.
Should weigh upcasting vs. discard-and-replay for large aggregates, and reason about consistency-gap failure symptoms in a decoupled snapshot-store architecture.
## Why the location is a real design choice The event store itself is usually optimized for append-heavy, ordered writes to many independent streams and range-scans by stream ID and version — that's the access pattern events need. Snapshots have a different access pattern: - single-key **point lookups** ('give me the latest snapshot for aggregate X'); - infrequent writes relative to event appends; - payloads that can be considerably larger than an individual event (a snapshot serializes the aggregate's entire current state, which can be many events' worth of data). Because the access patterns diverge, teams have a real design choice about where snapshots physically live, and the choice trades off consistency simplicity against read/write performance and operational surface area. ## Colocated with the events Keeping snapshots in the same store as events — for example, as a special event type appended to the same stream ('SnapshotTaken' as just another event, which some event-store products support natively), or as a table/stream colocated in the same database — has the strong advantage of **consistency simplicity**. If the snapshot write and the event append can be covered by the same transactional or at-least-ordered guarantees the store already provides, you largely sidestep the class of bugs where a snapshot's claimed version drifts from what the event log actually contains, because there's one system's consistency model to reason about rather than two. Event Store DB, for instance, supports storing snapshots as regular events in a dedicated stream, which lets it reuse the same durability, ordering, and read APIs the rest of the system already relies on. The cost is that you're bound by the primary event store's performance characteristics for a workload it may not be optimized for — a store tuned for high-throughput sequential append and stream-scan may not be the fastest thing for random point-lookups of large blobs by aggregate ID, especially at high read concurrency. ## A separate, purpose-built store Keeping snapshots in a separate, purpose-built store — a key-value store like Redis or DynamoDB, or blob storage like S3, keyed by aggregate ID — lets you pick a storage engine actually optimized for fast point-reads of a moderately large payload, and lets you scale snapshot storage independently of event-append throughput (which matters if the two have very different load profiles — event writes might be constant and moderate while snapshot storage needs to serve occasional-but-latency-sensitive reads at high fan-out during, say, a deploy that restarts many aggregate-hosting processes at once). The cost is operational: now there are two systems to keep available, back up, and monitor, and — the real hazard — a **consistency gap** between them. - If the snapshot store and event store are written by different processes without coordination (e.g., an async snapshot worker that reads from the event store and writes to Redis independently of the main write path), there's a window where a reader could see a snapshot that's newer than what a concurrently-in-flight event append will eventually settle to. - Or, if the snapshot writer crashes mid-write, a torn/partial snapshot. None of these are usually correctness-fatal for the aggregate itself (the event store is still authoritative and a bad snapshot triggers fallback to replay), but they can produce confusing, hard-to-reproduce staleness bugs if the fallback-on-mismatch logic isn't airtight. ## The two placements side by side | Placement | What it buys | What it costs | |---|---|---| | **Same store as the events** | consistency simplicity — one system's consistency model to reason about | bound by the primary event store's performance characteristics | | **Separate, purpose-built store** | a storage engine optimized for fast point-reads; snapshot storage you can scale independently of event-append throughput | operational — two systems to keep available, back up, and monitor, and a consistency gap between them | ## Schema evolution Schema evolution is the second major design axis, and it's largely orthogonal to the storage-location choice. An aggregate's serialized shape changes over the life of a system: fields get added, renamed, or removed; an enum value that used to have three cases now has four; a nested value object's internal structure changes. A snapshot written last month, before the change, will not deserialize cleanly into this month's expected shape. The disciplined approach is to tag every snapshot, at write time, with an explicit **schema or aggregate-version number** (distinct from the event-stream version number, which tracks event count, not shape) — a simple integer that the deserializer bumps whenever the aggregate's persisted shape changes incompatibly. On load, before attempting to deserialize the payload into the current type, the code checks this tag; if it doesn't match what the current code expects, the snapshot is treated as absent, and the loader falls back to full event replay from the beginning, then optionally writes a fresh, correctly-shaped snapshot to replace the stale one. The alternative — writing an in-place upcaster/migration function that transforms old-shape bytes into new-shape bytes on the fly — is sometimes worth it for aggregates so large that even full replay is prohibitively slow, but it's meaningfully more code and more surface area for subtle bugs than simply discarding and rebuilding, so most teams reserve it for aggregates where the full-replay fallback would genuinely be too slow to tolerate.
- Why is the schema/aggregate-version tag kept separate from the event-stream version number?The stream version tracks how many events have been applied — a purely quantitative count — while the schema version tracks the shape of the serialized payload, which changes due to code deployments, not event volume; conflating them would make it impossible to tell whether a mismatch is due to staleness (more events since snapshot) or incompatibility (different shape).
- When would writing an upcaster/migration function for old snapshot formats be worth the extra complexity?When the aggregate is large enough or long-lived enough that a full event replay fallback would itself be too slow to meet the read-latency budget — at that point, transforming old snapshot bytes in place avoids paying the full replay cost just because the schema moved on, even though it adds real migration-code complexity and testing burden.
- What's a concrete symptom of an uncoordinated separate snapshot store going stale relative to the event store?A reader loading the aggregate right after a burst of writes sees a snapshot that predates several just-appended events, so the tail-replay step correctly still catches it up — the symptom shows as slightly slower-than-expected loads or minor replay-tail-length spikes, not incorrect data, as long as fallback logic is sound.
Like choosing whether to keep your backup photos on the same hard drive as your working files (simple, always in sync) or on a separate cloud service optimized for fast browsing (faster to browse, but now you have to make sure the two don't drift apart).
saying these in an interview costs you the question
- assumes snapshot storage location has no performance/consistency implications
- conflates the event-stream version with a schema/format version
- proposes silently deserializing a mismatched schema version instead of falling back
- doesn't recognize the extra operational burden of a second storage system