skip to content

Event Sourcing

Storing the sequence of things that happened and deriving current state by replaying it, instead of storing the state itself. Interviewers probe it because it changes what fixing a bug means.

part ofEvent-driven architecture & messagingoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

In an architecture that combines Event Sourcing with CQRS (Command Query Responsibility Segregation), a client submits a command to change data. Walk through what happens on the write side before that change becomes visible to a query.

level: juniorimportance: must knowfreq 80%

answer

  1. load-replay-validate-append
  2. aggregate state is derived from replay, not stored
  3. command side never writes the read model directly
  4. optimistic concurrency via expected stream version
  5. publish then project asynchronously

basics

~10 s

The command side loads the past events for that thing, checks the business rules, creates one new event describing what happened, and appends it. Nothing is overwritten, only added.

solid answer

~30 s

A command handler loads the target aggregate by replaying (or restoring from a snapshot plus replaying) its historical events to get current state, validates the command against business rules, and if valid, appends exactly one new event to that aggregate's stream using an optimistic-concurrency check against the expected stream version. The event store is the append target, not the read database. Once appended, the event is published so downstream projectors can update read-model tables asynchronously. The command handler never writes to the query-side store directly, which is why the two sides can be scaled, modeled, and even stored independently.

go deeper

for a junior

Can describe the load-validate-append flow at a high level, ideally with an analogy like a ledger.

for a middle

Can name optimistic concurrency and stream versioning, and explain why snapshots help replay performance.

for a senior

Can discuss the transactional boundary (one aggregate equals one append), idempotent event publishing, and the trade-offs against plain CRUD.

for a principal

Can argue when Event Sourcing and CQRS are worth adopting at all versus simpler approaches, and the organizational cost of owning event schemas and projections long-term.

## What each half owns **Event Sourcing** stores state as an append-only sequence of immutable domain events rather than as mutable rows; the current state of an entity (an 'aggregate' in domain-driven design terms) is derived, not stored directly, by replaying its events in order. **CQRS** separates the model used to change state (commands) from the model used to read state (queries). Combined, the command side owns an aggregate and its event stream, while the query side owns one or more read-model projections built from those events. ## The flow, step by step Concretely, the flow is: - **(1)** a command arrives at a command handler naming a target aggregate by id; - **(2)** the handler loads that aggregate by fetching its event stream from the event store and replaying every event through the aggregate's apply logic to rebuild current in-memory state (for long streams this is sped up with a periodic snapshot, so replay only covers events since the last snapshot); - **(3)** the handler runs domain validation against that reconstructed state -- for example, refusing to ship an order that was already cancelled; - **(4)** if valid, it produces one new event (e.g. `OrderShipped`) and appends it to the stream, passing the version number it expected the stream to be at; - **(5)** the store either accepts the append, giving the stream a new version, or rejects it if another command changed the aggregate first (optimistic concurrency), forcing the handler to retry against fresh state; - **(6)** once appended, the event is durable and becomes the new source of truth; - **(7)** separately and asynchronously, one or more projectors subscribed to the event stream or a downstream bus consume the new event and update whatever read-model tables or search indexes serve queries. ## Why it exists This exists because a single mutable row model forces every reader and writer through one shape and one store, coupling scaling, schema, and technology choices that often have very different needs: writes need strong consistency and small, fast transactions per aggregate, while reads often need denormalized, search-optimized, or aggregated shapes serving far higher volume. Event Sourcing additionally gives a **full audit trail for free** (every state change is a permanent, inspectable fact) and lets you rebuild any number of new read models later by replaying history, since nothing was ever discarded on write. ## The trade-offs The trade-offs are real. On the plus side: - strong auditability; - the ability to add new projections after the fact by replaying history; - and a write path that is fast and uncontended because it only ever appends to one stream per command. On the minus side: - the read side is only eventually consistent with the write side (a projector takes some non-zero time to catch up); - the team must design and maintain event schemas that will live forever (since old events can never be silently rewritten); - and debugging requires reasoning about a system that has more moving parts than a single CRUD database. Complexity is the toll paid for those benefits, and teams that adopt Event Sourcing for entities with no auditing or historical-replay need often find the toll isn't worth it. ## Failure modes Failure modes show up mainly around concurrency and durability. - **Two commands race against the same aggregate.** The second one to attempt its append fails the optimistic-concurrency check and must retry -- if a handler doesn't implement retry, users see spurious command failures under load. - **The event store or the publish step fails** between appending and notifying projectors. The write is still durable (it's in the store) but projectors may lag until they catch up on their next poll or subscription resume; this is normal, not data loss, as long as projectors are built to resume from a durable checkpoint rather than relying on a live-only push. - **A snapshot mechanism is buggy.** Aggregate reconstruction can be slow (no snapshot, full replay every time) or, worse, wrong (a stale or corrupted snapshot silently producing incorrect state). ## Where it shows up A concrete, widely used real-world shape of this pattern is a banking ledger or an order-management system built with a framework such as **Axon Framework** or backed by a store such as **EventStoreDB**, where account balance or order status is never stored as a column to be updated in place; instead, every deposit, withdrawal, or status change is appended as its own event, and the balance or status shown to a user is a read-model column maintained by a projector that has consumed that account's or order's full event history.

  • What stops two concurrent commands against the same aggregate from producing conflicting events?
    The append call passes the version the handler expected the stream to be at. If another command already appended an event and moved the stream to a newer version, the store rejects the second append as a conflict. The handler then reloads the current state and either retries the command or surfaces a conflict to the caller.
  • Why not just update the read model synchronously inside the same operation as the event append?
    Doing so would couple the write path's latency and availability to every read model's store, defeating the point of separating them, and it often isn't even possible when the read model lives in a different storage technology such as a search index. Systems that truly need read-your-writes for a specific flow usually solve it by returning the updated data directly in the command's response, or reading straight from the just-updated aggregate, rather than making every projection synchronous.

Like a bank passbook: instead of erasing and rewriting your balance, each transaction is written as a new line in the book, and your current balance is just the sum of every line so far.

saying these in an interview costs you the question

  • describes commands as directly updating the read model or database row
  • thinks the aggregate's state is stored rather than derived by replaying events
  • no mention of any concurrency or version check on append
  • conflates the event store with a plain queue that has no durable, replayable log
  • assumes read-side updates happen synchronously with the write

context

open as a page

What is an event store, and how does its append-only log differ from a traditional database table that you update in place?

level: juniorimportance: must knowfreq 55%

basics

~20 s

An event store only lets you add new records, never change or delete old ones. Instead of overwriting a row to reflect current state, it keeps every change as its own permanent entry, in order.

open as a page

In an event-sourced system, why can't you just edit the shape of an already-stored event when your code's model changes, and what technique is normally used to let old and new event shapes coexist?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Old events are stored forever and get replayed to rebuild state. If you change what a field means without a plan, old events break the code that reads them. Upcasting is a small translator that turns an old-shaped event into the new shape before your code sees it, so replay still works.

open as a page

In an event-sourced system, what is a 'projection' and why do we build one instead of querying the raw event log directly for reads?

level: juniorimportance: must knowfreq 75%

basics

~20 s

A projection is a separate copy of your data, built by reading through the history of every change (events) and turning it into a simple, fast-to-search table - like building an index from a diary instead of re-reading the whole diary every time you need an answer.

open as a page

In an event-sourced system, an aggregate's current state isn't stored directly - it's derived by replaying its event stream. In plain terms, how does the system reconstruct that state, and what's this process called?

level: juniorimportance: must knowfreq 78%

basics

~20 s

The system starts with an empty/default state and applies each stored event for that aggregate, one by one, in the order they happened, updating the state a bit with each event. Doing this to rebuild state is called rehydration; the loop that folds each event into state is the fold/reduce pattern.

open as a page

In an event-sourced system, why do teams take periodic snapshots of an aggregate's state, and what exactly gets stored in a snapshot?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Event sourcing rebuilds an object's state by replaying every event that ever happened to it. If there are thousands of events, that replay gets slow. A snapshot is a saved copy of the state at some point, so you only replay events since then.

open as a page

In a system where the write side is event-sourced and a separate read side serves queries from projections built off those events, why is the read side typically eventually consistent, and what problem does that cause for a user who just submitted a change and immediately re-queries?

level: middleimportance: must knowfreq 85%

basics

~20 s

Events travel from the write side to the read side afterward, taking a little time. If you read right after writing, the read model may not have caught up yet, so you can briefly see stale data.

open as a page

When two commands try to append to the same event stream at nearly the same time, how does optimistic concurrency control in an event store prevent one write from silently clobbering the other?

level: middleimportance: must knowfreq 75%

basics

~20 s

Each write says 'I'm adding this assuming the stream currently has N events.' If someone else already added an event in between, the store rejects the write instead of accepting it blindly, so the app can retry with fresh data.

open as a page

Why do event-sourced systems typically partition the event log into one stream per aggregate ID rather than writing every event for every entity into a single global stream?

level: middleimportance: must knowfreq 65%

basics

~20 s

Grouping by aggregate id keeps each entity's own history together and in order, so you can load just that one entity's events fast, and two different entities' writes never block or interleave with each other.

open as a page

When evolving the schema of an event type that's already been written to an append-only event store, which kinds of changes are generally safe to make without breaking existing consumers or replay, and which are dangerous?

level: middleimportance: must knowfreq 65%

basics

~20 s

Adding a new optional field is usually safe because old readers can ignore it and new readers can default it. Renaming, removing, retyping, or changing the meaning of an existing field is dangerous because old events won't have the new shape and something will misread them.

open as a page

When a brand-new projection needs to be built for a system that already has years of events in its log, how does a 'catch-up subscription' take it from zero to fully up to date, and then keep it live afterward?

level: middleimportance: must knowfreq 65%

basics

~20 s

It reads through the entire history of past events first, from the very beginning, building up the data step by step, and once it reaches the present, it switches to just listening for new events as they happen - like binge-watching a show to get caught up, then watching new episodes live.

open as a page

During command handling in an event-sourced aggregate, rehydration produces the current state, and then the incoming command has to be decided against it. Concretely, what are the two responsibilities involved - deciding vs. applying - and why should they not be merged into one function?

level: middleimportance: must knowfreq 65%

basics

~20 s

There are two separate jobs: first, figure out if the command is allowed and what should happen (the 'decide' step, checking business rules against the current state you just rebuilt); second, once you know what happened, update the state to reflect the new event (the 'apply' step). Keeping them separate means the apply step never has to think about whether something is allowed - it just records it.

open as a page

Walk through, step by step, how an event-sourced aggregate is reconstructed when a snapshot exists: what does the loading code fetch, in what order, and how does it combine the snapshot with the event stream?

level: middleimportance: must knowfreq 70%

basics

~20 s

The code fetches the newest saved snapshot, turns it into an in-memory object, then asks the event store for only the events recorded after that snapshot's position, and applies those events one by one to bring the object fully up to date.

open as a page

What strategies are commonly used to decide when to take a new snapshot of an event-sourced aggregate — for example triggering on event count versus elapsed time — and what trade-offs does each have?

level: middleimportance: must knowfreq 65%

basics

~20 s

You can snapshot every N events (like every 500 changes), on a timer (like nightly), or when loading gets slow. Event-count triggers keep replay short and predictable; time-based triggers are simpler but can let busy aggregates drift far behind.

open as a page

A read model in an event-sourced CQRS system is built by a background projector that subscribes to the event stream and updates a query-optimized table. In production, list the concrete ways this projection pipeline fails, and how you'd detect and recover from each.

level: seniorimportance: must knowfreq 65%

basics

~20 s

The projector can crash, fall behind, receive duplicates, or choke on a bad event. Fix these with durable checkpoints, lag monitoring, idempotent updates, and the ability to replay events to rebuild the table from scratch.

open as a page

When a projection subscribes to read events from an event store to build a read model, what ordering guarantees can it actually rely on, and where do they break down?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Within one entity's own history, events always arrive in the exact order they happened. But if you're watching many entities at once, events from different ones can arrive interleaved in almost any order relative to each other, so you can't assume a global timeline unless the store gives you one explicitly.

open as a page

A user submits a form that appends an event to the log, then the client immediately navigates to a page that reads from a projection built off that log - and the change is missing. What is actually happening here, and what are the common ways to handle it?

level: seniorimportance: must knowfreq 70%

basics

~20 s

The projection hasn't caught up to the newest event yet, so the read is briefly stale - like refreshing a scoreboard a split second before it updates. Common fixes: make the client wait for the update, read the write directly instead of from the projection, or show a temporary version on screen immediately.

open as a page

Beyond rehydrating an aggregate to handle its next live command, replay is also used as a debugging and audit tool - for example, reconstructing what an order's state looked like right before a bad event caused a bug. How does 'replay for debugging' differ mechanically from normal rehydration, and what do you need to build to support it safely?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Normal rehydration always folds every event to reach 'now.' Debugging replay instead stops folding partway through - at a specific event, timestamp, or version - so you can see exactly what the state looked like at that moment, without touching the live system or writing anything back.

open as a page

In an event-sourced CQRS system, a single business workflow such as placing an order needs to update several independent aggregates -- say Order, Inventory, and Payment -- each with its own consistency boundary. What role does a process manager (saga) play here, and how does it drive the workflow forward?

level: middleimportance: should knowfreq 70%

basics

~20 s

A saga listens for events, and each time something happens it decides the next command to send to the next aggregate. It's like a conductor reacting step by step, instead of one big all-or-nothing transaction across everything.

open as a page

A team discovers that a projection's event-handling logic has been silently miscalculating a customer's lifetime-spend column for six months. What are the mechanics of fixing this by rebuilding the projection from the event log, and what must the event handlers guarantee for that rebuild to produce correct results?

level: middleimportance: should knowfreq 50%

basics

~20 s

You fix the buggy code, wipe out the broken table, and replay every past event from the start through the corrected logic to rebuild it fresh - like erasing a wrong total and re-adding every receipt correctly this time.

open as a page

A client retries an append call to an event store after a network timeout, but the original append had actually already succeeded server-side before the timeout. How do idempotency keys prevent this from creating a duplicate event?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Each write carries a unique ID. If the same ID shows up twice because of a retry, the store recognizes it already processed that exact write and just confirms success again instead of adding a second copy.

open as a page

What is the 'copy-and-replace' approach to migrating an event stream, how does it differ from upcasting-on-read, and when would you choose it despite it being more invasive?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Upcasting translates old events every time they're read, keeping the original bytes forever. Copy-and-replace instead writes a whole new, corrected copy of the event stream once, then switches everyone to read from that new copy — more work up front, but no more translation cost on every future read.

open as a page

What's the difference between a 'weak schema' approach and a 'strong schema registry' approach to versioning events within an event-sourced store, and what does each cost you?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A weak schema just serializes events as loosely-typed JSON with a version number and trusts application code to handle whatever shape shows up. A strong schema registry enforces a formal, checked schema for every event version up front, catching incompatible changes before they can even be written.

open as a page

A projection's catch-up subscription redelivers a batch of already-processed events after the consumer crashes and restarts. What must the projection's event handlers do to avoid corrupting the read model on redelivery, and how does at-least-once delivery change how projection code should be written?

level: seniorimportance: should knowfreq 50%

basics

~20 s

The handler must be safe to run twice on the same event without messing up the data - for example, setting a value directly instead of always adding to it - because most event systems can redeliver an event after a crash rather than guaranteeing it's processed exactly once.

open as a page

An aggregate's apply function was written when its ProductAdded event had fields (sku, quantity). Two years later the event is redesigned to (sku, quantity, warehouseId), and old events already in the stream only have the two original fields. What has to happen for replay to still correctly rehydrate aggregates whose streams contain the old-shaped events, and why can't you just edit the old events in place?

level: seniorimportance: should knowfreq 50%

basics

~20 s

You can't rewrite old events - they're an immutable historical record - so instead you add a translation step, often called upcasting, that transforms an old-shaped event into the new shape (e.g., filling in a default warehouseId) right before it's folded, so the apply function only ever has to deal with one, current shape.

open as a page

When designing where snapshots live relative to the event stream (same store vs. a separate store), what trade-offs come into play, and how should the design handle the aggregate's serialization format changing over time?

level: seniorimportance: should knowfreq 55%

basics

~20 s

You can keep snapshots in the same database as events, or in a different, faster store. Same-store is simpler and stays consistent; a separate store can be optimized for fast reads but risks getting out of sync. Either way, you tag snapshots with a format version so old ones can be safely ignored if the code changes.

open as a page

You're designing the aggregates for an event-sourced order-management system with CQRS. How do you decide where to draw each aggregate's consistency boundary, and what's the trade-off between making an aggregate bigger, to get more invariants enforced for free inside one transaction, versus splitting it and coordinating across the split via a process manager?

level: principalimportance: should knowfreq 45%

basics

~20 s

An aggregate should be just big enough to enforce the rules that must be true at every single instant. Anything that can tolerate a short delay to become consistent should live in a separate aggregate, coordinated afterward by a saga.

open as a page

A platform team decides to fix a bug in a widely-used read-model projection by triggering a full replay of every aggregate's event stream to rebuild it from scratch across the whole system. What can go wrong at scale, and how would you design the rebuild to avoid taking down the system or producing a projection that's subtly different from what live processing would have produced?

level: principalimportance: should knowfreq 40%

basics

~20 s

Replaying every stream at once can overload the database and downstream services (a 'replay storm'), and if the fold function isn't perfectly deterministic or the rebuild runs against a moving live system, the rebuilt projection can end up subtly different from what you'd get in production. You avoid this by rebuilding into a separate copy, throttling the replay, and only swapping it in once it's verified to match.

open as a page

In a high-throughput event-sourced system where multiple processes might concurrently write events and take snapshots for the same aggregate, what specific race conditions and failure modes can corrupt or invalidate a snapshot, and how do you design around them?

level: principalimportance: should knowfreq 35%

basics

~20 s

If two things try to save or read a snapshot for the same object at the same time, you can end up with a snapshot that doesn't match reality. Fix this by always deriving snapshots from a confirmed, already-saved event position, and by checking that the snapshot's numbers actually line up before trusting it.

open as a page

An event type OrderPlaced has been in production for two years with a fixed set of fields. The business now needs to add a required 'currency' field to how orders are represented, but the event store's old OrderPlaced events don't have it. How do you evolve the event schema without breaking existing projections, and what does this have to do with rebuilding read models?

level: seniorimportance: nice to knowfreq 35%

basics

~20 s

You never rewrite old events; they stay exactly as they were recorded. Instead you translate old events into the new shape on the fly when they're read, or version the event type, and any projector reading them handles both old and new shapes.

open as a page

showing 1–30 of 33