skip to content

questions

5

In an event-sourced system, an aggregate's current state isn't stored directly - it's derived by replaying its event stream. In plain terms, how does the system reconstruct that state, and what's this process called?

level: juniorimportance: must knowfreq 78%

answer

  1. fold/reduce over events
  2. apply(state, event) -> state
  3. log is truth, state is derived
  4. initial state + ordered events
  5. non-deterministic apply = bug

basics

~20 s

The system starts with an empty/default state and applies each stored event for that aggregate, one by one, in the order they happened, updating the state a bit with each event. Doing this to rebuild state is called rehydration; the loop that folds each event into state is the fold/reduce pattern.

solid answer

~40 s

Rehydration reconstructs an aggregate's current state without ever storing that state directly. You start from an initial/empty state object, load all events recorded for that aggregate's stream in the order they were appended, and pass each one through a pure 'apply' function that takes (currentState, event) and returns nextState - the classic fold/reduce pattern from functional programming. After the last event is applied, you have the aggregate's up-to-date in-memory state, ready to validate and execute the next command against. It's conceptually identical to git reconstructing a file by replaying commits, or a bank statement balance being the sum of all transactions. The key properties are: the apply function must be deterministic and side-effect-free, and events are immutable facts, so replaying them always yields the same result.

go deeper

for a junior

Should describe the basic loop: start from empty state, apply each event in order, end up with current state. Doesn't need to know about snapshots or performance implications yet.

for a middle

Should additionally explain why the apply function must be pure/deterministic and be able to write or read a simple apply(state, event) function for a toy aggregate.

for a senior

Should connect rehydration to command handling (load-decide-apply-persist cycle) and know that unbounded stream length is a real operational concern, even without implementing the mitigation.

for a principal

Should be able to reason about system-wide implications: how non-determinism in apply propagates into corrupted projections across a whole system, and how to design apply functions and event schemas defensively to keep replay safe.

## What rehydration is In an event-sourced system, you never store an aggregate's current state as a row you overwrite — state is treated as a **derived, calculated value**. What is stored, permanently and immutably, is the sequence of events that happened to that aggregate: an ordered, append-only log for a specific aggregate id. **Rehydration** is the process of turning that log back into an in-memory object you can query or run business logic against. Concretely: 1. Allocate an initial state (often the class's default constructor, or a sentinel "not yet created" state). 2. Load every event ever recorded for that aggregate id in the order they were appended. 3. Pass each one through a state-transition function — conventionally called `apply`, `when`, or `evolve` — that takes the current state and one event and returns the next state. Running that function across the whole list, carrying the accumulator forward, is exactly the **fold** (also called **reduce**) pattern most engineers already know from `Array.reduce` or functional programming. For a bank account aggregate, the events might be `AccountOpened(id, owner)`, `MoneyDeposited(amount)`, `MoneyWithdrawn(amount)`; applying them in order against an initial "no account" state yields, after three folds, an object with balance = deposit - withdrawal and owner set — the same object you'd get if you'd kept a mutable balance field all along, just computed on demand. ## Why the log is the source of truth This exists because the event log is treated as the single source of truth, and everything else — current state, read models, caches — is a projection of it. Storing only current state (the classic CRUD row-per-aggregate approach) destroys history: once you overwrite a balance, you can no longer answer "what was the balance last Tuesday" or "why did it change." By keeping the raw facts and deriving state through replay: - you get that **history for free** — you can always re-derive any past state by folding a prefix of the stream; - you can **build multiple different projections from the same log** — a balance view, a fraud-detection view, an audit view; - you get a **natural audit trail**, because nothing is ever destructively updated. ## The trade-off The trade-off is computational: replay costs time and CPU proportional to the number of events in the stream, every time you need the state. - A brand-new aggregate with three events rehydrates instantly. - An aggregate with 50,000 events (a long-lived IoT device, a years-old customer account) can take a noticeable amount of time to fold on every command. Production systems mitigate this with periodic **snapshots** of computed state so a rehydration only has to replay events since the last snapshot, but that's an optimization layered on top of the core replay mechanism, not a replacement for it — the log remains authoritative and the snapshot is just a cache. The other cost is conceptual: developers used to CRUD have to relearn "where is the current value" — it isn't in a column, it's a fold result, and that fold has to be run somewhere (in-process on read, in a background projector, or both). ## Failure modes The main failure mode at this level is writing an `apply` function that isn't a pure function of (state, event). If it reads the wall clock, calls an external service, generates a random id, or depends on anything other than its two inputs, then replaying the exact same event stream on two different days — or on two different read replicas — produces two different states, silently. - This is the **single most common bug** in early event-sourcing implementations: someone puts "set lastSeen = DateTime.Now()" inside an apply method instead of inside the command handler that produces the event, and now replay is non-deterministic. - A **second common failure** is applying events out of order (e.g., due to a bug in stream loading or a bad sort), which for most real state machines produces wrong or even invalid state (a withdrawal applied before its matching deposit could go negative when it shouldn't be allowed to). ## Where it shows up A concrete real-world example: frameworks like **Axon Framework** (Java) and **EventStoreDB-based** services expose exactly this pattern — an aggregate root class with an `on(Event)` or `apply(Event)` method for each event type, and a framework-provided repository that, on `load(id)`, reads the stream and folds it through those methods before handing you a ready-to-use aggregate to invoke a command against. The pattern is also visible outside formal event-sourcing frameworks: - **Git** reconstructs a file's contents by replaying commits. - A **bank statement's** running balance is the fold of all listed transactions. Both are the same rehydration idea in a different guise.

  • What happens if the apply function used during rehydration calls an external service or reads the system clock?
    Replay becomes non-deterministic: folding the same event stream at different times or on different nodes can produce different state, which breaks the core guarantee that state is a pure function of its events. Any time-dependent or environment-dependent data must be captured as a field on the event itself (e.g., an occurredAt timestamp) at write time, not recomputed during apply.
  • Does rehydration need to load the entire event stream every single time a command comes in?
    In the naive version, yes, which is fine for short streams but becomes a real cost as streams grow into the tens of thousands of events. Production systems typically add periodic snapshots so only events since the last snapshot need replaying, but that's an optimization on top of replay, not a substitute for it.
  • Where should validation logic like 'you can't withdraw more than the balance' live - in the apply function or the command handler?
    It belongs in the command handler (or a decide function), which checks invariants against the already-rehydrated state and decides whether to emit an event at all. The apply function should be a dumb, unconditional state-transition step - it just folds an event that has already been decided as valid, and must never reject or branch on business rules.

Like a bank statement's running balance: the statement doesn't store 'the balance' anywhere permanent, it's just the running sum of every listed transaction from zero - replay the transactions and you get the balance.

saying these in an interview costs you the question

  • Says the current state is stored in a database row and 'events' are just a log kept alongside it
  • Puts business validation inside the apply/fold function
  • Doesn't realize apply must be a pure function of (state, event)
  • Thinks replay only matters for debugging, not for normal command handling
  • Can't explain what 'initial state' means before any events are applied

context

open as a page

During command handling in an event-sourced aggregate, rehydration produces the current state, and then the incoming command has to be decided against it. Concretely, what are the two responsibilities involved - deciding vs. applying - and why should they not be merged into one function?

level: middleimportance: must knowfreq 65%

basics

~20 s

There are two separate jobs: first, figure out if the command is allowed and what should happen (the 'decide' step, checking business rules against the current state you just rebuilt); second, once you know what happened, update the state to reflect the new event (the 'apply' step). Keeping them separate means the apply step never has to think about whether something is allowed - it just records it.

open as a page

Beyond rehydrating an aggregate to handle its next live command, replay is also used as a debugging and audit tool - for example, reconstructing what an order's state looked like right before a bad event caused a bug. How does 'replay for debugging' differ mechanically from normal rehydration, and what do you need to build to support it safely?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Normal rehydration always folds every event to reach 'now.' Debugging replay instead stops folding partway through - at a specific event, timestamp, or version - so you can see exactly what the state looked like at that moment, without touching the live system or writing anything back.

open as a page

An aggregate's apply function was written when its ProductAdded event had fields (sku, quantity). Two years later the event is redesigned to (sku, quantity, warehouseId), and old events already in the stream only have the two original fields. What has to happen for replay to still correctly rehydrate aggregates whose streams contain the old-shaped events, and why can't you just edit the old events in place?

level: seniorimportance: should knowfreq 50%

basics

~20 s

You can't rewrite old events - they're an immutable historical record - so instead you add a translation step, often called upcasting, that transforms an old-shaped event into the new shape (e.g., filling in a default warehouseId) right before it's folded, so the apply function only ever has to deal with one, current shape.

open as a page

A platform team decides to fix a bug in a widely-used read-model projection by triggering a full replay of every aggregate's event stream to rebuild it from scratch across the whole system. What can go wrong at scale, and how would you design the rebuild to avoid taking down the system or producing a projection that's subtly different from what live processing would have produced?

level: principalimportance: should knowfreq 40%

basics

~20 s

Replaying every stream at once can overload the database and downstream services (a 'replay storm'), and if the fold function isn't perfectly deterministic or the rebuild runs against a moving live system, the rebuilt projection can end up subtly different from what you'd get in production. You avoid this by rebuilding into a separate copy, throttling the replay, and only swapping it in once it's verified to match.

open as a page