skip to content

A team running a CQRS system backed by Event Sourcing wants to add a brand-new read model over two years of historical order data, and separately needs to fix a bug in an existing projection's logic. What does deriving all read models from the event log make possible here that a 'read replica of the write database' approach would not, and what operational risks come with exercising that capability?

level: principalimportance: should knowfreq 50%

answer

  1. event log retains deltas, not just current state
  2. new projections can backfill over full history
  3. rebuild-from-scratch fixes historically-wrong rows
  4. blue-green projection cutover
  5. event schema evolution risk during long replay

basics

~30 s

Because every past change was saved as an event, not thrown away, the team can build a new report over years-old data by replaying all those old events through brand-new logic, as if they were just now happening — something a normal database copy can't do, since it only ever has today's current numbers. The risk is that replaying millions of old events takes real time and computing power, and can strain the system if done carelessly.

solid answer

~60 s

Because the event log preserves every historical state transition, not just current state, a team can write an entirely new projection today and run it once against two years of past events to backfill a read model that would otherwise only be able to start from 'now.' A read replica of a state-stored write database can only ever mirror current rows — it has no memory of the deltas that produced them, so a new read model built that way starts empty and can only accumulate history going forward. Fixing a buggy existing projection similarly benefits: the team can rebuild the read store from scratch by replaying the full stream through corrected logic, guaranteeing every row reflects the fix, not just rows touched after deployment. The operational risks are real: replaying years of events can take significant time and load on the event store, the new/rebuilt projection needs to be run alongside (not instead of) the live one until it's caught up and verified, and any point where events changed shape over time (schema evolution) has to be handled correctly by the replay logic or the backfill silently produces wrong results for older events.

go deeper

for a junior

Should grasp the basic idea: events keep history around, so new reports can be built over old data, unlike a plain copy of today's data.

for a middle

Should be able to explain both the new-read-model and the bug-fix rebuild use cases at a mechanical level.

for a senior

Should identify the concrete operational risks (load during replay, cutover safety) and propose a rebuild-alongside-live approach rather than an offline rebuild.

for a principal

Should design the end-to-end rebuild/cutover strategy (blue-green projection, reconciliation checks, rollback path) and account for event schema evolution as a first-class risk across a long replay window.

## What each write model retains The defining structural difference between an event-sourced write model and a plain state-stored one is what's actually retained over time. | Write model | What survives over time | |---|---| | A state-stored table, even with a read replica sitting behind it | Only ever holds current state — the total of Tuesday's order plus Wednesday's edit plus Thursday's cancellation collapses down to one final row, and the intermediate deltas that produced it are gone the moment they're overwritten | | An event log | Retains every one of those deltas, forever, as discrete, ordered facts | That single difference is what makes 'deriving read models from events' a genuinely different capability, not just a stylistic choice, from 'deriving read models from a database copy.' ## A brand-new read model over old data Concretely, when the team wants a brand-new read model — say, a report on time-from-order-placement-to-fulfillment broken out by product category, for the last two years, that nobody thought to build two years ago — an event-sourced system lets them write the new projection's handler logic today and then run it once as a **backfill job** that streams through every `OrderPlaced`, `ItemAdded`, and `OrderFulfilled` event from the last two years, in order, producing a fully populated, historically accurate read model on the first run. A read-replica-of-a-state-store approach has no equivalent: the replica only has today's rows, and a brand-new report built off it can only start accumulating history from the moment it's deployed forward — the previous two years are simply unrecoverable, because the granular events that would let you recompute them were never kept as such. ## The same mechanism fixes a buggy projection The same mechanism serves the bug-fix case just as directly. If an existing projection's handler had a logic error — say, it mishandled a particular event type and produced a slightly wrong running total — patching the handler code only fixes behavior for events that arrive after the fix ships; every row the old, buggy logic already wrote stays wrong forever unless something reprocesses the history. Because the event log is the durable, replayable source of truth, the standard remedy is to build a fresh copy of the read store from an empty state and replay the entire stream through the corrected handler, guaranteeing that every row — the ones from two years ago and the ones from this morning — was produced by the same, correct logic. This is qualitatively different from a typical database-migration bug fix, where you'd have to write a bespoke backfill script that tries to reverse-engineer what the correct historical values 'should have been,' often imperfectly, from whatever current state happens to still be sitting in the table. ## The operational risks of exercising it None of this is free, and the operational risks scale with exactly how much history exists. 1. **Load during replay.** Replaying two years of events for every order a business has ever processed can mean tens or hundreds of millions of events streaming through a new projection's handler, and that's real, sustained load on the event store — enough to require running the backfill as a rate-limited background job rather than a naive tight loop, and enough that teams need to actively monitor its progress and impact on the live system rather than assuming it'll quietly finish overnight. 2. **Correctness during the cutover.** A second risk: the new or rebuilt projection has to be built and verified alongside the existing live one — never by taking the live read model offline while the rebuild runs — and the team needs a clear, tested way to atomically swap traffic over once the rebuild has caught up to the live edge of the stream, or users can be served a half-built, still-catching-up read model mid-migration. 3. **Event schema evolution.** A third and often underestimated risk: over two years, the shape of an event or the meaning of a particular field may well have changed (a field renamed, an enum value's semantics shifted, an event split into two more specific ones), and the new projection's handler has to correctly interpret every historical version of every event type it consumes, not just the current one — getting this wrong doesn't throw an error, it silently produces a plausible-looking but incorrect backfilled read model, which is far more dangerous than a build failure because it can go unnoticed for a long time. ## The blue-green projection rebuild A well-known real-world pattern for managing exactly this is the blue-green projection rebuild: 1. stand up the new or corrected projection under a fresh name/table; 2. run it as a backfill plus live-tailing job until it has fully caught up to the current stream position; 3. run automated reconciliation checks comparing its output against expectations or against the old projection for overlapping data; 4. and only then flip application traffic to read from the new table — keeping the old one available briefly as a rollback path in case the reconciliation missed something.

  • Why can't a state-stored write model with a read replica just add change-data-capture (CDC) history logging retroactively to get this same capability?
    CDC can only capture changes from the moment it's turned on forward; it has no way to recover the deltas for changes that already happened and were overwritten before CDC existed, so the two-years-back backfill scenario is still unavailable for any history predating the CDC log's start. It closes the gap for future history, but the past is still unrecoverable.
  • Should a projection rebuild ever run by taking the existing live read model offline first?
    No — the standard approach builds the new or corrected version alongside the live one under a different name, verifies it's correct and caught up, and only then cuts traffic over, so users keep being served the working (if imperfect) old version throughout the rebuild instead of seeing an empty or partially-built read model.
  • What's the specific danger of event schema evolution during a two-year historical replay, compared to normal ongoing projection operation?
    Ongoing operation only ever has to interpret the current, latest shape of each event type, but a full historical replay has to correctly interpret every version of every event type that existed across the whole period, including old shapes that may no longer even be documented well; a handler written only against today's schema can silently misinterpret old events instead of failing loudly.

It's like the difference between only ever photographing today's finished jigsaw puzzle versus keeping every single piece ever placed, in the order it was placed. If someone later asks what the puzzle looked like halfway through, or wants to build a totally different picture using just the edge pieces, you can only answer that with the second approach — the photograph-only approach has already lost the information.

saying these in an interview costs you the question

  • Thinks a read replica of a state-stored database offers the same retroactive-read-model capability as an event log
  • Suggests taking the live read model offline while a rebuild runs
  • Doesn't mention event schema evolution as a risk during a long historical replay
  • Assumes replaying years of events is essentially free/instantaneous with no load or time cost

context