skip to content

A team discovers that a projection's event-handling logic has been silently miscalculating a customer's lifetime-spend column for six months. What are the mechanics of fixing this by rebuilding the projection from the event log, and what must the event handlers guarantee for that rebuild to produce correct results?

level: middleimportance: should knowfreq 50%

answer

  1. log is truth, projection is disposable
  2. reset checkpoint plus wipe store plus replay
  3. determinism required
  4. suppress side effects during replay
  5. blue-green cutover vs rebuild-in-place

basics

~20 s

You fix the buggy code, wipe out the broken table, and replay every past event from the start through the corrected logic to rebuild it fresh - like erasing a wrong total and re-adding every receipt correctly this time.

solid answer

~40 s

Because the event log is the source of truth and the projection is just a derived, disposable view, fixing a projection bug means: fix the handler code, then reset the projection - drop or clear its storage and reset its checkpoint to zero - and start a fresh catch-up subscription that replays the entire event history through the corrected handlers. This only produces correct results if the handlers are deterministic (same events in same order always yield the same output) and free of side effects that shouldn't repeat, since replay reruns every historical decision, including ones that originally triggered downstream effects like emails or webhooks, which must not fire again during a rebuild.

go deeper

for a junior

Should grasp the basic idea: fix the code, clear the table, replay events to rebuild it.

for a middle

Should describe the concrete steps - checkpoint reset, fresh catch-up subscription, cutover - accurately.

for a senior

Should identify determinism and side-effect suppression as hard requirements for safe rebuilds, and discuss blue-green cutover versus rebuild-in-place trade-offs.

for a principal

Should discuss this as an organizational practice: validating a rebuild before cutover (shadow comparison, sampling), versioning projections, and designing handler architecture upfront so future rebuilds are safe by construction.

## Why a rebuild is possible at all The single property that makes projection rebuilds possible at all is that **the event log, not the projection, is the source of truth**: a projection is explicitly a derived, disposable artifact, so if its logic was wrong, the fix isn't to patch the corrupted data in place - it's to discard the projection entirely and regenerate it from the immutable facts that never changed. ## The mechanics, step by step Mechanically, a rebuild has a few concrete steps. 1. **First, the handler bug is fixed and deployed** - in the example, whatever logic accumulates lifetime spend (perhaps it double-counted refunded orders, or missed a currency conversion) is corrected in code. 2. **Second, the projection's storage is reset**: the target table (or index, or cache) is truncated or dropped and recreated, and the projector's checkpoint is reset to position zero. 3. **Third, a fresh catch-up subscription is started** against the full event log, replaying every event from the beginning through the now-corrected handler, exactly the same mechanism used to bootstrap a brand-new projection. Once replay reaches the tail, the projection is back to serving live traffic with corrected data, and typically only then is the corrected projection cut over to serve production reads, often via a **blue-green swap** so the buggy version keeps serving until the corrected one is verified. ## Why not a migration script This mechanism exists because it is dramatically safer and simpler than trying to patch six months of accumulated errors with a targeted data-migration script. A migration script has to reverse-engineer exactly which rows are wrong and by how much, is itself new code that can have its own bugs, and offers no guarantee it matches what the corrected logic would have produced from scratch. Full replay sidesteps all of that: it doesn't matter how the projection got into a bad state, because the bad state is discarded and rebuilt from ground truth. This is also why immutability of the event log matters so much in event sourcing - it's precisely what makes 'just replay it' a reliable recovery strategy rather than a leap of faith. ## What the handlers must guarantee For a rebuild to actually produce correct results, the handlers must satisfy two properties. 1. **First, determinism**: given the same sequence of events in the same order, the handler must always produce the same resulting state. If a handler's logic depends on anything outside the event data itself - the current wall-clock time, a call to an external pricing service whose answer may have changed, a random ID generator - then a rebuild months later can produce different results than the original run did, and the 'fix' may itself be subtly wrong or non-reproducible. 2. **Second, and often overlooked**, handlers must be free of externally visible side effects that shouldn't be repeated: if the original event processing also sent a 'welcome' email or fired a webhook to a partner system, a naive full replay through those same handlers would resend every email and webhook ever triggered, which is a serious operational hazard. In practice, teams handle this either by keeping notification and side-effect logic in entirely separate handlers from pure state-projection handlers (so a rebuild only re-runs the latter), or by running rebuilds with external effects explicitly suppressed. ## The trade-off The trade-off is mostly about time and availability versus correctness confidence. Rebuilding is conceptually simple and trustworthy, but for a large event log it can be slow - hours for millions of events - during which either the projection is unavailable (if rebuilt in place) or a duplicate resource cost is paid (if rebuilt alongside the old version for a blue-green cutover). It also assumes the fix is actually correct; a rebuild with a still-wrong handler just produces a different wrong answer, faster than a hand-written migration would have, but wrong all the same, so rebuilds are usually validated against a sample or a shadow-traffic comparison before cutover. ## Failure modes in practice Failure modes in practice include: - forgetting to suppress side-effecting handlers during rebuild, which floods customers with historical notifications - non-deterministic handlers producing a rebuild that doesn't match expectations and requires a second investigation to figure out why - rebuilding in place without a blue-green swap, which takes the read API offline or serves partial data mid-rebuild A well-known concrete pattern here is the **projection-versioning** approach used by frameworks like Axon Framework, where a new projection version is built alongside the old one under a new name or table, validated, and only then does a router or feature flag cut reads over to it - avoiding any window where reads see either downtime or a half-rebuilt table.

  • What goes wrong if a rebuild replays historical events through handlers that also send customer-facing emails?
    Every notification ever triggered across the entire event history gets resent during the rebuild, which is a serious operational and reputational hazard. Teams avoid this by separating pure state-projection logic from side-effecting logic into different handlers, or by running rebuilds with side effects explicitly suppressed.
  • How would you avoid downtime on a customer-facing read API while a projection with millions of events is being rebuilt?
    Use a blue-green approach: build the corrected projection into a new table or store alongside the existing one, let it fully catch up via a fresh subscription, validate it, and only then atomically switch reads over to the new version. The old, buggy projection keeps serving traffic throughout, so users never see downtime or an empty table mid-rebuild.
  • Why does non-determinism in a handler make rebuilds risky?
    A deterministic handler guarantees the same event history always yields the same resulting state, which is what makes 'just replay it' trustworthy. If a handler depends on something outside the event data - current time, an external service call, randomness - a rebuild months later can silently produce different results than the original run, undermining confidence that the fix actually worked.

Like discovering your accountant used the wrong tax formula on this year's ledger: you don't try to hand-patch every affected line item - you fix the formula and re-run it against every original receipt from the start of the year to produce a correct ledger from scratch.

saying these in an interview costs you the question

  • proposes writing a manual data-migration script instead of considering replay
  • doesn't mention resetting the checkpoint alongside clearing the store
  • misses that side-effecting handlers like emails or webhooks would refire during replay
  • assumes rebuild always requires downtime with no mention of blue-green cutover
  • doesn't connect handler determinism to the safety of replay

context