When a projection's logic has a bug and produced incorrect data, how do you safely rebuild it by replaying the event history, without taking the read side offline or serving inconsistent results mid-rebuild?
answer
- projections are derived, disposable, rebuildable
- blue-green: new store + catch-up + cutover
- replay cost scales with retained history
- retention window bounds how far back you can replay
- old store stays live until cutover
basics
~20 sYou build a brand new copy of the projection from scratch by replaying all the past events into a fresh table, and only switch reads over to it once it's fully caught up — so users keep reading the old (working) version the whole time instead of seeing a half-rebuilt one.
solid answer
~50 sThe safe pattern is blue-green rebuild: stand up a new, empty instance of the projection store, run a fresh catch-up subscription from the beginning of the event history against it with the fixed handler logic, and only cut reads over to it once it has fully caught up to the live head — the old projection keeps serving traffic the whole time. This avoids the two failure modes of an in-place rebuild: readers seeing a partially-rebuilt, inconsistent dataset, and the rebuild racing with live writes to the same table. You need the source event history to actually be retained long enough and completely enough to replay (which constrains your event retention policy), the rebuild to run on separate infrastructure so it doesn't starve production capacity, and a clean cutover mechanism — a router, alias, or feature flag — to atomically switch traffic once caught up.
go deeper
Should understand that a broken projection can be fixed by reprocessing events rather than hand-editing the read data.
Should describe the basic blue-green idea: build a new copy alongside the old one and switch over once it's ready.
Should identify the concrete risks — retention limits, infra contention during replay, atomic cutover across consumers — and design around them.
Should set organization-level policy: retention requirements for replayability, standard rebuild tooling/runbooks, and when to invest in projection snapshotting versus accepting full-genesis replay cost.
## Why a projection can be thrown away A projection is **derived state**: everything in it can, in principle, be reconstructed by reapplying the full history of source events, from the beginning, through the handler logic. This 'derived and disposable' property is one of the core ideas in event-driven design that distinguishes a projection from the source of truth it's built from — the projection itself doesn't need to be trusted or backed up the way the underlying event log or write model does, because if it's ever wrong, correctness can be recovered by replaying rather than by manually patching data. ## The blue-green rebuild, step by step The standard mechanism for doing this safely is often called a **blue-green rebuild**, borrowing the term from blue-green deployment. 1. You provision a brand-new, empty instance of the projection store — a new table, a new search index, a new collection — completely separate from the one currently serving reads. 2. You then run a fresh catch-up subscription starting from the earliest retained position in the event history, through the corrected handler logic, writing into that new store. 3. You track its progress until it has fully caught up to the live head of the event stream. 4. Only then do you atomically switch the read path — via an alias swap, a router change, a feature flag, or a DNS/service pointer update — so that queries start hitting the new, correct store instead of the old one. 5. The old store is decommissioned after a safety window, in case a rollback is needed. ## Why replay rather than hand-editing This pattern exists because bugs in projection logic are effectively inevitable over the life of a system: - a schema change on the source event gets missed by a handler - an event type is mishandled - an aggregation has an off-by-one Since the write side (or the durable event log/broker feeding the projection) is the actual source of truth and the projection is purely derived from it, the natural way to recover correctness is to reprocess history through fixed logic — **not** to hand-edit rows in the projection store, which fixes the symptom for existing data but does nothing about the underlying bug and is easy to get subtly wrong. This same replay mechanism also covers the related, happier case of adding an entirely new projection after the fact: if a new query need arises, and the event history already exists, that history can simply be replayed once to build the new projection from scratch, with no special coordination with the write side. ## What it costs - The central trade-off is **cost**, and it scales directly with the size of the retained event history — a full replay against a huge stream can take hours or longer to reprocess, which is the price paid for keeping the write side/event log lean and pushing derived complexity onto replayable projections instead. - **Retention becomes a hard constraint** as a result: if old events are truncated or compacted away, replay-from-scratch is only possible back to whatever the retention window still contains — a limitation sometimes mitigated by periodically snapshotting projection state as a cheaper resume point, a lighter-weight technique than the aggregate-level snapshotting used in fully event-sourced write models. - Running a blue-green rebuild also **temporarily doubles infrastructure cost** (two live copies of the store), and it strictly requires the handler to be idempotent and deterministic, since the whole approach depends on reprocessing producing the same correct result regardless of how many times it's run. ## Failure modes - **Rebuilding in place** is the most common production mistake — truncating and refilling the very table the API currently reads from — which opens a window where reads see partial, empty, or otherwise inconsistent data instead of either a complete old view or a complete new one. - A related failure is **letting the rebuild race with the live write path** if both end up sharing the same handler code or table without proper isolation, corrupting counts mid-flight. - Running the rebuild against **shared production infrastructure** without rate limiting or backpressure can starve live traffic of I/O capacity exactly while you're trying to fix things. - If a projection has **multiple independent consuming services**, forgetting to cut all of them over consistently and at roughly the same time means some callers answer from the old projection and others from the new one, producing visibly inconsistent answers to the same question depending on who's asked. - And discovering that **a retention policy has already expired** the very history a rebuild needs is a particularly painful failure, usually only found the moment a rebuild is actually attempted. ## A worked example A concrete example: a search-index projection built with Elasticsearch develops a mapping bug after a new field is added to the source event without the handler being updated to index it. The fix is to: 1. deploy the corrected handler as a consumer writing into a brand-new index (say, `orders-v2`), 2. run it against the retained event topic from the earliest offset, 3. monitor its lag until it reaches the current tail, 4. and then flip the application's index alias from `orders-v1` to `orders-v2` — closely mirroring Elasticsearch's own recommended reindex-and-alias-swap workflow for schema changes, and directly analogous to blue-green deployment applied to a read projection instead of a whole service.
- What happens to your rebuild strategy if your event source only retains 30 days of history and a bug has been silently corrupting a projection for 90 days?A full replay-from-genesis rebuild can't recover the first 60 days of corrupted state because that history is already gone — you're limited to correcting the last 30 days via replay and have to fall back to manual reconciliation, a compensating data-fix script, or accepting the historical inaccuracy for anything outside the retention window. This is a strong argument for either longer retention on critical streams or periodic projection snapshots as a safety net.
- Why is running the rebuild on separate infrastructure from the live projection important?A full historical replay is a sustained, high-throughput read against the event source and a high-throughput write against the new store — if it shares the same database, consumer group, or network capacity as live production traffic, it can starve live reads/writes and degrade the user-facing system precisely while you're trying to fix it.
- If your projection consumers are multiple independent services rather than one, how does that complicate the cutover step?Each consuming service needs to be pointed at the new store consistently and roughly simultaneously; if the cutover isn't coordinated, some callers answer from the old projection while others answer from the new one, producing visibly inconsistent answers to the same question depending on which service is asked. A shared alias/router that all consumers go through, flipped once, avoids this.
It's like repainting a mural by working on a fresh wall next to the old one and only unveiling it once it's finished, rather than painting over the original wall while people are still looking at it — visitors always see a complete picture, either the old one or the finished new one, never a half-painted wall.
saying these in an interview costs you the question
- proposes truncating and refilling the same live table as the rebuild strategy
- doesn't mention that replay is bounded by event retention
- assumes rebuild is instantaneous regardless of history size
- no plan for cutting reads over atomically
- conflates fixing the projection with editing the source-of-truth event data