skip to content

A read-model projection has a bug that has been silently producing wrong denormalized data for three months. What is the safe way to fix and recover this in production, and what design decisions made earlier would have made this easier?

level: principalimportance: should knowfreq 35%

answer

  1. read model = disposable, rebuild from source of truth
  2. out-of-band build + validate + blue-green cutover, don't patch in place
  3. full rebuild vs targeted repair trade-off
  4. retention limits can block full replay
  5. 3-month-silent bug = missing reconciliation process

basics

~20 s

Fix the bug in the code that builds the read model, then re-run that building process from scratch so the read model gets recomputed correctly, instead of hand-editing the wrong data. This only works cleanly if the real source of truth was never lost, only the derived copy.

solid answer

~50 s

Because a read model is a derived, disposable copy — never the source of truth — the safe recovery path is: fix the projector bug, then rebuild the read model from the write side (replaying the event log, or re-scanning the write database) rather than trying to patch the corrupted data in place. In practice this needs a way to run the fixed projector against historical data without taking the live read model offline: build the corrected version into a fresh table/index, backfill it, validate it against a sample of known-good expected values, then atomically cut traffic over (blue-green swap). Decisions that make this tractable, made ahead of time, include: retaining the event log/change history long enough to replay from, versioning read models so old and new can coexist, and making projectors idempotent so a partial or repeated rebuild is safe.

go deeper

for a junior

Should have the basic instinct that you should fix the bug and somehow regenerate the correct data rather than hand-editing wrong values, even if not fully precise on mechanism.

for a middle

Should describe rebuilding the read model from the source of truth (event log or write DB) as the standard recovery approach, and know not to edit the read model directly.

for a senior

Should design the concrete rebuild-and-cutover process (out-of-band build, validation, blue-green swap) and weigh full rebuild versus targeted repair with real trade-offs.

for a principal

Should identify that a three-month-silent bug points to a missing reconciliation/detection process as a root organizational gap, and propose durable prevention (retention policy, idempotent projectors, versioned read models, periodic checksums) beyond just fixing this one incident.

## The two halves of the recovery When a read-model projection has been silently wrong for months, the recovery has two separate parts: fixing the code that produced wrong data going forward, and repairing the historical data that's already wrong. Because a read model is, by design, a **derived and disposable copy** of the write side rather than a source of truth in its own right, the safe move for the historical repair is not to hand-edit the bad rows but to regenerate the read model from whatever the authoritative history is — a fully event-sourced event log, a change-data-capture stream retained long enough, or, at minimum, the current state of the write-side database if per-event history isn't available. ## The practical sequence The practical sequence looks like: 1. identify and fix the root cause in the projector code; 2. decide what to replay from and confirm that source actually goes back far enough to cover the three-month window; 3. build the corrected read model **out-of-band**, in a new table, index, or namespace, rather than in place, so the currently-serving (wrong) read model keeps answering live traffic while the rebuild runs; 4. validate the new build against a sample of hand-checked expected values, or against an independent reconciliation source, before trusting it; 5. and then atomically switch reads over to the new version — a **blue-green cutover** — and only decommission the old one once the new one has been observed correct under real traffic for a while. ## Why regenerating beats hand-editing This whole approach only works because the write side, or the event log, is still the ground truth throughout the incident: the read model's job was always to be a cheap, fast, rebuildable projection of that truth, never a place where facts are created. If a team instead tries to 'just UPDATE the bad rows' directly in the read store, they're now doing manual, ad hoc data surgery against a live table under production load, with: - no systematic way to know they found every affected row - no audit trail proving the fix was complete - a real risk of making things worse And the fix has to be re-derived by hand for every future bug of this kind, instead of 're-run the (now-correct) projector.' ## Rebuild everything, or repair a slice The main trade-off is the cost and blast radius of a full rebuild versus a narrower, targeted repair. | Full rebuild | Targeted repair | |---|---| | Replaying the entire event history or fully re-scanning the write database. Conceptually simple and provably correct if the projector logic is now right, but at scale it can take hours or days, requires double the storage during the cutover window, and needs the replay source to still exist and be complete going back far enough. | Compute exactly which read-model rows were affected by the bug (e.g., events of a specific type in a specific date range) and reprocess only those. Much faster and cheaper, but it's harder to prove correct: it requires being certain the bug's blast radius was fully and correctly characterized, and any mistake in that scoping leaves silently-wrong data behind, which is exactly the failure mode that caused the three-month incident in the first place. | Senior/principal-level judgment here is choosing based on data volume, replay-source availability, and confidence in the bug's scope, not defaulting to either option reflexively. ## Failure modes Several failure modes can make this materially harder than the clean version above. 1. **Retention.** If the event log has a retention policy or has been compacted (common in stream-based systems, where topics are often retained for days or weeks, not months), a full replay from three months ago may simply be impossible, forcing a hybrid strategy of 'restore from the oldest available snapshot, then replay whatever incremental history remains.' 2. **A moving write side.** If the rebuild takes long enough, the write side keeps moving during the rebuild, so a naive 'replay to now, then cut over' can miss events that arrived during the rebuild window itself — the rebuild process needs to either catch up in a tight loop until it's within an acceptably small lag, or the cutover needs to briefly pause writes or use a dual-write/backfill overlap strategy. 3. **The missing safety net.** The deepest failure mode is the one implied by 'silently producing wrong data for three months': there was no reconciliation process comparing the read model against the source of truth, so nothing caught the drift until a human noticed by accident — the bug being three months old is itself evidence of a missing safety net, not just a code defect. ## The decisions made long before the bug The design decisions that make this whole class of incident cheaper to recover from, and faster to detect, are ones made well before the bug ever ships: - retaining the write-side event history (or CDC log) for long enough to replay from, even if it means a cheap cold-storage tier rather than the live topic - writing projectors idempotently so any replay — full, partial, or repeated — produces the same result without needing careful sequencing - versioning read models (`read_model_v2` alongside `read_model_v1`) so a rebuild can happen entirely out-of-band with a clean cutover instead of a risky in-place migration - running a periodic reconciliation job that checksums or spot-checks the read model against the source of truth, so drift is caught in days, not quarters

  • How would you decide between a full rebuild and a targeted repair when you're not fully certain you've identified every event or row affected by the bug?
    Default to the full rebuild when you can't prove the affected scope with high confidence — an incomplete targeted repair silently leaves bad data behind, which is a worse outcome than a slower but provably-complete rebuild. Reserve targeted repair for cases where the bug's blast radius is mechanically derivable, e.g., 'every event of type X between these two timestamps,' and you can verify completeness against a count or checksum.
  • What's the risk of doing the historical repair by directly UPDATEing the bad rows in the live read-model table instead of rebuilding out-of-band?
    You're modifying production data under live traffic with no atomic cutover, no easy rollback if the fix is itself wrong, and no independent verification step before customers see the result — and you have to be certain you've found every affected row, with no systematic way to prove that, unlike a rebuild driven mechanically from the source of truth.
  • If the event log has already been compacted or expired for part of the affected window, what options remain?
    Fall back to whatever coarser-grained history still exists — a nightly write-side database snapshot, an audit log, or backups — to reconstruct a starting state as close as possible to the start of the gap, then replay whatever fine-grained event history remains from that point forward; accept that any gap with truly no recoverable history may require manual reconstruction or a documented, disclosed data-quality caveat for that window.

It's like discovering a printed report has had a formula error for months — you don't go through and hand-correct each printed copy in circulation; you fix the spreadsheet formula, regenerate a fresh report from the underlying ledger, verify the new report against a few known totals, and then replace the old report with the new one, keeping the old one on hand just in case something's still off.

saying these in an interview costs you the question

  • Suggests directly patching bad rows in the live read model as the primary fix, with no mention of rebuilding from source
  • No mention of validating the rebuilt data before cutover
  • Assumes the event log/history is infinitely retained with no discussion of retention limits
  • Doesn't address the fact that a three-month-old undetected bug implies a missing reconciliation/monitoring process
  • Treats full rebuild as free/instant with no discussion of duration, storage, or catch-up during rebuild

context