A projection backing a high-traffic, customer-facing read API takes four hours to fully rebuild from the event log after a handler bug fix. What techniques let you deploy the corrected projection without taking that read API offline during the rebuild?
answer
- blue-green for read models
- new projection, new storage, own subscription
- atomic cutover after full catch-up
- parallel replay across streams for speed
- shadow validation before flipping traffic
basics
~20 sBuild the new, corrected version in a completely separate table while the old one keeps serving traffic, and only switch customers over to the new one once it's fully caught up and checked - like renovating a store in a new building instead of closing the old one down for the day.
solid answer
~40 sThe core technique is blue-green (or side-by-side) projection rebuilding: deploy the corrected handler as a brand-new projection writing to its own table or index under a new name, start a fresh catch-up subscription for it from position zero while the existing, buggy projection keeps serving all live reads unaffected, let it fully catch up to the live tail, validate its output through spot-checks or a shadow comparison against the old one, and then atomically flip a router or feature flag so new reads go to the rebuilt projection. Related techniques include parallelizing the replay across partitions or streams to shrink the four hours, and versioned or aliased read-store names so the cutover is a metadata change rather than a data migration.
go deeper
Not expected to originate this; should be able to follow the basic idea that you build the new version separately and switch over once it's ready.
Should understand blue-green or side-by-side rebuilding as a named pattern and why it avoids downtime, even if unfamiliar with parallelization or shadow-validation details.
Should propose the side-by-side pattern independently and reason about the storage and compute cost trade-off and cutover atomicity concerns.
Should design the full rollout: parallelized replay strategy, load-impact management on the shared event store, pre-cutover validation methodology, and a rollback plan, as a repeatable organizational playbook rather than a one-off fix.
## The fundamental move The fundamental move that makes a zero-downtime rebuild possible is **refusing to touch the projection currently serving traffic at all**: instead of truncating and refilling the existing table in place - which either takes reads offline or serves a half-rebuilt, inconsistent view mid-process - the corrected handler is deployed as an entirely new, independent projection instance, writing into its own storage (a new table, a new index, a new collection) under a distinct name or version tag. - This new instance starts a completely fresh **catch-up subscription** from position zero against the same event log, and because it's physically separate storage, it can take however long it needs - the four hours in this scenario - without any customer-visible impact, since every request during that window continues to be served by the old, still-running projection exactly as before. - Only once the new projection's subscriber reports it has caught up to the live tail does **cutover** happen: typically an atomic metadata change - flipping a router's target table name, updating a feature flag, or repointing an alias - that redirects new reads to the rebuilt projection. This is commonly called **blue-green deployment applied to a read model**: the old version keeps serving while the new one is built and verified, and only the final switch is user-visible, and even that switch is designed to be instantaneous and reversible. ## Why rebuilding in place is the wrong choice This pattern exists because the naive alternative - rebuild in place - forces an unacceptable choice for high-traffic systems: either take the API offline for the rebuild duration, which is rarely tolerable for a customer-facing service, or leave it serving reads against a table that is actively being truncated and repopulated, which produces genuinely wrong answers, missing rows, partial aggregates, for the entire rebuild window, arguably worse than an honest outage. Side-by-side rebuild sidesteps both bad options by decoupling 'building the correct new state' from 'serving current requests' entirely, at the acceptable cost of temporarily running two copies of the data. ## The trade-offs The trade-offs are real and worth naming precisely. - **Storage and compute cost** roughly double during the rebuild window, since both the old and new projection exist simultaneously, plus the new one is under active heavy write load from replay while the old one continues normal traffic. - Four hours is also **an estimate, not a guarantee** - replay throughput can be affected by contention on the shared event store between the new subscription's replay reads and every other live consumer reading from that same store, so a poorly throttled rebuild can either take longer than expected or degrade the latency of unrelated live subscribers pulling from the same log. - **Cutover itself**, even though conceptually atomic, has to be designed carefully: if the router flip isn't instantaneous across all serving instances, for example a rolling deploy of a config change rather than a single shared flag read, different instances can briefly serve different projections for the same logical request, which needs its own consistency consideration, especially if the client makes several dependent requests in a short window. ## Shrinking the four hours, and validating the result Two complementary techniques shrink the four hours itself, rather than just tolerating it: 1. **Parallelizing the replay** across independent partitions or aggregate streams - since ordering only needs to be preserved within a single aggregate's stream, not globally across all aggregates, many event stores allow concurrent replay of different streams by different worker threads or processes, which can turn a four-hour serial replay into a fraction of that with enough parallelism. 2. **Starting the new projection's replay from a recent full-state snapshot** rather than position zero, if the source data supports snapshotting, which avoids re-deriving state that was already correct and unaffected by the bug. Validation before cutover is the other half of doing this safely at principal level: shadow-comparing a sample of the new projection's output against the old one's, accounting for the expected differences from the bug fix, and running the corrected projection against a slice of live traffic in read-only shadow mode before it's trusted with real customer requests, catches a wrong fix before it becomes customer-visible rather than after. ## Failure modes at this scale Failure modes at this scale of operation include: - cutting over before replay has genuinely reached the live tail, a race similar to the snapshot-then-listen problem in catch-up subscriptions, which serves a projection that's subtly behind and looks like a regression - underestimating the replay-load impact on the shared event store, degrading unrelated live consumers during the rebuild window - and, most costly, discovering post-cutover that the fix introduced a new bug, which is why some organizations keep the old projection warm and readable for a rollback window after cutover rather than immediately decommissioning it Axon Framework's projection-versioning support and Kafka Streams' state-store rebuilding are both concrete, named implementations of this side-by-side rebuild-then-cutover pattern in widely used event-driven frameworks.
- Why not just rebuild the existing projection table in place, e.g., truncate and refill it?That forces a choice between real downtime for the rebuild duration, or serving reads against a table that's actively half-empty and being repopulated, which produces visibly wrong answers the entire time - both worse than the side-by-side approach, which keeps the old, complete projection serving throughout and only switches over once the new one is fully built and verified.
- What determines whether replay for the new projection can be parallelized across multiple workers?Ordering guarantees typically only need to be preserved within a single aggregate's own event stream, not globally across all aggregates in the log, so independent streams or partitions can usually be replayed concurrently by separate workers without breaking correctness, which is what shrinks a long serial replay into a much shorter parallel one.
- What's a concrete way to validate the new projection before flipping traffic to it, beyond just 'it finished replaying'?Shadow comparison against the old projection's output on a sample of records, accounting for the expected deltas from the bug fix, and running the new projection against a slice of real, live read traffic in a read-only shadow mode to see how it behaves under production-like conditions before it's trusted with actual customer requests.
Like renovating a shop by building the improved version next door rather than closing the original for the afternoon: customers keep shopping in the old location the whole time construction happens, and only once the new shop is fully stocked and checked do you swap the sign and unlock the new front door.
saying these in an interview costs you the question
- proposes truncating and refilling the live projection table in place with no mention of downtime or serving inconsistent data
- treats replay completion as automatically safe to cut over without validation
- doesn't consider the load impact of a heavy replay on the shared source event store
- assumes parallelizing replay is always safe regardless of ordering scope
- no rollback plan if the rebuilt projection turns out to still be wrong post-cutover