Once a team has CQRS in production — separate write and read models kept in sync via projections — what ongoing operational burdens does that dual-model setup create that a single shared model doesn't have, and how do teams typically manage them?
answer
- consumer lag monitoring
- replay/rebuild from source
- dead-letter events
- schema version skew write vs read
- on-call surface roughly doubles
basics
~20 sNow there are two models to deploy, monitor, and fix bugs in, plus a pipeline that keeps them in sync — and if that pipeline breaks or falls behind, the read side can go stale or wrong without anything obviously crashing.
solid answer
~40 sThe dual model creates ongoing costs beyond the initial build: schema/contract changes now have to be coordinated across write model, event schema, and every read projection consuming it; the projection pipeline itself needs monitoring (consumer lag, dead-letter handling, replay/rebuild tooling) because it can silently fall behind or drop events; and on-call surface roughly doubles since incidents can originate in either model or in the sync layer between them. Teams manage this with event schema versioning/contracts, replay-from-source tooling to rebuild a read model from scratch when it drifts, lag alerting/dashboards, and treating projections as disposable — cheap to tear down and regenerate rather than patched in place.
go deeper
Should recognize that two models means more to build and that they can get out of sync, without needing to detail lag monitoring or replay mechanics.
Should name concrete operational mechanisms — lag monitoring, replay/rebuild, schema versioning — as the specific ongoing costs, not just 'more complexity.'
Should be able to design the coordination discipline (event contracts, deployment ordering across write model and projectors) that prevents the version-skew failure mode described above.
Should treat this as a staffing/ownership problem too — who owns projector on-call, how many read models the org can sustainably operate — and know when to consolidate or retire projections that no longer earn their operational cost.
## Why the burden is underestimated The operational burden of CQRS is easy to underestimate at design time because most of the pain shows up after launch, once the system has to evolve and be kept healthy under real traffic. Mechanically, a running CQRS system has at least **three moving parts** that all have to stay coherent: 1. the write model and its schema; 2. the event stream (or whatever mechanism carries changes from write to read side); 3. one or more read-model projections, each potentially with its own schema and even its own storage technology. Every one of these can independently break, and a production incident can originate in any of the three — which means the operational surface isn't additive, it's closer to multiplicative, because a bug in the read model might actually be caused by an event schema change nobody propagated to the projector. ## Burden one — change coordination The first burden is **change coordination**. In a single-model system, changing a field is one migration in one place. In CQRS, changing what a write emits (say, adding a new order status) means: - updating the write model; - updating the event schema (and deciding whether that's a breaking change for existing consumers); - updating every projector that builds a read view touching that field. Until all of those are deployed and caught up, you have **version skew**: old events with the old shape sitting next to new events with the new shape, which the projector must handle gracefully (usually via explicit schema versioning and backward-compatible event contracts, or a migration/dual-write period). Skipping this discipline is how teams end up with a projector silently crash-looping on the first event with a new field it doesn't recognize. ## Burden two — the pipeline as infrastructure The second burden is the projection pipeline itself as a piece of infrastructure to run. It needs its own monitoring: - **consumer lag** (how far behind is the read model from the latest write); - **dead-letter handling** for events a projector can't process; - and — critically — **the ability to rebuild a read model from scratch** by replaying the event history, because at some point a projector will have a bug, get out of sync, or need a new denormalized shape for a new query, and 'fix the data in place' is rarely safe or complete. Teams that treat this well design projections as **disposable**: cheap and fast to drop and regenerate from the write-side source of truth, which in turn requires the event store or write-side change log to actually retain enough history to replay from (or a periodic snapshot/backfill job if it doesn't). ## Burden three — doubled on-call and debugging surface The third burden is roughly doubled on-call and debugging surface. A support ticket saying 'my data is wrong' now requires figuring out which side is wrong: - is the write model's business logic incorrect; - did an event fail to publish; - did the projector drop or mis-transform it; - or is the read model just lagging and will self-correct in a few seconds? Each of those has a different fix and a different owner. Teams typically mitigate this with distributed tracing that follows a change from command through event to projection update, and with clear ownership boundaries (who's paged for projector lag vs. who's paged for write-model errors) so an incident doesn't stall on 'whose problem is this.' ## What it looks like when it goes wrong A concrete illustration: a team running an order-processing write model with a search-oriented read projection adds a new 'gift-wrapped' field to orders. They ship the write-model change and the event schema update together, but the search-index projector — a separate deployable owned by a different subteam — doesn't get updated until a week later. In between, the projector either drops the new field silently (read model quietly missing data, discovered only when someone builds a report on gift-wrapped orders) or, worse, crashes on an unexpected field depending on how strictly it validates events, taking the entire search index update pipeline down and causing all read-model updates — not just the new field — to stall and lag further behind every minute the projector is down. Recovering requires both fixing the projector and then replaying the backlog of events it missed, which is only possible because the event log retained history long enough to replay from. ## The lesson The overall lesson for a mid-level engineer is that CQRS's operational cost isn't just 'more code' — it's an ongoing tax paid in: - **coordination discipline** (schema/versioning contracts across models); - **infrastructure** (lag monitoring, dead-letter queues, replay tooling); - **incident response complexity** (which layer owns a given class of bug). A team should only take that tax on when the asymmetric-scaling or shape benefit clearly outweighs it — otherwise a single model, or read replicas of that single model, is operationally far cheaper to run.
- How would you detect that a read-model projector has silently stopped processing events?Instrument consumer lag directly — track the gap between the latest event offset/timestamp in the source stream and the last one the projector successfully processed, and alert when that gap exceeds a threshold. Relying on user complaints or absence of errors is not sufficient, because a stalled consumer often fails silently rather than crashing loudly.
- Why is 'rebuild the read model from scratch by replaying events' considered safer than patching stale or wrong data directly?Direct patches only fix the symptom you noticed and can introduce further drift from the true source of truth, while a full replay from the write-side event history regenerates the projection deterministically and consistently. It requires the event log to retain sufficient history (or snapshots), which is why retention policy is itself an operational decision teams have to make deliberately.
Like running a print newspaper (write side, one authoritative edit) alongside a live news ticker (read side) fed by a wire service — if the wire feed drops a story or garbles formatting, the ticker is wrong even though the newspaper's own record is fine, and someone has to notice the ticker is stale and know how to re-feed it from the newspaper's archive.
saying these in an interview costs you the question
- Thinks CQRS's cost is 'just more code to write once'
- No mention of monitoring lag or handling stuck/failed projections
- Assumes schema changes on the write side don't need coordinating with read-side consumers
- Can't describe how they'd recover a broken/drifted read model
- Treats the read model as safe to hand-edit to fix bugs