You're advising several teams across an organization on whether to adopt CQRS for their services. What decision framework would you give them to evaluate the trade-off consistently, and how would you handle a team that already adopted CQRS but the benefit never materialized?
answer
- three quantified signals: ratio, shape mismatch, independent scaling
- default to single model + cheaper mitigations first
- un-adoption needs same rigor as adoption
- measure on-call cost vs. actual benefit
- ADR template forces evidence not hunches
basics
~20 sGive teams a checklist of concrete signals (read/write ratio, shape mismatch, independent scaling need) instead of letting them decide on gut feel, and if a team adopted it without those signals, help them measure the actual cost/benefit and consider simplifying back to one model.
solid answer
~50 sA useful framework asks teams to quantify, not guess: what's the measured read/write ratio and has it caused a specific incident or bottleneck; does the read query shape genuinely diverge from what the write schema supports without expensive JOINs or write-schema pollution; and is there a real need to scale or deploy reads and writes independently (different regions, different SLAs). If none of these are true, default to a single model plus cheaper mitigations (read replicas, caching, materialized views) first. For a team that already adopted CQRS without benefit, don't assume the fix is always 'rip it out' — measure the current operational cost (on-call load, incident rate tied to sync issues) against any latent benefit, and if the cost clearly dominates, plan a deliberate, incremental collapse back to a single model rather than leaving a permanently unjustified architecture in place.
go deeper
Not generally expected to own this; a reasonable answer is recognizing that different teams might make inconsistent choices and that's worth having a shared guideline for.
Should be able to list a few concrete signals (ratio, shape, scaling need) that would justify CQRS, even without designing the full organizational process.
Should be able to apply the framework rigorously to a single team's decision and reason about un-adopting a specific instance if the cost doesn't pay off.
Should design the cross-team governance process itself (ADR template, external review, revisit cadence) and reason about the organizational cost of both over- and under-adoption at scale, including the staged-collapse plan for reversing a bad adoption.
## Why this becomes a governance problem At the organizational level, the CQRS trade-off question stops being about any one service's technical merits and becomes a **governance problem**: many teams, making the same build-vs-simplicity decision independently, tend to converge on inconsistent answers unless given a shared framework — - some **over-adopt** it because it's fashionable or was used successfully on one high-profile project; - others **under-adopt** it and quietly suffer contention problems a split model would have fixed. A principal engineer's job here is less 'decide for every team' and more 'give every team the same good decision procedure and the same vocabulary to justify their answer.' ## The three questions to answer first A workable framework has teams answer three quantifiable questions before considering CQRS, deliberately framed to require evidence rather than intuition. 1. **First**: what is the measured (not guessed) read/write ratio, and is there a specific, attributable incident or bottleneck caused by it — a write timeout under read load, an over-provisioned database sized mostly to serve reads it wouldn't otherwise need? If the answer is 'we assume reads will be much higher eventually' with no current data, that's not yet a justification; it's a prediction, and predictions should be revisited when they become measurements. 2. **Second**: does the read side need a genuinely different query shape than the write schema naturally provides — specifically, would satisfying the read pattern on the write schema require either expensive multi-table JOINs at read time or denormalizing fields into the write schema that would slow or complicate writes? If a straightforward index or a read replica of the existing schema would serve the read pattern fine, the shape mismatch isn't real yet. 3. **Third**: is there an actual operational need to scale, deploy, or geographically distribute reads and writes independently — for example, needing read replicas in multiple regions for latency while keeping a single write region for consistency — as opposed to a hypothetical future need. ## The default when none of the three hold If none of the three hold, the framework's default is explicit: don't adopt CQRS, and instead reach for cheaper mitigations first — - **read replicas** for volume; - **caching** for hot reads; - **a materialized view or scheduled rollup table** for one awkward reporting query. Each of which solves a narrower version of the problem at a fraction of the ongoing cost (no dual codebase, no event-sync pipeline, no org-wide staleness reasoning). This ordering matters because it keeps CQRS as an **escalation, not a default**, which directly counters the 'resume-driven architecture' failure mode where an interesting pattern gets applied because it's interesting rather than needed. ## The harder half — the benefit that never showed up The harder half of this question is the team that already adopted CQRS and the benefit never showed up — a real and common situation, because architecture decisions are made ahead of actual load, and sometimes the load never materializes the way the team predicted, or the read-shape divergence turns out to be smaller in practice than it looked on paper. The instinct to immediately 'rip it out' is as dangerous as the instinct to never revisit it: ripping out a read model that other services or reports now depend on is itself a real migration with its own risk, and premature reversal can cost more than just leaving a slightly-over-engineered system running. The right move is to **measure before deciding**: quantify the ongoing operational cost concretely — how much on-call load and how many incidents in the last quarter trace back to projector lag, event-schema coordination, or staleness bugs, versus how much value the split is actually delivering (is the read model genuinely serving traffic a single model couldn't have, or has traffic never grown into the scenario it was built for). If the cost clearly and durably dominates with no offsetting benefit, plan a deliberate, incremental collapse: - freeze new dependents on the read model; - migrate existing read traffic back to direct queries (possibly via a replica) one consumer at a time; - and only decommission the projection pipeline once nothing depends on it — treating the un-adoption with the same rigor and staged rollout as the original adoption, rather than a rushed reversal. ## Making the decision reviewable A concrete organizational pattern that works well is codifying this as a lightweight architecture decision record (ADR) template specific to CQRS adoption — forcing whoever proposes it to fill in the three quantified signals above with actual numbers or a stated plan to measure them, reviewed by someone outside the immediate team. This doesn't remove judgment from the process, but it replaces 'I have a hunch CQRS is right here' with a paper trail that can be revisited later — which is exactly the artifact needed six months on to evaluate the team whose benefit never materialized, without relitigating the decision from scratch on memory and vibes.
- How would you avoid this framework itself becoming a rubber-stamp exercise where teams fill in numbers to justify a decision they've already made?Require the quantified signals to be reviewed by someone outside the proposing team — an architecture review board or a peer principal — specifically empowered to push back on weak evidence, such as a 'ratio' based on projected rather than measured traffic. Pairing the framework with a mandatory revisit date (e.g. re-review the decision against actual data six months post-launch) also catches cases where the original justification didn't hold up.
- What's a risk of being too aggressive about collapsing an underused CQRS setup back to a single model?If other services or reports have quietly come to depend on the read model's specific denormalized shape, collapsing it without first identifying and migrating those consumers can break them outright. The safe approach is to freeze new dependents, inventory existing ones, and migrate each off before decommissioning, treating it as a staged deprecation rather than a single cutover.
Like a company-wide policy on when to build a custom in-house tool versus buying off-the-shelf software: without a shared checklist, some teams over-build bespoke systems for problems a standard tool would solve, others under-build and suffer — the fix isn't a blanket rule but a consistent, evidence-based decision procedure everyone uses, plus a willingness to decommission a custom tool later if it turns out the standard option would've been fine.
saying these in an interview costs you the question
- Gives a blanket 'always/never use CQRS' rule instead of a decision framework
- Treats an already-adopted CQRS system as untouchable once built, regardless of measured cost
- Proposes ripping out an underused read model immediately without checking for existing dependents
- Uses projected/hoped-for future traffic as sufficient justification instead of requiring measured evidence
- Has no process for someone outside the proposing team to review the justification