A team is about to apply CQRS with an eventually-consistent read model to a feature that checks a user's account balance before authorizing a withdrawal. As the architect reviewing the design, what makes this a case where you'd push back and require strong (immediate) consistency instead, and what are the concrete ways to get it without abandoning CQRS entirely?
answer
- read for display vs read for authorization
- check-then-act / TOCTOU race
- push the invariant into the write-side transaction
- narrow exception, not abandoning CQRS system-wide
- a faster projector shrinks the race, it doesn't close it
basics
~20 sMoney is a case where showing the wrong, stale number can cause real harm, letting someone overdraw because the read side hadn't caught up yet. For that one check, read the current true value directly instead of trusting the lagging copy, even if the rest of the app still uses the fast, eventually-consistent read model.
solid answer
~60 sThe deciding factor is the cost of acting on stale data versus the cost of strict consistency: an authorization decision made against a stale balance can allow an overdraft, a real correctness violation, not just a cosmetic UI glitch. Contrast with a 'like count', where staleness is harmless. Concrete ways to get strong consistency for this one check without discarding CQRS system-wide: (1) route the balance-check read directly against the write-model database or a synchronously-replicated strongly-consistent replica instead of the async read model; (2) push the invariant into the write side itself, making 'withdraw' a single command whose handler checks and decrements balance atomically inside its own transaction, so the read model is used only for display, never for authorization; (3) use a strongly-consistent read API some databases expose, such as a leader or quorum read, scoped narrowly to this one query. The general principle: keep CQRS's async read model for reads that tolerate lag, and carve out specific invariant-critical decisions to read and write against the authoritative source directly.
go deeper
Not expected to make this call unprompted; should be able to recognize, when told, that checking money against stale data sounds risky.
Should be able to explain why a stale balance read before authorizing a withdrawal is dangerous, i.e. the TOCTOU risk, when prompted.
Should proactively flag this pattern during design review and propose at least one concrete strong-consistency mechanism scoped to the specific check.
Should set the organizational principle - read-for-display versus read/write-for-authorization - once, as reusable review guidance, and weigh it against the org's actual risk tolerance and compliance obligations across many features, not just this one.
## The question that decides it The decision hinges on a single question: is this read purely informational, used only for display, or does it feed a subsequent decision that changes state? - **If a read result is used to decide whether to authorize a command** — 'read the balance, then decide whether to allow the withdrawal' — it sits squarely in the danger zone, because it compounds a classic **check-then-act** race with the added risk of the checked value being stale on top of merely being racy. - **A display-only read**, by contrast, being briefly wrong just means a number on screen is a few seconds out of date, which self-corrects and costs nothing real. ## Why the distinction exists This distinction exists because CQRS's core benefit — scaling reads independently of writes by decoupling them — necessarily reintroduces a staleness window as its price. Some domain invariants are strict and must never be violated: - financial balances - unique-username reservation - inventory oversell prevention - seat-booking uniqueness For these, the cost of a violation — real money lost, a double-booked seat, a compliance incident — vastly outweighs the read-scaling benefit for that one specific decision path, even though the same eventually-consistent read model remains perfectly appropriate for the other 95% of the application's reads. ## The concrete fixes The concrete fixes trade off differently. 1. **Routing the authorization read directly to the write-side database** (or a synchronously-replicated, strongly-consistent replica) defeats part of CQRS's scaling purpose for that one path — it couples read load back onto the write database and forfeits the fast denormalized store — but it's a narrow, deliberate, well-scoped exception, not an abandonment of CQRS elsewhere. 2. **Pushing the invariant fully into the write side** is architecturally the cleanest option: the write model's own command handler re-reads the current authoritative state and makes the decision atomically inside its own transaction, so it doesn't matter at all what a lagging read model displays to the user, because the read model is never consulted for the actual decision. This does mean duplicating some 'read' logic (checking sufficiency) inside the write model, which can feel redundant with what the read model already computes for display, but that redundancy is the point — the write side must never trust a value it didn't itself just read inside its own transaction. 3. **A quorum or leader read**, a narrower option some databases support directly, scoped to just this query, trading a small latency cost for a strong-consistency guarantee without rearchitecting anything. ## What getting it wrong looks like Getting this wrong produces a specific, well-known failure mode: two concurrent withdrawal requests both read a stale-but-locally-sufficient balance from the read model at nearly the same instant, both get authorized against that stale value, and the account goes negative — a textbook **time-of-check-to-time-of-use (TOCTOU)** race, made worse here because the 'check' itself was reading a value that could already be behind the true, currently-committed balance. A common but ultimately false fix is simply tightening the projector pipeline to reduce lag — making the read model catch up faster narrows the race window and reduces how often it's hit, but any nonzero lag, however small, still leaves the race mathematically possible; speed is a mitigation of frequency and severity, not a structural guarantee, and treating it as a fix creates a false sense of safety instead of actually closing the hole. The structural fix has to be moving the check into the write model's own atomic transaction, not making the async pipeline faster. ## Where it shows up This dual-speed resolution — fast, loosely-consistent reads for display; strict, authoritative checks inside write-side transactions for anything that authorizes a state change — is the standard guidance in CQRS literature, and shows up concretely in real systems: - banking core systems and payment processors enforce balance checks inside the ledger's own transaction rather than trusting any denormalized reporting view; - e-commerce checkout flows commonly re-validate stock inside the order-placement write transaction even though the 'in stock' indicator a shopper sees on the product page comes from a fast, potentially-stale read model. The principle generalizes well beyond this one example: any time a CQRS read feeds an authorization decision rather than merely rendering a screen, that specific read deserves scrutiny for whether it needs to be pulled out of the eventually-consistent path entirely.
- Why isn't 'just make the projector faster' a sufficient fix for the balance-check scenario?Any nonzero lag window, however small, still leaves a race where two concurrent operations can both read a not-yet-decremented balance and both get authorized. Speed reduces how often the window gets hit but doesn't close it structurally, so it's a mitigation of severity and frequency, not a guarantee - a high enough concurrency or throughput can still hit it.
- Is it ever acceptable to authorize against the read model if you add a compensating step afterward?Sometimes - for example, allowing a withdrawal optimistically and reconciling or flagging it for reversal if the write-side ledger later shows it was actually insufficient, similar to how some payment holds work. But this shifts risk toward needing to claw back money from a user after the fact, which is often worse than a brief added latency cost, so it's typically reserved for cases where reversal is cheap and socially acceptable, not core banking withdrawals.
- How would you explain this trade-off to a product manager who just wants 'the fast version'?Frame it as a narrow, well-justified exception rather than an all-or-nothing choice - the vast majority of the product keeps the fast, eventually-consistent read model and its scaling benefit, and you're asking for a direct check on just the one authorization decision where being wrong costs real money. Quantify the cost of a bug here, such as overdrafts or compliance exposure, against the marginal latency cost of that one direct read.
Like a bouncer checking ID against a guest list that's synced from the front office every few minutes, versus calling the front office directly for the one VIP-only, high-stakes decision - for the regular line, the slightly-stale printed list is fine; for the one decision where letting the wrong person in is a real problem, you pick up the phone and ask the authoritative source directly instead of trusting the printout.
saying these in an interview costs you the question
- treats every read the same regardless of whether it feeds an authorization decision
- believes reducing projector latency alone eliminates the race condition
- proposes making the entire system strongly consistent everywhere as the only fix
- doesn't recognize the check-then-act/TOCTOU risk of combining a stale read with a state-changing command
- can't name a concrete mechanism (direct write-side read, invariant inside the write transaction, quorum read) to get strong consistency for the narrow case