skip to content

questions

6

In a CQRS (Command Query Responsibility Segregation) system, commands update a write model and queries read from a separate read model that is synchronized afterward. Why does a user sometimes not see their own change immediately after saving it, and what do engineers call the time window during which this can happen?

level: juniorimportance: must knowfreq 80%

answer

  1. write model vs read model
  2. async projector/CDC
  3. outbox pattern
  4. convergence, not immediacy
  5. lag = commit time to read-model-updated time

basics

~20 s

The read copy of the data updates a little after the write copy, through a background sync step. Until that catches up, queries can show old data. That gap in time is called the lag window (or replication lag).

solid answer

~40 s

CQRS splits the model that handles writes from the model optimized for reads; they're kept in sync asynchronously (via events, a message bus, or change-data-capture) rather than in the same transaction. That asynchronous propagation takes non-zero time - the 'lag window' or 'replication lag' - during which the read model can return data that doesn't yet reflect the latest write. This is eventual consistency: guaranteed to converge, not guaranteed to be immediate. Window size depends on event-bus latency, projector throughput, and batching, ranging from single-digit milliseconds (in-process handlers) to seconds or minutes (queued, cross-service projections). Teams accept this trade-off because it lets the read side scale and be modeled independently, at the cost of needing explicit handling for read-after-write expectations.

go deeper

for a junior

Should recognize that write and read sides are separate stores updated asynchronously, and be able to name 'eventual consistency' or 'lag window' as the reason for stale reads; not expected to design mitigations yet.

for a middle

Should describe the propagation mechanism (events, CDC/outbox) and estimate what drives lag size, and know this needs explicit handling in UX (spinners, optimistic UI, read-your-writes).

for a senior

Should design and instrument the propagation pipeline, choose synchronous vs asynchronous per use case, and put SLOs and alerts on lag.

for a principal

Should set org-wide guidance on when eventual consistency is acceptable versus when a use case demands synchronous reads, and own the trade-off across many read models and services.

## How a write reaches the read model **CQRS** separates command handling, which mutates the authoritative 'write model' (typically a normalized schema enforcing business invariants), from query handling, which reads from a separate 'read model' — often denormalized or materialized specifically for a UI screen or report, such as a document store, a search index, or a set of precomputed aggregates. Crucially, the two are not updated inside the same database transaction. When a command commits, the write side typically emits a domain event, or a row change is captured via **change-data-capture (CDC)** or a **transactional outbox**, which travels through some channel — an in-process event bus, a Kafka topic, an outbox-polling job — to one or more **projectors** that translate the event into an update against the read store. Each hop adds latency: - event publish - transport - consumer pickup - projection write - cache invalidation, if fronted by a cache The cumulative delay between the write committing and the read model reflecting it is the **lag window**. ## Why the two sides were split apart This split exists because write and read workloads have fundamentally different shapes: | Side | What it needs | |---|---| | Writes | need strict invariant enforcement, and often benefit from normalization and locking | | Reads | need to be fast, denormalized, and horizontally scalable, frequently serving far higher volume than writes ever do | Coupling both concerns into one schema and one transaction, as classic CRUD does, forces compromises on both sides: - either the write schema gets denormalized for read convenience (risking inconsistent invariants), - or the read queries stay slow and awkwardly shaped around normalized tables. CQRS lets each side scale and evolve independently — read models can be reshaped or entirely rebuilt without touching write logic, and multiple specialized read models (a search index, a cache, an analytics warehouse) can all be derived from the same stream of write-side events. ## The trade-off The trade-off is explicit: - **You gain** independent scalability and read models tailored exactly to query needs. - **You pay** with eventual consistency. Any code path that assumes 'if I just wrote X, reading X back will show the new value' breaks unless it's specifically designed around the lag window (see read-your-writes techniques). - **You also pay** with more moving infrastructure — an event bus or CDC pipeline, one or more projector processes, and the operational burden of monitoring and occasionally rebuilding/replaying read models from scratch when their shape changes or they get corrupted. ## Failure modes Several concrete failure modes show up in production. 1. **A projector can simply fall behind under load**, so the backlog and observed lag both grow monotonically until someone scales it out. 2. **A single malformed or unexpected event** — a 'poison message' — can stall a consumer entirely, freezing the read model at a fixed point in time until an engineer intervenes, which is a much worse failure than ordinary lag because it doesn't recover on its own. 3. **Events delivered out of order** (common across partitions in a distributed log) can, if the projector isn't order-aware, apply updates in the wrong sequence and leave the read model in an incorrect end state rather than just a temporarily stale one. 4. **Duplicate delivery**, which most at-least-once messaging systems guarantee will eventually happen, can double-apply an effect (e.g., incrementing a counter twice) unless the projector logic is idempotent. 5. **The write-side outbox or CDC pipeline itself crashes mid-batch**, and a read model can end up with a partially-applied event, inconsistent even within its own fields. ## Where it shows up A concrete, widely recognizable example: an e-commerce catalog where product edits go through a write-side service backed by a relational database, and a Debezium-style CDC pipeline streams row changes through Kafka into an Elasticsearch index that powers site search. A seller updates a product's price, and the write commits instantly in the source database — but the change only appears in search results once: 1. the CDC connector picks it up, 2. Kafka delivers it, 3. the Elasticsearch indexer applies it — typically a few hundred milliseconds to a few seconds later under normal load. Engineers on such systems commonly expose an 'index freshness' or 'connector lag' dashboard specifically to keep this window bounded and alertable, because customers and internal support staff will otherwise report it as a bug rather than recognize it as the expected, monitored trade-off CQRS makes by design.

  • What determines whether the lag window is milliseconds or minutes?
    Mainly the transport used - an in-process synchronous event handler is near-instant, while an external queue with a polling projector adds real delay. Projector throughput, batching interval, and any backpressure under load also directly control the size of the window.
  • Is eventual consistency the same as the read model being simply unreliable?
    No - eventual consistency is a guarantee that the read model will converge to the correct value given no further writes, whereas 'unreliable' implies no guarantee at all. The distinction matters: you can design around a bounded, monitored lag window, but you can't design around an undefined consistency model.
  • How would you actually observe the lag window in a running system?
    Emit a timestamp when the write-side event is produced and another when the projector applies it to the read model, then track the difference as a latency metric (ideally a p50/p99 histogram). A synthetic canary write, probed repeatedly against the read API until it appears, gives an outside-in measurement of what users actually experience.

Think of a company's live sales dashboard fed from the checkout database via a streaming pipeline: the moment a sale closes it's already true in the ledger, but the dashboard tile only ticks up once the pipeline processes that event a few seconds later - same real-world fact, two clocks.

saying these in an interview costs you the question

  • says CQRS always requires two separate physical databases
  • claims read-model updates happen instantly by default
  • conflates eventual consistency with no consistency guarantee at all
  • can't name any mechanism (events/CDC/outbox) connecting the write and read sides
  • thinks the lag window is a fixed constant regardless of load or deployment

context

open as a page

In a CQRS front-end that shows an 'optimistic UI' after a user submits a command - for example, liking a post updates the like count instantly, before the read model confirms it - what has to happen if the command later fails validation, or the eventual read-model update disagrees with what the UI guessed?

level: middleimportance: must knowfreq 70%

basics

~20 s

The UI has to notice the mismatch and fix itself - either roll back to the old value and show an error, or quietly swap in the real value once the read model catches up, so the screen never keeps showing a fact that turned out to be wrong.

open as a page

A user submits a profile update through a CQRS-based app, then is redirected to a 'view profile' page that queries the read model. Name two concrete techniques for making sure that page shows the just-saved change even though the read model may not have caught up yet, and explain how each works.

level: middleimportance: must knowfreq 75%

basics

~20 s

Two options: (1) right after saving, read straight from the write side (or a cache of it) for that user's own data instead of the lagging read model; (2) have the client remember a version number from its last write and make the read model wait until it has caught up to at least that version before answering.

open as a page

An order-placement flow in a CQRS/event-driven system needs to reserve inventory, charge payment, and update a read-model 'order status' projection, each owned by a different service with its own write model. No single ACID transaction spans all three. What coordination pattern keeps this consistent, and what happens if the payment step fails after inventory has already been reserved?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Use a saga: a sequence of local transactions where each step tells the next one to proceed, coordinated either by a central orchestrator or by services reacting to each other's events. If payment fails after inventory was reserved, the saga runs a compensating action - releasing the reserved inventory - to undo that earlier step.

open as a page

A team is about to apply CQRS with an eventually-consistent read model to a feature that checks a user's account balance before authorizing a withdrawal. As the architect reviewing the design, what makes this a case where you'd push back and require strong (immediate) consistency instead, and what are the concrete ways to get it without abandoning CQRS entirely?

level: principalimportance: must knowfreq 60%

basics

~20 s

Money is a case where showing the wrong, stale number can cause real harm, letting someone overdraw because the read side hadn't caught up yet. For that one check, read the current true value directly instead of trusting the lagging copy, even if the rest of the app still uses the fast, eventually-consistent read model.

open as a page

You're on call for a CQRS system where a projector consumes a Kafka topic to update a read-model database. Support reports that some users see stale search results for several minutes after editing a listing. What would you measure to confirm and bound the lag window, and what are two concrete mitigations if the backlog keeps growing under load?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Measure the gap between when an event was written and when the projector actually applied it (processing lag), plus how many unprocessed events are queued up (offset lag). If the backlog keeps growing, either process events faster with more parallel consumers, or make each projector write cheaper, for example by batching, so it can keep pace.

open as a page