skip to content

In a CQRS system where the write side uses Event Sourcing, a user submits a command that changes data, and the client immediately re-queries a read model to display the result — but the change is missing or shows stale data. Why does this happen structurally, and what are the common ways teams handle it?

level: seniorimportance: must knowfreq 70%

answer

  1. read-your-own-writes problem
  2. command path and query path have no shared transaction
  3. version-aware polling using stream version
  4. optimistic UI from command response
  5. route to strongly-consistent source as fallback

basics

~20 s

Saving the change (writing the event) and updating what the screen shows (the read model) are two separate steps done by two separate pieces of code, and the second one takes a little time to catch up after the first. If you check the screen data too fast, you can catch it before it's updated.

solid answer

~50 s

This is the classic 'read-your-own-writes' problem, and it's structural, not a bug: the command handler's job ends the moment it durably appends an event; a separate, asynchronous projector then has to notice that event and update the read model, and that hop always takes some non-zero time, even if it's usually milliseconds. Common mitigations: have the command response include the new event's stream version and have the client (or an API gateway) poll or block the next query until the read model's tracked version is at least that high; update the UI optimistically from the command's own response data rather than re-querying immediately; or route the specific just-written entity's next read to a strongly-consistent source (the aggregate itself, or a synchronously-updated cache) for a short window. None of these eliminate the lag system-wide — they just hide it for the one user who just acted.

go deeper

for a junior

Should recognize that writes and reads happen through separate paths and that a short delay before the read side catches up is expected, not a bug.

for a middle

Should name at least one concrete mitigation, like not re-querying and instead using the command's own response to update the UI.

for a senior

Should describe multiple mitigations (version-aware polling, optimistic UI, routing to a strongly consistent source) and know when each is appropriate.

for a principal

Should be able to define a system-wide consistency contract for read models, set SLOs for projector lag, and design monitoring/alerting so this failure mode is caught before it becomes a user-facing incident at scale.

## Why the stale read is structural This is usually called the **read-your-own-writes** problem, and in a CQRS+Event-Sourcing system it isn't a bug to be fixed so much as a direct, unavoidable consequence of how the two sides are wired together. The command path and the query path are, by design, two separate pipelines with no shared transaction between them. - When a command handler processes a request, its entire job — reconstituting the aggregate, running the business logic, and appending the resulting event(s) to the event store — ends the instant that append durably succeeds. Nothing in that call path waits for, or even knows about, any read model. - Separately, and asynchronously, a projector process is subscribed to the event stream; at some point after the append, it notices the new event, applies it, and updates the read store the UI actually queries. The gap between 'event durably appended' and 'projection updated' is real wall-clock time — usually milliseconds under healthy conditions, but it can stretch to seconds or longer under load, projector lag, or an outage — and any read that lands inside that gap sees the old state. ## How this differs from replica lag This differs meaningfully from the read-your-own-writes problem in, say, a simple primary/replica relational database, where the same underlying failure mode (replication lag) exists but is often narrower and more bounded, because the primary and replica run near-identical logic over near-identical data. In CQRS+ES, the read model can be an entirely different shape from the event stream — sometimes involving nontrivial aggregation, joins across multiple aggregate streams, or external lookups — so the lag isn't just network replication delay, it's real processing time, and it can vary a lot depending on what the specific projection has to do with that event. ## The three mitigations 1. The most direct mitigation is **version-aware polling**: the command handler's response includes the stream version (or a global sequence position) of the event it just appended, and the client — or, better, a thin layer in the API — then either polls the read model's own tracked 'last processed version' until it's caught up to at least that number, or the query endpoint itself blocks briefly until the projection has caught up before returning. This gives a precise, per-request consistency guarantee without making every read wait on every write system-wide — only the specific request that just wrote has to wait for its own change to land. 2. A second, cheaper mitigation is **optimistic UI update**: the client doesn't re-query the read model at all after a successful command; it locally merges the change it already knows it just made (from the command's own request/response) into what it displays, and lets the background projection catch up invisibly for everyone else. This sidesteps the lag entirely for the acting user's immediate experience, at the cost of the client needing to know how to render its own optimistic state. 3. A third approach, used when the UI genuinely needs to read the freshest possible state right after writing, is to **route that specific follow-up read to a strongly consistent source** instead of the eventually-consistent projection — reconstituting the aggregate directly (the same path the command handler used) or reading from a synchronously-updated cache keyed by aggregate ID — accepting the cost of bypassing the purpose-built read model for that one call. ## Failure modes that compound it The failure modes that show up in production tend to compound this problem rather than being separate from it. - **Projector lag under load** is the big one: if event volume spikes (a batch import, a traffic surge) faster than a projection can keep up, the read-your-writes window stretches from milliseconds to potentially minutes, and a mitigation designed around 'usually fast enough' polling can start timing out or degrading the UX badly right when the system is under the most stress. - **A second failure mode is inconsistent expectations across a team**: without an explicit, agreed-upon consistency contract for a given read model, some engineers write client code assuming synchronous consistency (breaking under normal async lag) while others over-defensively poll everywhere (adding needless latency to reads that never needed it), and the inconsistency itself becomes a source of subtle bugs and support tickets. ## How checkout screens sidestep the lag A concrete, well-known real-world instance of this pattern: e-commerce checkout flows commonly show an 'order confirmed' screen driven entirely from the command's own response payload (order ID, items, total) rather than by immediately re-querying an 'order history' read model, precisely so the confirmation page can render instantly and correctly regardless of how far behind the order-history projection happens to be at that moment — the read model catches up in the background, and by the time the user navigates to 'my orders,' it almost always has.

  • Why doesn't the command handler just update the read model itself, synchronously, to avoid this problem entirely?
    Doing that would couple the write path to every read model's schema and failure mode, defeating the core purpose of separating them in the first place — a slow or failing projection would then also slow down or fail commands, and you'd lose the ability to have many independently-scaled read models built asynchronously off the same event stream.
  • Is version-aware polling free, latency-wise, for the acting user?
    No — it trades read-your-writes correctness for a small added wait on that specific follow-up read, since the client (or gateway) has to block or retry until the projection's tracked version catches up. It's a deliberate, scoped cost taken only by the request that needs strong consistency, not a system-wide slowdown.
  • How would you detect, in production monitoring, that this lag has become a real problem rather than a theoretical one?
    Track projector lag directly — the difference between the latest event position in the stream and the position each projection has processed up to — and alert when that gap exceeds a threshold, rather than trying to infer the problem indirectly from user complaints. Correlating lag spikes with event-volume spikes usually points at the root cause quickly.

It's like mailing a letter and then immediately calling the recipient to ask if they got it yet — the letter is genuinely sent and on its way (the event is durably appended), but the delivery (the projection catching up) hasn't happened yet, so asking too soon just gets you 'not yet.'

saying these in an interview costs you the question

  • Calls this a bug rather than a structural consequence of the write/read separation
  • Suggests fixing it by having the command handler synchronously update the read model
  • Doesn't distinguish between per-user mitigations (optimistic UI, version polling) and the underlying system-wide lag, which never fully disappears
  • Has no idea how to monitor or detect projector lag in production

context