A cloud service uses CQRS with a read store that's updated asynchronously from an event-sourced write side. A user updates their profile, immediately reloads the page, and still sees the old value for a couple of seconds. What's causing this, and how would you address it for a production user experience?
answer
- projection lag, not a bug
- read-your-writes routing
- optimistic UI reconciliation
- version/sequence token wait
- monitor lag as a metric
basics
~20 sThe read copy of the data hasn't caught up to the write yet — there's a short delay while the update flows through to it. Fixes include showing the user their own change right away without waiting for the read copy, or making that specific read wait until it's caught up.
solid answer
~50 sThis is ordinary projection lag — the write succeeded and was appended to the event log, but the read-model projector hasn't yet consumed and applied that event to the read store the page is querying. It's the expected cost of eventual consistency in CQRS with async projections, not a bug. Standard mitigations: (1) read-your-writes routing — after a write, route that user's immediate follow-up read to the write side or a synchronously-updated path instead of the async projection; (2) optimistic UI — apply the change client-side immediately without waiting for a server read, reconciling silently if the eventual server value differs; (3) version/sequence-token consistency — pass a version from the write and have the read wait until the replica reaches at least that version; (4) reduce and monitor projection lag itself so the window stays imperceptible. The choice depends on how strict the correctness requirement is for that specific field.
go deeper
Should be able to say the read copy just hasn't caught up yet and that this is a normal, expected delay rather than data loss.
Should name at least one concrete mitigation (e.g., showing the change immediately in the UI) and know it's a form of eventual consistency.
Should compare multiple mitigation strategies (read-your-writes routing, optimistic UI, version tokens) and their tradeoffs, and mention lag monitoring as a production practice.
Should weigh which mitigation fits which class of data/consistency requirement across a whole platform, and factor in how growing lag becomes a systemic incident, not just a single-user cosmetic issue.
## What is actually happening In CQRS with event sourcing, a write doesn't update the read store directly — it's appended as an event to the write-side log, and a separate asynchronous projector consumes that event and updates the read-optimized store the UI actually queries. The gap between 'event appended' and 'read model updated' is **projection lag**, and it is the direct, expected mechanism behind the scenario described: the profile update succeeded and is durably recorded, but the specific read the reloaded page issued hit a read model that hadn't processed that event yet. This isn't a bug in the traditional sense — it's the eventual-consistency tradeoff the whole pattern makes deliberately, in exchange for the read/write scaling and shaping benefits CQRS provides. What matters in practice is whether the lag window is short enough, and handled gracefully enough, that users don't notice or aren't harmed by it. ## The mitigations There are several established mitigation strategies, and choosing among them is a judgment call based on how strict the correctness requirement is for the specific piece of data. 1. **Read-your-writes routing.** The simplest and most common is read-your-writes routing: for a short window right after a user's own write, route their reads for that specific data back to the write side (or a path known to be synchronously consistent with it) instead of the async read model, then fall back to the normal projected read model afterward. This guarantees the user always sees their own change immediately without requiring the entire read path to become synchronous. 2. **Optimistic UI.** A second approach is optimistic UI: the client applies the change locally as soon as the write request succeeds, without waiting for any subsequent read to confirm it, and only reconciles if a later read genuinely disagrees, which should be rare and usually indicates a real failure, not just lag. 3. **Version/sequence-token consistency.** A third is version/sequence-token consistency: the write response includes a version or event-sequence number, and a subsequent read is required to wait until the read replica's applied version reaches at least that number before answering — this gives an explicit, boundable consistency guarantee rather than an unbounded 'eventually.' 4. **Shrinking the lag itself.** A fourth, complementary approach is simply shrinking and monitoring the lag itself — tracking each projector's consumer lag as a first-class production metric, alerting when it exceeds a threshold, and treating persistent lag growth as an incident, since a growing gap makes every one of the above mitigations less effective or more visibly broken. ## The tradeoff: complexity versus strictness The tradeoff across these options is complexity versus strictness. Read-your-writes routing and version tokens both require plumbing — the write path needs to expose a version, and the read path needs logic to either target the write side or wait for a specific version, which adds coupling between the two models that CQRS otherwise tries to avoid. Optimistic UI is simpler to implement but pushes the burden onto client code and creates a usually-rare reconciliation edge case if the eventual value diverges from what was optimistically shown — for example if a concurrent write from another session or device changed the same field in between. None of these options make the system strongly consistent end-to-end; they narrow where and when a user is exposed to staleness, which is usually the actually useful goal rather than eliminating eventual consistency altogether, which would defeat much of the reason CQRS was adopted. ## Failure modes in production Failure modes in production tend to cluster around lag growing unexpectedly: - a projector falling behind under load; - a downstream dependency of the projector throttling writes; - a deployment briefly pausing a projector's consumption. When lag grows past what any read-your-writes or version-token mechanism was designed to tolerate, users start seeing genuinely broken-looking behavior — not just their own recent change missing, but potentially other users' visible state looking stale too, which is a much bigger deal than the single-user 'did my own save work' case in the original scenario. ## Where it shows up A concrete real-world shape: a profile-editing feature in a consumer app writes the update through the command side, immediately re-renders the new value optimistically in the UI so the user sees their own change instantly, while the background projector updates the durable, search-indexed profile read model used by other parts of the app (e.g., another user viewing this profile) — that other-user view is allowed to lag by a second or two, monitored via projector-lag alerting, but the editing user's own view never depends on that projection catching up at all.
- Why not just make every read synchronous with the write to avoid this entirely?That would eliminate the independent-scaling and denormalized-shaping benefits CQRS exists to provide — the read path would be coupled back to the write path's consistency and throughput characteristics for every consumer, not just the user who just wrote, which reintroduces the exact bottleneck CQRS is meant to remove.
- How would you detect this problem before users report it?Instrument each projector's consumer lag (how far its cursor trails the head of the event log) as a metric, alert when it crosses a threshold tuned to what's user-perceptible, and correlate lag spikes with deploys or load events so regressions are caught quickly rather than discovered via user complaints.
- What's the risk with pure optimistic UI and no reconciliation check at all?If the optimistic value silently diverges from what the server eventually persists — due to a rejected write, a conflicting concurrent update, or a bug — and the client never checks back, the user can be shown a confidently wrong state indefinitely, which is worse than a brief, honest loading delay.
Like mailing a change-of-address form to the post office and then immediately checking an old printed phone directory — the post office (write side) already has your update, but the directory (read model) is only reprinted periodically, so it briefly shows the old address even though the real record is already correct.
saying these in an interview costs you the question
- Calls this a bug rather than expected eventual consistency
- Only solution offered is 'make everything strongly consistent'
- Doesn't distinguish the writing user's own read-your-writes need from other users' staleness tolerance
- No mention of monitoring/alerting on projection lag
- Assumes optimistic UI needs no reconciliation with the eventual server state