A user submits a profile update through a CQRS-based app, then is redirected to a 'view profile' page that queries the read model. Name two concrete techniques for making sure that page shows the just-saved change even though the read model may not have caught up yet, and explain how each works.
answer
- session/sticky read to write side
- version or offset token echoed on next read
- write-through cache
- own-write guarantee vs everyone-sees-it-immediately
- Cosmos DB session consistency, DynamoDB ConsistentRead
basics
~20 sTwo options: (1) right after saving, read straight from the write side (or a cache of it) for that user's own data instead of the lagging read model; (2) have the client remember a version number from its last write and make the read model wait until it has caught up to at least that version before answering.
solid answer
~40 sStandard read-your-writes techniques: (a) session/sticky routing - route the requesting user's subsequent reads for a short window to the write-side database or a synchronously-updated cache, bypassing the async read model entirely, often gated by a 'just wrote' flag or short TTL; (b) version/token-based read-your-writes - the write returns a monotonic version (an LSN, event offset, or ETag) that the client attaches to its next read; the read tier either blocks/polls until its applied-offset reaches that version, or falls back gracefully if it can't. Trade-off: (a) is simple but only helps the writer's own next read and adds read-path branching; (b) generalizes across devices/clients but needs version plumbing and can add wait latency to the read.
go deeper
Should recognize the problem exists and that redirecting to a read page right after a save can show stale data; naming 'refresh again' as a workaround is acceptable but not yet a systemic fix.
Should be able to name and implement at least one of the two techniques - sticky read-through or version token - end to end for a single feature.
Should choose between the techniques based on multi-client/multi-device needs, own the token propagation contract, and reason about the added latency/complexity budget.
Should establish the pattern once as a reusable platform primitive (e.g. a shared session-consistency middleware) so individual feature teams don't reinvent it inconsistently.
## Technique one — sticky, write-side reads The first technique is sticky, write-side reads. Immediately after a command succeeds, the client (or server on the client's behalf) is routed to read that specific data directly from the authoritative write-side store, or from a synchronously-updated write-through cache, rather than the asynchronously-projected read model. This is usually scoped narrowly: only the writer's own next request for their own data is redirected, often via a short-lived flag or a brief TTL window (e.g., 'for the next 5 seconds, this user's profile reads hit the primary DB'). - **It's simple** to reason about and requires no changes to the read-model pipeline itself. - But **it only solves the problem for the writer** — other users viewing that same profile still see the plain eventually-consistent read model. - And **it partially defeats the read-scaling purpose** of CQRS if applied broadly or left on too long. ## Technique two — version- or token-based read-your-writes The second technique is version- or token-based read-your-writes, sometimes called **session consistency** or **causal consistency**. When a command commits, the system returns a monotonic marker identifying how far the write model (or the event stream) has progressed — a log sequence number, a Kafka offset, or an application-level version counter. The client stores this token and attaches it to its next read request. The read tier compares the token against its own 'applied up to' watermark: - if it has already caught up past that version, it answers immediately with fresh data; - if not, it either blocks briefly (bounded-staleness read), polls until it catches up, or explicitly signals 'not yet ready' so the client can retry. This generalizes across devices and sessions far better than sticky routing, because the guarantee travels with the token rather than being pinned to a server or session, but it requires plumbing a version through every write response and every subsequent read call, and it can add real wait latency to reads that arrive right after a write. ## Why the problem exists at all This problem exists specifically because CQRS intentionally decouples the write and read paths for scalability, and that decoupling has a UX and correctness cost that must be handled explicitly rather than assumed away: a system that is technically 'eventually consistent' but shows a user their own action apparently failing (their comment vanished, their setting reverted) erodes trust even though the data is not actually lost, just delayed. Read-your-writes patterns exist to give a narrow, well-scoped consistency guarantee — 'you will always see your own writes' — without paying the cost of making the entire system strongly consistent for every reader. ## How the two compare The trade-offs run in opposite directions for the two techniques. | Approach | Where it wins | What it costs | |---|---|---| | Sticky/direct reads | are cheap to build for a single feature | but don't generalize (a second browser tab, a mobile app reopened later, or another user viewing the same data get no such guarantee) and add branching logic to the read path that has to be remembered and maintained | | Version-token reads | generalize well and can be built as reusable middleware | but require every write endpoint to return a token and every read endpoint to accept and honor one, plus a mechanism (blocking, polling, or a rejected-with-retry response) for what happens when the read side genuinely hasn't caught up — and that wait, however bounded, is still latency the user feels | ## Failure modes Failure modes show up when the scoping breaks down. - **Sticky routing pinned to a server instance** can silently stop working after a failover or load-balancer reroute, quietly reverting to plain eventual consistency without anyone noticing. - **A version token that isn't propagated correctly** through client caching layers (e.g., a mobile app that caches an old token after being backgrounded) can either hang waiting for a version the read side will never specifically 'catch up to' if it expired, or simply be ignored, silently downgrading the guarantee. - **Clock skew or an incorrectly-implemented monotonic counter** can make version comparisons meaningless, causing either false 'not ready' responses on data that's actually current, or false 'ready' responses on data that's actually stale. ## Where it shows up in shipped databases A concrete, well-known real-world instance of this pattern is Azure Cosmos DB's 'session consistency' level, which is built exactly this way: each write returns a session token, the SDK automatically attaches it to subsequent reads in that session, and Cosmos DB guarantees the session will never see data older than what that token represents, while other sessions still get the (cheaper) default consistency level. Amazon DynamoDB offers a related but simpler mechanism — a per-request `ConsistentRead=true` flag that forces a read from the leader replica instead of a possibly-lagging follower, which is the direct-write-side-read technique built into the database itself, at extra read cost.
- Does read-your-writes have to apply globally, or can it be scoped just to the writer?It's normally scoped to the writer or their session - that's the entire point of session consistency. Other clients still see the plain eventually-consistent read model; demanding a global strong-consistency guarantee for every reader is a much bigger and more expensive ask that CQRS is specifically trying to avoid.
- What happens with the version-token approach if the client never sends a token, e.g., on a fresh page load?The read model falls back to normal eventual consistency for that request. Read-your-writes is opt-in per request via the token; without one attached, the server has no way to know it needs to wait for anything, so it just answers with whatever it currently has.
- How would you implement this in a system fronted by DynamoDB specifically?DynamoDB exposes a ConsistentRead=true option on GetItem and Query operations, which reads from the leader replica instead of a potentially-lagging follower - a built-in version of the 'read directly from the write side' technique, at higher read cost and latency, and scoped only to the same table and region.
Like a courier who, right after dropping your package at the sorting hub, hands you a receipt with a batch number - when you check status, the system compares your receipt's batch number to what's already been processed and either shows you the fresh answer or says 'give it a second', while everyone else just sees whatever the public tracking board currently displays.
saying these in an interview costs you the question
- assumes eventual consistency automatically self-resolves in time for the writer's own next read
- proposes making all reads strongly consistent to fix this one case
- doesn't distinguish 'the writer sees their own write' from 'everyone sees the write immediately'
- no mechanism to correlate a specific write with a specific subsequent read (no token, session, or flag)