skip to content

A user submits a form that appends an event to the log, then the client immediately navigates to a page that reads from a projection built off that log - and the change is missing. What is actually happening here, and what are the common ways to handle it?

level: seniorimportance: must knowfreq 70%

answer

  1. async gap between write and projection catch-up
  2. return position, wait for projection to reach it
  3. optimistic UI hides the lag
  4. bypass projection: read aggregate directly
  5. sticky routing across replicas

basics

~20 s

The projection hasn't caught up to the newest event yet, so the read is briefly stale - like refreshing a scoreboard a split second before it updates. Common fixes: make the client wait for the update, read the write directly instead of from the projection, or show a temporary version on screen immediately.

solid answer

~40 s

This is the read-your-writes problem, a direct consequence of eventual consistency between the write side (the event log) and the read side (an asynchronously updated projection): there's an unavoidable window between an event being appended and the projection reflecting it. Common mitigations: return the write's resulting position or version to the client and have the next read either poll until the projection's position is at least that version, or block server-side until the projection catches up before responding; use optimistic UI updates so the client renders the expected result immediately without waiting on the projection at all; or, for cases that truly need strong consistency, read from the aggregate or write side directly for that one request.

go deeper

for a junior

Should recognize this as 'the update needs a moment to show up' rather than concluding data was lost.

for a middle

Should name the pattern (eventual consistency / read-your-writes) and describe at least one concrete mitigation like optimistic UI or polling.

for a senior

Should compare multiple mitigations (position tracking, optimistic UI, bypass-read, sticky routing) with their trade-offs and pick appropriately per use case.

for a principal

Should reason about this as a systemic UX and reliability concern across a whole product - setting conventions for which flows get read-your-writes guarantees versus which tolerate visible lag, and how that's communicated to the frontend team as a contract.

## What is actually happening **Read-your-writes** is the specific consistency guarantee people expect by default - 'after I successfully save something, if I immediately look at it, I should see my own change' - and it is exactly what asynchronous, log-derived projections do not give you for free. The mechanism producing the gap is straightforward: a write appends an event to the log and returns success as soon as that append is durable; the projection, however, learns about that event only when its subscriber next reads and processes it, which happens on its own schedule - milliseconds later under light load, longer under a backlog. Between those two moments, any read against the projection reflects the state before the new event, even though the write has already succeeded from the client's point of view. This isn't a bug; it's the direct cost of decoupling reads from writes, which is the entire reason projections are fast and scalable in the first place. ## Why the trade-off is accepted The reason this trade-off is accepted is that strict, always-consistent reads would mean either: - not decoupling the read model from the write model at all (losing the query-shape and scaling benefits), or - synchronously updating every projection inside the same transaction as the write (which reintroduces tight coupling and doesn't scale to multiple projections or multiple read replicas). Accepting a small inconsistency window in exchange for a decoupled, horizontally scalable, purpose-built read side is the deliberate bet event sourcing with projections makes. ## The mitigations Several concrete techniques mitigate the read-your-writes problem without abandoning the async projection architecture. - One is **version or position tracking**: the write endpoint returns the event's resulting stream position (or a monotonic version number) to the client, and the subsequent read either polls the projection until its checkpoint is at or past that position, or the server blocks the read request momentarily until the projection catches up, bounded by a timeout to avoid hanging forever under heavy lag. - A second is **optimistic UI**: the client renders the expected post-write state immediately from the data it just submitted, without waiting on the projection at all, and reconciles silently if the eventual server read differs - this is the most common approach in consumer-facing apps because it hides the lag entirely from the user's perceived experience. - A third, used sparingly, is **bypassing the projection** for the specific read that needs strong consistency: reading the aggregate directly from the write-side event stream (replaying just that one aggregate's events, which is cheap because a single aggregate's history is small) instead of the shared projection, trading a bit of extra latency on that one call for correctness. - A fourth is **session or sticky consistency**: routing a given user's subsequent reads to the same projection replica they just wrote through, common when multiple projection replicas exist behind a load balancer and different replicas may lag by different amounts. ## Failure modes Failure modes in production tend to cluster around forgetting this gap exists rather than not knowing the mitigations. 1. The most common is a UI that shows **a jarring flicker** - the item is missing for a moment then pops in - because no mitigation was applied and the projection lag happened to be visible under load, which erodes user trust even though nothing was actually lost. 2. A second is choosing a **naive fixed 'sleep 200ms before redirecting' hack**, which works in testing under low load but breaks under production traffic spikes when actual lag exceeds the hardcoded delay, producing intermittent, hard-to-reproduce bug reports. 3. A third, more serious failure is silently reading from **a stale replica** in a way that looks like data loss - support tickets say 'my order disappeared' when actually a replica behind a load balancer just hasn't caught up, and without version tracking or sticky routing, it's hard to distinguish real data loss from ordinary catch-up lag during incident triage. 4. A fourth is **over-correcting**: making every read block on projection catch-up 'just to be safe,' which quietly reintroduces the very latency and scaling costs projections exist to avoid. ## A real-world pattern A concrete, widely used real-world pattern: many e-commerce checkout flows use exactly the optimistic-UI approach - after a successful 'place order' call, the confirmation page renders directly from the data the client just submitted rather than re-querying the orders read model, while the `orders_view` projection catches up in the background for the order-history page the user will see later, by which time the lag is imperceptible.

  • Why not just make every projection update synchronous with the write, inside the same transaction, to avoid this problem entirely?
    That reintroduces the tight coupling and scaling limits event sourcing with async projections is meant to escape - it ties every write to the availability and latency of every downstream projection, doesn't scale cleanly to multiple projections or read replicas, and defeats the purpose of decoupling read shape from write shape. It trades away the architecture's main benefit to fix an edge case that has cheaper mitigations.
  • What's the risk of using a hardcoded delay like 'wait 200ms then redirect' instead of tracking projection position explicitly?
    It works only as long as actual projection lag stays under the hardcoded value, which is true in light testing but breaks under real production load spikes when lag legitimately exceeds 200ms, producing intermittent failures that are hard to reproduce and debug. Explicit position tracking, polling or blocking until the projection reaches a known event version, is robust to variable lag instead of gambling on a fixed number.
  • How does sticky or session routing help with read-your-writes when there are multiple projection replicas behind a load balancer?
    Different replicas can lag by different amounts, so a write followed by a read that happens to land on a slower replica can appear to have been lost even though a different, more caught-up replica has it. Routing a user's writes and their subsequent reads to the same replica, or routing to a replica confirmed at or past the write's version, avoids that specific inconsistency.

Like mailing a letter and then immediately calling the post office to ask if it arrived - the letter is genuinely sent and on its way, but the tracking system hasn't updated yet, so the answer 'not yet delivered' doesn't mean the letter was lost.

saying these in an interview costs you the question

  • treats a missing just-written item as data loss rather than lag
  • proposes making all projection updates synchronous as the default fix
  • suggests a fixed sleep or delay as a robust solution
  • doesn't mention returning a version or position from the write to correlate with the read
  • assumes all read replicas of a projection are always equally caught up

context