A user saves their profile, the application immediately re-reads it from an asynchronous read replica, and the old values come back. Explain why this happens and what the standard fixes are.
answer
- commit ack precedes replica apply
- read-your-writes = per-session guarantee
- return what you wrote; do not re-read
- sticky primary window after write
- LSN/GTID token instead of a guessed window
basics
~20 sThe primary acknowledged the commit before the replica applied it, so the follow-up read hit a copy that does not yet contain the change. Fixes: read that user's data from the primary for a short window after their write, or make the read wait until the replica has applied that write's position.
solid answer
~50 sThis is the classic read-your-writes violation. With asynchronous replication the primary confirms the commit as soon as it is durable locally, then ships the change; the replica applies it a moment later. A read issued in that window sees the pre-write state, so the user is told the save succeeded and then shown the old value. Fixes, cheapest first. Render the response from what you just wrote instead of re-reading. Pin the session to the primary for a short window after a write, typically by stamping the session or a cookie with a last-write timestamp and routing that user's reads to the primary until it expires. More precisely, capture the replication position of the commit (a PostgreSQL LSN, a MySQL GTID) and either wait on the replica until it has applied that position or route to the primary if it has not. The anti-patterns are a fixed sleep before reading, and simply sending every read to the primary, which throws away the replica.
go deeper
Say the commit was acknowledged before the replica applied it, and give one fix: read the just-written data from the primary, or return it from the write itself.
Name the read-your-writes property, describe the sticky-primary window and how you would choose it, and mention the position-token alternative.
Compare the fixes on blast radius and cost, discuss window sizing versus worst-case lag, cross-device and background-job cases, and timeouts and fallbacks for a wait-based approach.
Decide where the guarantee belongs: per-endpoint routing policy versus changing the commit contract for the whole cluster, and what that choice costs in write latency and availability.
## The mechanism With asynchronous replication the sequence is: the primary applies your change and makes it durable in its own log, it answers COMMIT to the client, and only afterwards does the change record travel to the replica and get applied there. Between the second and third step, the primary and the replica disagree. If your application answers the user's next request by reading a replica, it reads a copy of the database that predates the write it just confirmed. The user experiences the worst kind of bug: the system said saved and then showed the old data, which reads as data loss even though nothing was lost. The property being violated has a name: read-your-writes (also called read-after-write consistency). It is a per-session guarantee: a session that performed a write must not subsequently observe a state older than that write. Note it is weaker than requiring globally fresh reads. Other users may keep seeing the old profile for a moment; that is usually acceptable. It is the author of the change who must never see it disappear. ## Fix 1: do not re-read at all The cheapest fix is to stop asking the database. The write path already knows the new values, so return them (or the row returned by the write statement) in the response, and update the client-side state from that. This removes the round trip entirely. It fails when the write triggers derived data that only the database computes, or when the next page is a different request that the client cannot pre-populate. ## Fix 2: sticky primary reads after a write Stamp the session with the moment of its last write, for example in the session store or in a cookie the router can see. While now minus that stamp is less than a chosen window, route that session's reads to the primary; afterwards, let them go back to replicas. The window must exceed your realistic worst-case lag, which is why teams usually pick a few seconds rather than a few hundred milliseconds. Strengths: simple, needs no replication internals, works with any driver. Weaknesses: it is a guess. Too short and the bug survives under lag spikes; too long and a write-heavy user permanently pins their traffic to the primary, undoing the offload. It also only helps the writer's own session; a different device or a background job for the same user is not covered unless you key the stamp by user rather than by session. ## Fix 3: wait for the write's replication position The precise version. After committing, ask the primary for the position of that commit in the change stream: an LSN in PostgreSQL, a GTID in MySQL. Carry that token in the session or the request. On the read side, either compare it to the replica's applied position and choose the primary if the replica is behind, or ask the replica to block until it has applied that position. This gives exactly the guarantee you want with no arbitrary window, and it lets reads keep using replicas as soon as they have caught up. The cost is plumbing: the token must be threaded through the application, and the wait needs a timeout and a fallback to the primary. ## Fix 4: change the durability contract You can also make the commit itself wait until a replica has applied the change (synchronous replication at the apply level). That removes the anomaly for every reader of that replica, but it puts replica latency and replica availability into every write. It is a heavier, system-wide answer to what is usually a per-endpoint problem. ## Anti-patterns Sleeping for a fixed delay before reading is a race dressed as a fix; it slows every user and still breaks under a lag spike. Retrying the read until the value looks new is worse, because the application cannot distinguish stale from legitimately changed by someone else. Routing all reads to the primary solves it by deleting the feature you were paying for. Finally, do not confuse this with caching: an application cache in front of the read path creates the same symptom for the same reason, and needs its own invalidation on write. ## How to present it Name the anomaly, state the cause in one sentence (commit acknowledged before apply), then give a ladder of fixes with their trade-offs, and note that the guarantee needed is per-session, not global freshness.
- Why is a fixed sleep before the read a bad fix?It does not synchronise with anything: it slows every request by the sleep, and under a lag spike the replica is still behind when the sleep expires, so the bug reappears exactly when the system is stressed. Correct fixes either avoid the read, route it to a node known to have the write, or wait on the actual replication position.
- Does read-your-writes for one user guarantee other users see the write too?No. It is a per-session property. Another user reading a lagging replica may still see the old value for a while, which is normally acceptable. If a second actor must see the first actor's write, that is a causality requirement between sessions and needs the position-token approach, passing the token along with the request, or reading from the primary.
saying these in an interview costs you the question
- Blaming caching or the ORM instead of naming replication lag
- Adding a sleep or a retry loop before the read
- Sending all reads to the primary as the only fix
- Believing synchronous shipping of the log alone fixes it, when the replica still has to apply the change before a read sees it
- Pinning to the primary forever after any write, so replicas end up idle