A request writes a row, then a read-only unit routed to a replica reads it back stale - what causes this and what fixes it?
answer
- asynchronous replication, per-unit routing
- the second unit knows nothing of the first
- request-scoped pin to the primary
- cache keyed by datasource, or it leaks stale
basics
~20 sReplication is asynchronous, and each unit is routed on its own declared attributes with no memory that an earlier unit in the request wrote. The fix is a request-scoped pin sending later units to the primary.
solid answer
~40 sThe write committed on the primary; the following read-only unit was routed to a replica that had not applied it yet, so the user sees their own change missing. Nothing is broken - routing looks only at what the unit declares, and each unit is decided independently. The standard remedy inside the layer is a **pin**: request-scoped state set when a unit writes, which the routing layer reads at every later boundary in that request and uses to force the primary. It costs primary read load for the tail of that request, usually a good trade. Two related traps: the pin must survive whatever carries the request, so async continuations need it propagated explicitly, and any identity map or shared cache must be keyed by the datasource it was loaded from.
go deeper
Recall that replicas receive changes after the primary commits, so a read sent to a replica right after a write can miss it. The row is not lost; that server is simply behind.
Explain why per-unit routing produces this - each boundary is decided on its own declared attributes - and describe pinning the rest of the request to the primary as the standard remedy.
Show the diagnosis: check the primary to separate lag from a lost write, log which datasource served each unit, correlate with replication delay, and check whether a cache is answering ahead of the query. Then discuss pinning and its cost.
Make freshness an explicit per-endpoint contract rather than a side effect of a repository marking, and decide where the pin lives, how it crosses async boundaries, and what the system does when replication delay exceeds the budget.
## What actually happened Three independent facts combine into the bug: 1. **Replication is asynchronous.** A commit on the primary returns before replicas have applied it. The window is usually milliseconds and occasionally seconds, and it widens exactly when the system is busy. 2. **Routing is per unit of work.** The layer chooses a datasource from what the unit declares at its boundary. A read-only unit declares that it will not write; it does not declare anything about how fresh its data must be. 3. **Units in one request do not know about each other.** The second unit has no idea the first one wrote the very row it is about to read. So the write lands on the primary, the read is served from a replica that is a few hundred milliseconds behind, and the response shows the pre-write value. The user, who has just pressed save, sees their change vanish. Nothing errored, and every component did its job. ## Why this is the layer's problem, not just the database's A replica being behind is expected and normal. What turns it into a defect is the routing rule: **read-only was treated as a synonym for stale-tolerant**. Those are different properties. A list of last month's invoices is both; a re-read of the row the user just edited is read-only but absolutely not stale-tolerant. ## Pinning: the request-scoped fix The mechanism that data-access layers use is a pin held for the life of the request: 1. A unit of work commits a write. 2. The layer sets a **request-scoped flag** - conceptually "this request has written". 3. At every later boundary in that request, the routing layer sees the flag and returns a primary connection, whatever the unit's own marking says. 4. The flag dies with the request. It is deliberately coarse. It does not track which rows were written, so it over-pins: an unrelated read later in the same request also goes to the primary. That is the point - the cheap version is correct and the precise version needs to know what the read will touch, which the layer does not. | approach | freshness it buys | what it costs | |---|---|---| | pin the rest of the request to the primary | own writes always visible | primary carries the tail of the request | | pin a user for a short window after a write | own writes visible across requests too | needs shared state and a window guess | | wait on a replica for a recorded replication position | fresh reads on the replica itself | added latency, and a fallback when it lags | | route this one read explicitly to the primary | precise | someone must know which reads need it | ## The traps around the pin - **It must travel with the request.** If the pin lives in state attached to the handling thread and the work continues on another, it is gone and the read is silently routed away again. Anything asynchronous needs it propagated deliberately. - **Background work has no request.** A job triggered by a write has no request scope to carry a pin; it either reads from the primary by default or tolerates lag by design. - **The pin cannot fix cross-request reads.** The next request from the same user starts clean. If own-write visibility must survive that, the window has to be held somewhere shared, keyed by user, with all the expiry questions that follow. - **Caches and identity maps are the second door.** A copy loaded on one datasource must not answer a read routed to another; if a shared cache is not keyed by datasource, pinning to the primary buys nothing because the stale copy is served without any query at all. The mirror case also bites: a cache populated from a replica keeps serving the pre-write value after the pin has already sent queries to the primary. ## Diagnosing it in production The signature is a read-your-own-write failure that is intermittent, load-correlated, and impossible to reproduce on a single-node development setup. Useful evidence: - replication delay metrics correlated with the complaint timestamps; - which datasource served the failing read, which is worth logging per unit; - whether the value appears correct on a retry seconds later, which distinguishes lag from a genuinely lost write; - whether the row is right in the database but wrong in the response, which points at a cache rather than at routing. That last check matters, because a dropped write inside a mis-marked read-only unit produces the same user complaint with a completely different cause: there, the row on the primary is wrong too. ## The judgement to state out loud Pinning trades replica offload for correctness on the one path where users notice most. Take it by default and remove it only where an endpoint has an explicit, agreed tolerance for staleness. Freshness is a property of the endpoint's contract, not of the repository method it happens to call.
- How do you tell this apart from a write that was silently dropped?Look at the row on the primary. With replication lag the primary holds the new value and only the replica is behind, so a retry moments later succeeds. With a write dropped inside a mis-marked read-only unit the primary is wrong too, and it stays wrong forever. Same user complaint, opposite cause.
- Why not make the read wait until the replica has caught up?You can, where the layer can record a replication position at commit and have the read wait for the replica to reach it. It keeps the read off the primary, which is the whole point of routing, but it adds latency to that read and needs a timeout with a fallback to the primary when lag is large. Pinning is coarser and simpler; waiting is finer and more machinery.
- What breaks the pin most often in real systems?Losing the request scope. Work handed to another thread, a queued continuation, or a scheduled job carries no pin, so a read that logically belongs to the same interaction is routed to a replica again. The fix is to propagate the flag explicitly across those hops, or to declare such work primary-only.
saying these in an interview costs you the question
- Blames the write, when the primary actually holds the new value
- Assumes replicas are synchronous unless someone configured otherwise
- Fixes it by retrying the read and hoping the replica caught up
- Pins per request but never propagates the pin across async hops
- Forgets a shared cache can serve the stale copy even after pinning