A caller updates a key on the primary of an in-memory store, then reads it from a replica and sees the previous value — why?
answer
- two machines, one of them behind
- the copy answers from what it has
- your own write, read elsewhere
- lag wider than your round trip
- pin that read, or tolerate it
basics
~20 sA replica answers from its own copy, which trails the primary by the propagation lag, so a read issued inside that window returns the value the key held before the write. The write is not lost, only not yet visible there.
solid answer
~50 sWrites land on the **primary**; a **replica** serves reads from its own copy of the whole keyspace, which the primary feeds. Between the moment the primary accepts a write and the moment a given copy applies it there is a propagation lag, and a read routed to that copy inside the window answers with the previous value. Nothing malfunctioned — the read simply went to a machine that had not been told yet. Acknowledgment does not settle it either: it proves the primary accepted the write, and where the store acknowledges only after at least one copy holds it, the copy your read happens to reach may not be that one. The fixes are to send that caller's next reads of that key to the primary for a bounded interval, or to decide the read tolerates being behind and say how far behind is acceptable.
go deeper
Recall that a replica is a copy that trails the primary, and that reading your own just-written key from it can hand back the old value. Say clearly that the write still exists.
Explain the window: the primary accepted the write, and the copy had not applied it when the read arrived. Say why being told yes does not prove that any particular copy is current.
Show how you would confirm this in production — correlate the read path's routing with measured propagation lag — and how you would pin the affected read without pinning everything the service reads.
Frame it as policy: which read classes may be answered by a copy at all, what staleness budget each carries, and who reviews a new read that claims it needs the primary.
## The two machines and the gap between them A **primary** is the node that accepts writes for a key. A **replica** is a copy of the whole keyspace kept on another node and fed by the primary; it can answer reads, which is the entire reason for having one. The delay between the primary accepting a write and a particular copy applying it is the **propagation lag**. How copies are fed, and what widens that lag, is a subject of its own; for this question only two properties of it matter: - it is not zero, because the write has to travel and be applied; - it is not constant, because it moves with write volume, network conditions and what else the machines are doing. A read answered by a copy therefore comes from a snapshot that is some amount of time old — usually smaller than anything the caller notices, and occasionally larger than the caller's own round trip, which is when this shows up. ## The failure, staged 1. The caller sends a write for a key to the primary. The primary applies it and tells the caller yes. 2. The caller immediately issues a read of the same key. Routing sends that read to a copy, because that is what the fan-out was built to do. 3. The copy has not applied the write yet. It answers, correctly and quickly, with the value the key held before. 4. A few milliseconds later the copy applies the write, and the identical read would now return the new value. Nothing malfunctioned. Each component did exactly its job, and the bug is in the assumption that "the store told me yes" and "every machine in the store can tell me so" are the same statement. This reaches you as "the write didn't save", and it is almost never about a lost write. ## Why acknowledgment does not settle it | Acknowledgment posture | What the caller knows when told yes | Can a read from a copy still be old? | |---|---|---| | Acknowledge-then-propagate | The primary holds the write | Yes — no copy is promised to hold it yet | | Wait-for-a-copy acknowledgment | The primary and at least one copy hold it | Yes — if the read lands on any other copy | | Wait for every copy | All copies hold it | No for that key, but the write path now stalls on the slowest copy and on any unreachable one | The honest statement is that stores in this class differ here, and the difference is exactly what decides the answer. Some acknowledge before any copy has the write; some acknowledge only once a copy does; some offer the choice per call rather than fixing it for the whole deployment; and some stores of this kind ship no replication at all, so the question does not arise until something outside the store copies the data. Where reads are routed by an intervening proxy rather than by the caller, the caller may not even learn which node answered it. ## What is actually at risk - **Advisory reads** — a count rendered on a page, a value read again in a second. Being a little behind is cosmetic, and tolerating it is why the copies earn their keep. - **The caller's own read-back** — the common case, because the same code path frequently writes and then reads what it wrote. The caller is the one party guaranteed to know the new value exists. - **Reads that decide something** — where the answer is used on the spot to decide whether to proceed, an old answer is not old, it is wrong. Classifying those is a larger subject than this one. One framing matters here. This staleness is the tier's own copy of the tier's own state being behind. There is nothing to invalidate, no other system to ask, and no refresh to trigger: the copy will catch up on its own, and the only decision is whether the caller can wait. ## The remedies, and what each costs - **Pin the caller's follow-up read.** After a caller writes a key, route its reads of that key to the primary for a bounded interval sized against the lag you actually measure. Cheap, narrow, and it works. - **Pin the read class.** Where the read's correctness depends on being current regardless of who issues it, route the whole class to the primary rather than chasing individual callers. - **Bound the staleness.** Measure the propagation lag continuously, name a budget, and decide what the read path does when the budget is exceeded. - **Accept it.** For most reads on most tiers this is the right answer, and it is also the only one that keeps the capacity the copies were bought for. Every pin hands back some of the read capacity the fan-out provided, which is the trade this whole arrangement is made of. ## What this is not It is not a lost write — that is a primary failing before its work reached any copy, a different failure with a different remedy. It is not a formal consistency model either: naming the session guarantees that describe it is a distributed-systems subject. And it is not cache invalidation, since nothing here holds a copy of another system's data. ## A protocol-neutral timeline of one read-your-own-write failure: the caller's read reaches the copy seven milliseconds before the write does. The numbers are illustrative — the window is whatever the propagation lag happens to be at that moment, and it is neither zero nor constant ``` t0 caller -> primary : put(user:42, "b") (key currently holds "a") t0+1ms primary -> caller : accepted t0+2ms caller -> replica : get(user:42) t0+2ms replica -> caller : "a" <- the previous value t0+9ms primary -> replica : write propagated t0+9ms replica applies it; the same read would now answer "b" ```
- If the store acknowledges a write only once at least one copy holds it, can this still happen?Yes. That posture proves one copy held the write, not that the copy your read reached is that one; with several copies, a read routed to any other can still trail. It narrows the window rather than closing it, and some stores let the posture be chosen per call rather than fixing it for the deployment.
- What is the cheapest fix for the one caller that just wrote?Route that caller's next reads of that key to the primary for a bounded interval, sized against measured propagation lag. How the pin is expressed varies: some callers choose per call, some per connection, and where a proxy decides routing the caller may not be able to express it at all. Each pin returns some of the read capacity the copies were added for.
You file a change at head office and are told it is recorded, then phone a branch that still works from this morning's printed bulletin. The branch is not wrong and your change is not lost — the bulletin it reads from has not been reprinted yet.
saying these in an interview costs you the question
- Says the write was lost, when it is only not yet visible there.
- Assumes every copy is current the instant the caller is acknowledged.
- Calls it a cache invalidation problem rather than a copy that is behind.
- Believes waiting for one copy to hold a write makes all copies current.
- Claims re-reading the same copy immediately always returns the new value.