An in-memory store holding both re-derivable values and sole-copy records serves all reads from replicas; which reads must return to the primary, and how do you bound the rest?
answer
- classify the read, not the caller
- advisory answers tolerate being behind
- presence or absence, right now
- budget the staleness, then measure it
- pinning costs the capacity you bought
basics
~20 sReads whose correctness is the presence or absence of a key at this instant cannot be answered by a lagging copy: route that class to the primary. Every other read gets a named staleness budget, measured against propagation lag rather than assumed.
solid answer
~50 sClassify reads by tolerance, not by caller. A read whose answer is advisory — rendered, counted for display, or re-derivable from a system of record — is fine on a copy that is slightly behind, provided you say how far behind is acceptable. A read whose correctness *is* the presence or absence of a key right now is not: asking whether another caller currently holds a claim, whether an allowance is exhausted, or whether a piece of work was already recorded as done turns propagation lag into a wrong decision rather than an old value. Route that class to the **primary**, pinned by class rather than by caller, because a class is reviewable and a caller is not. For the rest, bound the staleness: measure propagation lag continuously, name the budget, and decide what the read path does when the measurement exceeds it.
go deeper
Know that a copy can answer with an older value, so a read that decides something on the spot should be sent to the primary instead of a copy.
Separate reads whose answer is advisory from reads whose answer is a decision, and say which of the two tolerates a copy that is behind.
Pin by read class, defend the staleness budget with measured propagation lag at the tail, and state what the read path does when the measurement exceeds it.
Own the classification as policy: what may be answered from a copy by default, who approves a new pinned class, and what those pins cost in capacity.
## The classification is the answer Once reads are fanned out across copies, every read is answered from a snapshot that is some amount of time old. The engineering question is not "is this acceptable?" in the abstract; it is "which reads is it acceptable for, and how do I know?" The useful axis is what the caller does with the answer. | Read class | What an old answer costs | Where it should go | |---|---|---| | Advisory — rendered, displayed, counted for a human | A slightly old screen | A copy, with a named staleness budget | | Re-derivable — the value has another home | Extra work, a slower path | A copy; the system of record settles ties | | Decision on presence or absence, right now | A wrong decision, not an old one | The primary | | The caller's own read-back after its own write | Confusing behaviour, usually a bug report | The primary, for a bounded interval | The third row is the one worth naming carefully, because it is a whole family: a check on whether someone else currently holds a claim on a resource, a decision about whether an allowance has been spent, a lookup asking whether a unit of work has already been recorded as done. In each of them the answer is the *existence* of a key at this moment, and existence is precisely what the lag moves. Those reads do not get a little bit wrong as the lag grows; they get the decision backwards. How each of those records should be designed is a workload question in its own right — what belongs here is only the routing consequence. ## Why the state's other home matters The same lag produces an annoyance or an incident depending on a fact about the keys, not about the tier: whether anything else can reproduce the value. - If the value is derived and another system holds the truth, an old read costs recomputation or a slightly stale render. - If the tier holds the only copy of the record — a claim, an allowance, a note that something already happened — nobody can correct the answer, because there is nothing to correct it against. So the classification has to be applied per key family, not per tier. A single store routinely holds both kinds, which is what makes "just point reads at the copies" such a durable source of incidents. ## Pin the class, not the caller Pinning individual callers feels cheaper and ends badly. A caller is a deployment artefact: it gets split, renamed, replaced by two services, and the reasoning that justified the pin does not travel with it. The alternative is durable: - A read class — "is this claim held?", "is this allowance spent?" — is a property of the decision, and stays true wherever the code lives. - A class is **reviewable**: someone can ask whether it genuinely needs the primary, and the list is short enough to argue about. - A list of pinned callers only ever grows, because nobody who did not write the code dares unpin it. ## Bounding the staleness, honestly For everything else, the discipline is to convert a hope into a measurement: 1. **Measure propagation lag continuously** for the copies you route reads to, as time behind or as un-propagated work. 2. **Name a budget** per read class — the amount of staleness that class tolerates. Derive it from what the answer is for, not from what the tier currently achieves. 3. **Compare the tail, not the median.** The defensible bound is what the distribution does at the high percentiles, because those are the reads that generate the reports. 4. **Decide the behaviour on breach.** When the measured lag exceeds the budget, the read path must do something — route that class to the primary, or fail the read — because an unenforced budget is an intention. Exactly how lag is exposed differs across stores in this class, which is a reason to build the measurement into the read path's own view of the world rather than depending on a particular reporting surface. ## What bounding cannot do A staleness budget changes how *often* a read is behind, not whether it can be. That is fine when the cost is a stale render and useless when the cost is a wrong decision: a key that appeared or vanished within the budget is still missed. Which is why decision reads are routed rather than budgeted — the two remedies are not interchangeable and mixing them up is the most common error here. ## The bill Every read sent back to the primary is a read the copies do not serve, so it hands back part of the capacity the fan-out was bought for. That is not an argument against pinning; it is the reason to classify precisely rather than defensively, and it is the number to put in front of anyone who proposes more copies as a fix. ## Two independent reasons to route a read to the primary — a read class whose answer is a decision, and a bounded pin after this caller's own write. The pin lookup reads a default so an unwritten key takes the copy branch. This shape is expressible only where the caller chooses the route; behind an intervening proxy the same policy has to live in the proxy instead ``` # caller-side routing decision on write(key, value): primary.put(key, value) pinned_until[key] = now() + pin_window on read(key, read_class): if read_class is "decides-on-presence": return primary.get(key) # never a copy if now() < pinned_until.get(key, default=0): return primary.get(key) # this caller's own write return some_replica.get(key) # advisory: budgeted ```
- Why pin a read class rather than the caller that issues it?A caller is a deployment artefact that gets split, renamed and replaced, and the reasoning behind its pin does not travel with it. A read class is a property of the decision being made, so it stays true wherever the code lives, and it is short enough to review. Pinning callers ends with everything pinned, because nobody dares unpin someone else's.
- How would you prove a staleness bound rather than assert one?Measure propagation lag continuously and defend the tail of the distribution rather than its median, because the tail is what generates the reports. Then define the behaviour on breach — route that class to the primary, or fail the read. How lag is exposed differs across stores, so build the measurement into the read path's own view.
- Does bounding the staleness help a read whose answer is presence or absence?Rarely. An answer that is old by less than the budget can still be wrong, because the key may have appeared or vanished inside the window. Bounding changes how often that happens, not whether it can, so those reads are routed to the primary rather than budgeted.
saying these in an interview costs you the question
- Pins whole services to the primary instead of classifying reads.
- Assumes a stale read is always just a slightly old value.
- Names a staleness budget without ever measuring propagation lag.
- Thinks re-reading a copy until it agrees proves the value is current.
- Treats a lag-driven wrong decision as a cache freshness problem.