skip to content

How would you set and verify a maximum-staleness budget for a data-access layer's cache, and why is hit rate the wrong headline number?

level: principalimportance: should knowfreq 42%

answer

  1. hit rate measures savings, not truth
  2. it improves when invalidation breaks
  3. consequence of acting on the old value
  4. the weakest link sets the bound
  5. probe convergence, chart the tail

basics

~20 s

Derive a per-dataset maximum age from the consequence of acting on an old value, then verify it with convergence probes and divergence sampling. Hit rate measures work avoided and rises when invalidation breaks, so it hides the risk.

solid answer

~50 s

Hit rate reports how much database work the cache avoided; it says nothing about whether the values were true, and it actually improves when invalidation stops working, because nothing is being removed. The number worth publishing is the worst-case time between a row changing and the last copy of it disappearing. Derive it per dataset from the consequence of acting on a stale value: anything feeding an irreversible action gets no cache, a screen the user just edited needs a very small bound, reference data tolerates a large one. The achievable bound is set by the weakest link — commit-to-removal delay, propagation between instances, removals that never arrive, writes that bypass the layer — and where one can be missed outright it falls back to the entry's maximum age. Then verify by measurement: convergence probes, divergence sampling, and invalidations issued versus applied.

go deeper

for a junior

Learn the distinction first: hit rate says how much work was saved, staleness says how wrong the answer can be. They are different questions and only the second one is about correctness.

for a middle

Be able to say where staleness comes from: the delay from commit to removal, propagation between instances, invalidations that never arrive, and writes that bypass the layer — and that the last two fall back to the entry's maximum age.

for a senior

Show the verification, not the intent. Convergence probes charted at the maximum, divergence sampling against fresh reads, and counters for invalidations issued versus applied after each deployment.

for a principal

Own the two decisions configuration cannot make: which datasets are ineligible for caching at all, and who is accountable for the budget when the writers belong to another team.

## Hit rate answers the wrong question Hit rate is the number every cache reports and the number most teams put on the dashboard. It measures **work avoided**: the fraction of lookups that did not become a database read. It says nothing whatsoever about whether the values handed back were still true. A cache whose invalidation has been broken since the last deployment reports a *better* hit rate than a correct one, because nothing is being removed. As a headline it is actively misleading: it goes up as correctness goes down. The number that describes the risk is the **maximum staleness**: the worst-case time between a row changing and the last cached copy of it disappearing. That is the quantity a reader of your API cares about, the quantity an incident review will ask for, and the quantity you can actually design against. ## Where the budget comes from A staleness budget is not chosen from a menu; it is derived from what happens if someone acts on the old value. Work per dataset, not per cache: - What decision is made from this data? Displaying a description tolerates minutes. Deciding whether a request is authorised, or whether an account has funds, tolerates nothing — that data belongs on a live read. - Who notices, and how fast? A user who has just edited a value and is looking at the result notices instantly; a nightly report notices never. - Is the wrongness recoverable? A late-updating page corrects itself. A message sent, a payment taken, or an entitlement granted on an old value does not. - What does the write rate imply? For rarely-written data almost any budget is safe. For hot rows, the budget effectively decides how much of the cache is wrong at any moment. The output is a stated number per dataset — "at most N seconds" — and a decision for anything that cannot name one: do not cache it. ## Where the delay comes from The achievable bound is set by the weakest link in the invalidation path, not by the intent of the design. Wherever a removal can be missed outright, the bound falls back to the entry's maximum age: | Contribution | Typical size | What removes it | |---|---|---| | Delay between commit and the removal running | small, unless it is tied to method exit rather than commit | tie the hook to the commit | | Propagation to other instances | network and queue delay | one shared copy instead of per-instance copies | | Removals that never arrive at all | unbounded until something else clears the entry | a bounded maximum entry age as the backstop | | Writes that bypass the layer entirely | unbounded | ownership of the write path, or a change signal | The last two rows dominate most real systems, which is why a budget expressed as a bounded entry lifetime is usually honest, and one expressed only as an invalidation usually is not. ## Verifying the number instead of asserting it A staleness budget nobody measures is a comment. Three measurements make it real: 1. **Convergence probes.** Write a marker row on a schedule, then read it back through the normal path on every instance, and record the time until all of them return the new value. Chart the maximum, not the mean — the tail is the budget. 2. **Divergence sampling.** Compare a random sample of served values against fresh reads and alert on both the rate of mismatches and the age of the worst one found. 3. **Invalidation flow counters.** Count removals issued and removals applied per instance. A ratio that changes after a deployment is the earliest warning that the design has quietly stopped working. Alert on the tail of the convergence probe crossing the stated budget, not on hit rate moving. ## The judgment a lead actually owns Two decisions cannot be delegated to configuration. The first is which datasets are **ineligible** — the ones where an old value produces an irreversible action — because that decision has to survive people later trying to speed those paths up. The second is who owns the budget when the writers live in another team: a cache whose correctness depends on another team's write path is a coupling that must be written down, given a signal, or given up. Absent that, the budget silently becomes "however long until someone redeploys", which is not a number anyone can defend after an incident.

  • Why can hit rate improve when invalidation breaks?
    Because removals are what create misses. If the code that removes entries stops running, entries live longer, more lookups hit, and the chart moves in the direction everyone reads as good. That inversion is the reason it is a poor headline: the metric rewards exactly the failure you care about.
  • What does a convergence probe actually measure?
    End-to-end staleness. Write a marker row, then read it back through the normal serving path on every instance and record the time until all of them return the new value. It captures the whole path — hook placement, propagation, delivery losses — rather than the entry lifetime you configured, and the maximum is the number to chart.
  • Which datasets should be declared ineligible for caching outright?
    Those where acting on an old value is irreversible or unsafe: authorisation and entitlement decisions, balances used to approve something, anything that triggers a payment or a message. The decision has to be written down, because later performance work will otherwise reintroduce a cache on exactly those paths.
  • How does the ownership of the writers change the budget?
    If another team can write the rows, your bound is whatever their signal gives you plus your backstop age — and with no signal it is just the backstop. That coupling has to be documented and agreed, or the data should not be cached; otherwise the real budget quietly becomes 'until someone redeploys'.

saying these in an interview costs you the question

  • Puts hit rate on the dashboard as the cache's health metric
  • Assumes an invalidation that was coded is an invalidation that arrives
  • Sets one global staleness number for every dataset
  • Caches authorisation decisions for convenience
  • Alerts on averages rather than the convergence tail
  • Never measures convergence, only reasons about the design