skip to content

When an invalidation cannot be guaranteed to reach every stored copy, how do you decide which copies to trust it for and which to bound with a short window?

level: principalimportance: should knowfreq 42%

answer

  1. partial reach, not total staleness
  2. two truths at once
  3. can I name it, is it acknowledged
  4. the weakest copy sets the promise
  5. measure time to converge per surface

basics

~10 s

Classify content by the cost of being stale, then per copy ask whether you can name it, whether the purge is acknowledged, and what the worst case is. Elsewhere the window is the promise.

solid answer

~40 s

Partial reach is the failure that defines this area: a purge lands on some copies and not others, so the site serves two truths at once — one route new, another old, two users disagreeing. That is worse than being uniformly a minute behind, because it is irreproducible and destroys trust in what the reader is looking at. Decide per class of content, not per cache. Name the maximum staleness each class may carry, and for each copy ask three questions: can I address it, do I get an acknowledgement that the purge applied, and what is the worst case if the signal is lost? Where all three answers are good, a signal may carry the guarantee. Where any is not, the window **is** the guarantee.

go deeper

for a junior

Understand that one value can exist as several stored copies, and that an update reaching some of them is why two pages in the same site can disagree.

for a middle

Explain the three questions to ask of each copy — can it be addressed, is the purge acknowledged, what happens if the signal is lost — and why the weakest of those answers sets the real guarantee.

for a senior

Demonstrate measurement: a canary written through the real path, time-to-convergence recorded per surface, and alerting on surfaces that only converge when their window elapses.

for a principal

Own the promise per content class, align windows so surfaces agree rather than racing each one to be freshest, name the values that must not be cached at all, and make sure somebody is paged when convergence breaks.

## Two truths at once The characteristic incident in this area is not `everything is stale`. It is **partial reach**: the invalidation ran, it succeeded, and it landed on some of the copies of the changed value. The result is a site that contradicts itself — a detail page showing the new price while a listing shows the old one, a first load disagreeing with what a visitor sees after clicking around, one visitor seeing the update and their colleague not. This is worse than uniform staleness for reasons that are about people rather than systems: - it is **irreproducible** — whoever investigates hits a copy that happens to be correct; - it destroys the reader's model of the product, because no explanation they can form covers what they saw; - it makes every subsequent report untrustworthy, since nobody can say which surface was which. So the design goal is not maximum freshness. It is **consistency of the promise across surfaces**, at whatever level of freshness you chose. ## The three questions, asked per copy For each place a copy of a value can end up, ask: 1. **Can I address it?** Is there a handle — a path, a label — that a purge can match, or is the copy anonymous? 2. **Do I learn that the purge applied?** A local store answers synchronously. A remote one may accept the request and apply it later. Something in front of your origin may or may not report success. A payload already delivered to a visitor's open tab cannot be reached at all; only a fresh request replaces it. 3. **What is the worst case if the signal is lost?** If the answer is `stale until the window ends`, that window is what you are actually promising. Any copy that fails question 1 or 2 must not be carrying a freshness guarantee based on signals, no matter how reliable the signal appears in testing. ## Turning that into a policy | content class | tolerable staleness | who carries the promise | |---|---|---| | money, entitlements, permissions | none | do not cache the value; derive it per request | | inventory, availability, status | seconds to a minute | a short window; treat signals as a bonus | | editorial and catalogue content | minutes | signals, with a window as the backstop | | marketing pages, documentation | hours | a long window; signals for editor satisfaction | Two rules make the table hold in practice. First, **align the surfaces**: if one copy of a value can only be bounded by a two-minute window, giving another copy a one-second window buys nothing except the disagreement between them — a uniformly two-minute-old site is better than a self-contradictory one. Second, **the weakest link sets the promise**; the guarantee you may state publicly is the worst reachable copy's worst case, not the best one's. ## Verify reach, or you do not have it An invalidation strategy nobody measures is a habit, not a mechanism. The cheap instrument is an end-to-end canary: 1. write a recognisable value through the real write path; 2. poll every surface that should reflect it — server-rendered output, an in-app navigation, the machine-readable feed, a request that goes through whatever sits in front of the origin; 3. record time-to-convergence per surface; 4. alert on any surface that fails to converge inside its stated window. The output is the number this whole discussion is about: not `did the purge return success` but **how long each surface actually takes to tell the truth**. It also catches the two failures nobody else catches — a surface that never converges (a dependency nobody declared) and a surface that converges only because its window elapsed (a signal that has been silently broken for months). ## The judgment calls a lead owns - **Where to spend.** Extending reach to an additional tier costs engineering time and operational risk. Shortening a window costs capacity. Price both against the cost of staleness for that content class, and be willing to say that hours-old marketing copy is fine. - **Uniform versus tiered windows.** Uniform is easier to explain and keeps surfaces in agreement; tiered wins real capacity on high-traffic routes. Choose deliberately, and write it down. - **What is simply not cacheable.** The mature answer to `this must never be stale` is not a cleverer invalidation scheme; it is a value that is computed per request. Deciding which values those are is an architectural call, not an optimisation. - **Who is paged.** When propagation breaks, the failure is silent. Somebody must own the convergence metric, or the first report will come from a customer comparing two pages. ## The mental model Invalidation makes the common case fast; the window is what you promise. Say the promise out loud per class of content, check that every surface can keep it, and prefer a site that is consistently slightly behind over one that is fresh in places.

  • Why can a shorter window on one surface make the overall experience worse?
    Because disagreement between surfaces is the damaging symptom, not age. If one copy can only be bounded at two minutes, refreshing another every second guarantees a two-minute span in which the two contradict each other. Aligning the windows produces a site that is uniformly slightly behind, which users can understand and support can reproduce.
  • What is the honest answer to a requirement that a value must never appear stale?
    Do not cache it. Money, entitlements and permission checks should be derived per request, because every caching scheme trades some staleness for speed and no amount of invalidation reduces that to zero. Reserve the cache for values whose cost of being briefly wrong you can state, and spend the capacity you saved on the ones you cannot.
  • What single metric best tells you whether your invalidation actually works?
    Time-to-convergence per surface, from a canary written through the real write path. A purge returning success proves only that a request was accepted. Convergence time proves every surface eventually tells the truth, exposes surfaces that only converge when their window elapses, and turns a stated guarantee into something you can alert on.

A recall notice that reaches three of five warehouses is worse than one that reaches none: stock that was never shipped is at least uniformly wrong, while a half-delivered notice means two branches confidently contradict the other three.

saying these in an interview costs you the question

  • Treats a successful purge call as proof of reach
  • Optimises freshness per surface and accepts disagreement between them
  • Promises zero staleness for a cached value
  • Has no measurement of how long propagation actually takes
  • Assumes a payload already in an open tab can be reached
  • Drops freshness windows once invalidation signals exist