skip to content

How can a meta-framework's cached copies be addressed for invalidation, and what does each scheme require at read time?

level: middleimportance: must knowfreq 64%

answer

  1. you can only purge what you can name
  2. path, label, or nothing
  3. labels are attached when stored
  4. no label means the clock only
  5. name the data, not the page

basics

~20 s

You can only invalidate what you can name. Copies are addressed by the route path that produced them, by labels attached when the value was stored, or by nothing, leaving elapsed time as the only instrument.

solid answer

~50 s

Invalidation is an addressing problem before it is a freshness problem. Addressing **by path** targets the stored output of a URL: cheap to reason about, but the write path must enumerate every affected path, including parameterised ones, and it knows nothing about *which* data a page happened to read. Addressing **by label** targets every copy whose production recorded that label — the read says `this result depends on the products collection`, and a later purge of `products` finds all of them, including pages added after the purge code was written. The catch is that labels are attached at store time, so you cannot label retroactively: a value stored without one can never be targeted by it. If neither is available, the copy is **unnamed**, and the only instrument left is its freshness window. That is why unlabelled caching and long windows are a bad pairing.

go deeper

for a junior

Remember the three ways a stored copy can be addressed — by its route path, by labels recorded when it was stored, or not at all — and that an unnamed copy can only be outlived by its freshness window.

for a middle

Explain why labelling cannot be retroactive, why a path names a page but never the data behind it, and why a hand-written list of paths in the write handler drifts as soon as a new page reads the same value.

for a senior

Demonstrate that you decide the addressing scheme when you write the read, keep label names data-shaped, and can audit which cached values in an existing system are currently reachable by anything other than the clock.

for a principal

Own the naming scheme as a cross-team convention: labels are a shared vocabulary between writers and readers, and without a rule for how they are formed you get overlapping names, unreachable copies and purge code that nobody dares change.

## The addressing problem Every invalidation feature in every meta-framework answers the same question: **which stored copies do I mean?** A purge is a query over the cache, and a query needs something to match on. So the design question is not `how do I make this fresh` but `what handle did I leave on this copy when I stored it`. There are three answers in practice, and they are not interchangeable. ## 1. By path The copy is the rendered output of a route, and you name the route: purge the stored output for `/products/42`. - **What it requires at read time:** nothing. The path is inherent to the artifact. - **What it requires at write time:** that the write path can *enumerate* every affected path. For a single detail page that is easy. For the list page, the category pages, a feed, a sitemap and a navigation menu baked into a shared layout, it is a hand-maintained list that drifts the moment someone adds a new page reading the same data. - **The blind spot:** a path says which page, never which data. Two routes reading the same record look unrelated, and the same route may have thousands of parameter values of which you only know the ones you thought of. ## 2. By label attached at read time When a value is read or a route is rendered, the code records one or more labels describing **the data it depended on** — the entity, the collection, the tenant. The cache stores those labels alongside the copy. Later, purging a label invalidates every copy that recorded it. - **What it requires at read time:** the label must be attached *then*. This is the rule that catches people: labelling is not retroactive, so a value already stored without a label is unreachable by that label forever, no matter what you purge. - **What it buys:** the mapping from data to pages is built by the readers, not maintained by the writer. A page added next quarter that reads the same collection inherits reachability with no change to the write path — which is the whole reason labels exist. - **The discipline:** label the **data**, not the page. A label named after a screen reproduces the path scheme's drift with extra steps. ## 3. By nothing — time as the only instrument A copy can be stored with no path handle you can use and no labels. It is not addressable, so the only thing that ends it is its freshness window, or a blunt wholesale discard of the store. | scheme | handle | write path must know | typical failure | |---|---|---|---| | path | the URL of the artifact | every affected route, by hand | a page nobody listed stays stale | | label | names recorded when stored | which data changed | a read that forgot to attach the label | | none | — | nothing | staleness equals the window, always | ## The key is not the handle A related idea worth separating: a cached read also has a **key**, usually derived from its inputs, which decides *who may be served this copy*. The key is about matching requests to copies; the handle is about finding copies to destroy. A value can be keyed precisely and still be impossible to invalidate, because nobody attached anything you can purge by. Conflating the two is a common interview stumble. ## Choosing, in practice 1. Default to labels for anything derived from mutable domain data, and name them after the data. 2. Keep path addressing for genuinely page-shaped invalidation — a route whose content is the page itself, such as a single editorial document. 3. Give every copy a window regardless, sized to the staleness you could defend if the signal never arrives. 4. When you find yourself writing a list of paths in the write handler, treat that as the smell that a label belongs on the read instead. ## The mental model Invalidation reach is decided **at read time, not at write time**. By the time the write happens, the set of copies you can still name is already fixed by the handles the reads left behind — which is why `we will figure out purging later` quietly means `we will restart the process or wait for the clock`.

  • Why is the reach of an invalidation decided at read time rather than at write time?
    Because the handle is attached when the copy is stored. A purge can only match on the path of the artifact or on labels the read recorded, so by the time the write runs, the addressable set is already fixed. Adding labelling code later helps only copies stored after that deployment; everything already in the cache remains reachable solely by its window.
  • What is the difference between a cached read's key and the handle you invalidate it by?
    The key decides which later requests may be served that copy, and is usually derived from the call's inputs. The handle is what a purge matches on — the artifact's path, or labels recorded at store time. A value can be perfectly keyed and still unreachable by any purge, because nothing was ever attached to name it.
  • Why should labels describe the data rather than the page they appear on?
    Because the write path knows what data changed and usually not where it is displayed. A data-shaped label lets new pages that read the same value inherit reachability automatically, while a page-shaped label puts the writer back in the business of enumerating screens, which is the exact drift that path addressing already suffers from.

saying these in an interview costs you the question

  • Thinks a label can be applied to copies already stored
  • Assumes purging a path also reaches other pages using that data
  • Names labels after screens rather than after the data
  • Confuses the cache key with the handle a purge matches on
  • Believes every framework can enumerate parameterised paths for you
  • Caches without labels and then chooses a very long window