One write changes a value that many routes read. How do you determine what to invalidate, and how do over-wide and under-wide labels fail differently?
answer
- one write, many consumers
- who owns the mapping, writer or reader
- readers declare, writers purge by name
- too wide costs money, too narrow costs truth
- creates and deletes fan out wider
basics
~20 sLet the readers declare what they depended on, so one write purges by data name and reaches every consumer. Over-wide labels fail loudly as wasted work and origin load; under-wide labels fail silently as content nobody notices is stale.
solid answer
~50 sA single write usually touches far more than the obvious detail page: the list it appears on, category pages, a feed, a count rendered into a shared layout. There are two ways to know that set. You can maintain it in the write path as a list of routes, accurate only until someone adds another reader. Or each read records a label for the **data** it consumed, so the cache holds the mapping and one purge fans out to every copy that recorded it. Label width is the real knob, and the two directions fail asymmetrically: **too wide** means an unrelated write empties far more than it should — costly, visible in hit ratio and origin load. **Too narrow** means some consumer never recorded the label and quietly keeps serving the old value. Bias wide, narrow with evidence.
go deeper
Understand that changing one record can make many pages wrong — the item page, the lists it appears on, counts and feeds — and that each of those stored copies has to be reached somehow.
Explain why letting each read declare the data it depended on beats keeping a list of routes in the write handler, and what a label's width covers.
Show the asymmetry explicitly: a wide label wastes work in ways your graphs already show, while a narrow one leaves content quietly wrong, so you default wide and narrow only when you can measure the cost.
Treat label width as a budget across teams: agree the vocabulary, decide who may introduce a wide label, and put measurement in place for the silent direction, because no individual reviewer can see a dependency nobody declared.
## The fan-out is bigger than the page you edited Editing one record rarely invalidates one artifact. A realistic set for a single product change: - the product's own detail route; - every list, category and search result that included it; - any aggregate rendered from it — a count, a price range, a `latest` strip in a shared layout; - machine-readable outputs such as a feed or a sitemap; - stored results of the underlying data reads, separately from the rendered routes above. The engineering question is how the write path comes to know that set, and it has exactly two shapes. ## Two ways to know the set **Writer-maintained mapping.** The handler that performs the write names the routes to purge. It is explicit and easy to read, and it is correct only as long as somebody updates it every time a new page reads that data. Nobody does, because the person adding the new page is not looking at the write handler. **Reader-declared dependencies.** Each read or render records a label naming the data it consumed. The cache accumulates the mapping as a side effect of normal traffic, and the write purges by data name. A page added later inherits reachability with no change to the write path. This inverts responsibility onto the party that actually knows the dependency — the reader — which is why it scales. ## Label width, and why the two errors are not symmetrical Width is how much a single label covers: one entity, a collection, a whole tenant, everything. | | over-wide | under-wide | |---|---|---| | what goes wrong | unrelated writes invalidate unrelated copies | some consumer never recorded the label | | user-visible symptom | slower responses after every write; nothing incorrect | two surfaces disagree; one shows old data indefinitely | | who finds it | your dashboards — hit ratio drops, origin load rises | a user, a support ticket, or nobody | | how it scales | badly under write bursts, when the cache is repeatedly emptied | badly with team size, as new readers accumulate | | correctness | preserved | violated | That asymmetry drives the default. An over-wide label costs money and latency and announces itself in graphs you already have. An under-wide label costs correctness and announces nothing at all: the page looks fine, it is simply wrong, and it stays wrong until the window elapses — which, in a system that leans on signals, may be a very long time. ## Practical width rules 1. **Attach more than one label.** A read can record both the entity it fetched and the collection it came from; the write then purges the narrow name for an update and the wide one for a create or delete. 2. **Creates and deletes are wider than updates.** Editing an existing item leaves the lists intact in shape; adding or removing one changes every list, count and page boundary that contained it. 3. **Narrow only with evidence.** Move from a collection-wide label to per-entity labels when you can show the wide purge is hurting — not because it feels imprecise. 4. **Watch for the hot wide label.** One label attached to almost every copy turns any write into a full flush, and the worst moment for that is a bulk import, when writes arrive in the thousands. 5. **Keep a window underneath.** The window bounds how long an under-wide mistake can survive, which is the only cheap protection against the failure you cannot see. ## Diagnosing the silent direction Under-wide failures do not appear in cache metrics, because from the cache's point of view nothing went wrong. The signals that do work are comparative: a canary value written through the real write path and then polled on every surface that should show it; a scheduled job that compares a rendered aggregate against the source; and the support pattern of `page A shows the new price, page B does not`, which is a dependency the readers never declared rather than a purge that failed. ## The mental model Pick the failure you can see. Both width errors are certain to happen in a large codebase; only one of them tells you it happened.
- Why does a create or a delete usually need a wider purge than an update?An update changes an item that every list already contained, so the shape of those pages is unchanged. A create or delete changes membership: lists gain or lose a row, counts move, pagination boundaries shift and aggregates change. Copies that never recorded the new entity's own name still have to go, which is what a collection-level label is for.
- A bulk import fires thousands of writes. What does that do to a wide label?Each write purges everything the label covers, so the cache is emptied repeatedly while the import runs and the origin absorbs the full traffic with nothing stored to answer from. Treat imports as a distinct path: coalesce the signals into one purge at the end, or suppress per-row invalidation and rely on a single wider one.
- How would you detect that a label is too narrow, given the cache reports nothing wrong?Compare surfaces rather than watching the cache. Write a canary value through the real write path and poll every route that should reflect it, recording how long each takes; alert on any that never converge. Support reports of the form one page shows the new value and another does not are the same signal arriving late.
saying these in an interview costs you the question
- Keeps the list of affected routes in the write handler
- Treats a narrow label as safer than a wide one
- Assumes purging the detail page also fixes the lists
- Ignores that creates and deletes change list membership
- Fires one purge per row during a bulk import
- Expects cache metrics to reveal a missed dependency