skip to content

Invalidation & Staleness

How a mapper's cache goes wrong: writes that go around it, two instances holding different copies, entries written before a commit that never happened. The ordering rule is easy to get backwards.

on this pageshow

questions

5

In a data-access layer that caches loaded rows, what makes a cached entry stale, and which writes leave one behind?

level: juniorimportance: must knowfreq 70%

answer

  1. a copy that outlived its row
  2. who issued the write matters
  3. set-based statements hand over no identifiers
  4. writers outside the layer never invalidate
  5. a stale hit looks exactly like a correct one

basics

~20 s

A cached entry goes stale when the row it copies changes and nothing removes the copy. Any write the caching layer never sees can leave one behind: a set-based statement, a native statement, another service, or an operator at a console.

solid answer

~40 s

A cache in a data-access layer stores a copy of a row taken at read time; the copy is only valid while the row is unchanged. An entry goes stale when the row changes and nothing removed or replaced the copy. That happens whenever the write does not pass through the code that maintains the entries: a set-based `UPDATE ... WHERE ...` that never materialises objects and so hands the layer no identifiers, a hand-written native statement, a second service or scheduled job, a data fix run by an operator, or a change arriving underneath the engine such as a restore or a replication stream. The critical property is that a stale hit is indistinguishable from a correct one — no error, no log line, just a plausible value returned quickly.

go deeper

for a junior

Be able to say in one sentence what stale means: the copy outlived the row it was made from. Then name two writes that leave one behind, such as a set-based statement and a write from another service.

for a middle

Explain why the bypass happens mechanically: a set-based statement never materialises objects, so the layer has no identifiers to invalidate and can only evict by table, and only if told which tables are involved.

for a senior

Show how you would find this in production. Staleness has no error signal, so you compare served values against fresh reads and report the worst observed age rather than reading hit-rate charts.

for a principal

Frame it as an ownership question. A dataset with writers outside your process cannot be invalidated at all, so caching it means choosing and publishing a maximum age instead of claiming freshness.

## What a stale entry is A data-access layer that caches rows keeps a **copy** of what a row looked like at the moment it was read — sometimes the raw column values, sometimes a whole materialised object. That copy is correct only while the row it was made from is unchanged. As soon as the row changes and nothing removes or replaces the copy, the layer holds a **stale entry**: a value that was true once and is now false. The property that makes staleness dangerous is not that the value is wrong. It is that a stale hit is **indistinguishable from a correct hit**. There is no exception, no timeout, no retry, no log line. The lookup returns quickly with a plausible object, and the calling code has no signal at all that it is looking at history. ## Which writes leave one behind An entry survives a write only when the write happened somewhere the caching code was not watching. So the question is never "did the row change?" but "did the change pass through the code that maintains the entries?" The bypass routes fall into a small number of shapes: - **A set-based statement.** A single `UPDATE orders SET status = 'CLOSED' WHERE closed_at < ?` changes thousands of rows in the engine without materialising one object. The layer typically has no list of the identifiers the statement matched, so it cannot remove those entries one by one; it either evicts everything associated with the affected tables, or — if it was never told which tables the statement touches — evicts nothing. - **A hand-written native statement.** The same problem, with less information: the layer may not be able to work out which tables are involved without being told. - **Another writer entirely.** A second service, a scheduled job, a one-off data-fix script, a bulk import, an operator typing at a console. None of them run your invalidation code, and none of them can be made to. - **A change that arrives underneath the engine.** A restore from backup, a failover to a replica that had fallen behind, or a replication stream applying rows that were written somewhere else. The row changes without any statement your process ever issued. | Write route | Does the layer learn the identifiers? | Usual effect on entries | |---|---|---| | Object-by-object write through the layer | yes | the entry is removed or replaced correctly | | Set-based statement through the layer | no — at best the table names | broad eviction, or nothing at all | | Native statement | often not even the table names | nothing, unless declared explicitly | | Another service, job, or operator | no | no invalidation of any kind | | Restore, failover, replication apply | no | no invalidation of any kind | ## Why the layer's own map is not a defence Inside a single unit of work, most layers keep an identity map: read the same identifier twice and you are handed the same instance. That hides a class of surprises, but it is **consistency within one unit of work**, not freshness — the instance is as old as the read that created it. A cache whose entries outlive the unit of work that populated them extends the same trick across requests, and extends the staleness with it. The longer an entry can live, the wider the window in which a bypassing write can make it wrong. ## What follows for design Three consequences drive nearly every decision in this area: 1. **Cache only what you can invalidate.** If a dataset has writers outside your reach, the correctness of the cache is bounded by a maximum age you accept deliberately, not by an invalidation you are unable to write. 2. **Declare what a bulk statement touches.** Layers that let you name the affected tables can turn a silent bypass into a coarse eviction. A coarse eviction costs performance; a missed invalidation costs correctness, and the two are not comparable. 3. **Measure age, not hits.** The hit rate tells you what the cache saved. The number that describes the risk is the worst-case age of a value the cache can still serve after its row changed. ## How it shows up in practice The classic incident looks cheap at the time. Someone replaces a loop that loads a million objects with one set-based statement — an unambiguously good change for the database — and a screen starts showing values that are hours old, on some instances, some of the time. Nothing failed, nothing was logged, and the change looks unrelated. It is found by comparing what the application serves with a direct read of the row, which is the only reliable way to observe staleness: the cache will never volunteer it.

  • Why can a set-based statement issued through the same layer still leave stale entries?
    Because it changes rows in the engine without materialising them. The layer never learns which identifiers matched, so it cannot remove those entries individually. At best it can evict everything associated with the tables involved, and only if it has been told which tables the statement touches.
  • Does the layer's identity map inside one unit of work protect against staleness?
    No. Handing back the same instance for a repeated read gives consistency within that unit of work, not freshness — the instance is as old as the read that created it. If anything it hides the change, because a second read that would have shown the new value never reaches the database.
  • If staleness produces no error, how do you find out it is happening?
    By comparison, not by logs. Sample values the application serves and read the same rows directly, then look at how often they differ and how old the worst difference is. Counters the cache reports itself, such as hit rate or evictions, do not move when entries are wrong.

A cached row is a printed price list. It was accurate when it went to the printer, and nothing about holding the sheet tells you the shop changed a price this morning.

saying these in an interview costs you the question

  • Thinks a stale cached read fails loudly or logs a warning
  • Believes every write updates the cached copy automatically
  • Assumes a set-based statement refreshes the rows it touched
  • Treats a high hit rate as evidence the cache is correct
  • Says only other services can cause stale entries
  • Thinks the identity map keeps loaded objects fresh
open as a page

In a layer whose cache outlives one unit of work, why invalidate a row's entry both before the write and after the commit?

level: middleimportance: must knowfreq 62%

basics

~20 s

Removing before the write stops further hits on a copy about to become wrong; removing again after the commit drops the pre-image that a concurrent reader reloaded and put back during the write. Populate only after a real commit.

open as a page

Your service runs on several instances that each cache rows in memory; after one instance writes, why do the others keep serving the old value, and what makes them agree?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Each process holds its own entry, so an invalidation on one instance is invisible to the rest. Making them agree needs a broadcast removal, a single shared out-of-process copy, or not caching mutable rows per instance at all.

open as a page

Rows your cache holds are also written by another service and by operators; how do you stop it serving values no invalidation will clear?

level: seniorimportance: should knowfreq 48%

basics

~20 s

You cannot invalidate a write you never see, so either consume a change signal from the other writer, check a cheap freshness marker before serving, or accept and publish a bounded maximum entry age. Otherwise do not cache that data.

open as a page

How would you set and verify a maximum-staleness budget for a data-access layer's cache, and why is hit rate the wrong headline number?

level: principalimportance: should knowfreq 42%

basics

~20 s

Derive a per-dataset maximum age from the consequence of acting on an old value, then verify it with convergence probes and divergence sampling. Hit rate measures work avoided and rises when invalidation breaks, so it hides the risk.

open as a page