skip to content

Phil Karlton famously said cache invalidation is one of the two hard problems in computer science. What are the main strategies for invalidating stale cache entries — TTL expiration, explicit invalidation/purge on write, and versioned or keyed cache-busting — and what can go wrong with each in a distributed system with multiple cache nodes?

level: seniorimportance: must knowfreq 70%

answer

  1. TTL=passive, no coordination, worst-case bound
  2. purge-on-write=active delete, needs fan-out to all nodes
  3. cache-busting=new key per version, no delete needed
  4. combine purge + TTL backstop
  5. content-hash filenames = cache-busting for static assets

basics

~20 s

When the real data changes, the cached copy becomes wrong until it's removed or updated. You can let it expire on its own after a set time, tell the cache directly to delete it when the data changes, or change the cache's key/version so old copies are never looked up again.

solid answer

~50 s

TTL expiration removes entries automatically after a fixed duration, requiring no coordination but bounding freshness only probabilistically — data can be stale for up to the full TTL window. Explicit invalidation (purge-on-write) has the writer actively delete or update the cache entry the moment the source data changes, giving tighter freshness, but it requires the writer to know every cache/key that might hold a copy, and in a multi-node cache cluster or multi-region CDN, that purge has to reach every node or you get partial staleness — some readers see fresh data, others don't. Versioned/keyed cache-busting sidesteps deletion entirely by embedding a version identifier (a content hash, a monotonic version number) into the cache key itself, so a data change simply produces a new key that nobody has cached yet — old versions age out naturally via normal eviction/TTL rather than requiring an active purge, at the cost of temporarily storing multiple versions and needing every reader to know the current version pointer.

go deeper

for a junior

Should know that when data changes, the cache needs to be told somehow, and that TTL (auto-expiry) is one basic way this happens.

for a middle

Should describe TTL and purge-on-write as two distinct mechanisms and give a rough sense of their freshness trade-off.

for a senior

Should explain the multi-node fan-out risk in purge-on-write, why TTL is used as a backstop, and how versioned/content-hash keys sidestep the fan-out problem for static assets.

for a principal

Should design an invalidation strategy across a real multi-region system (e.g., CDN + app cache + DB), including how the 'current version' pointer itself is resolved and kept fresh cheaply, and how to detect/monitor incomplete purge propagation.

## Why invalidation is hard Cache invalidation is hard precisely because a cache is a deliberately-created second copy of data, and the moment the original changes, that copy becomes a lie unless something actively corrects it. The difficulty compounds in distributed systems because 'the cache' is rarely one thing — it's often many independent nodes (a cache cluster, CDN edge nodes spread across continents, per-instance local caches), and a single logical piece of data can have copies sitting in several of them simultaneously, each of which needs to learn about a change independently. ## TTL expiration The simplest strategy is **TTL (time-to-live) expiration**: every cache entry is stamped with an expiration time when written, and once that time passes, the entry is treated as gone (either actively swept or simply ignored/refetched on next access). - **TTL's appeal** is that it requires zero coordination — no writer needs to know who's caching what; the cache entry just quietly stops being trusted after its window. - **The cost** is that TTL gives only a probabilistic, worst-case freshness bound, not an actual freshness guarantee tied to when the data really changed: if a value changes the instant after it was cached, it can be served stale for the entire TTL duration, and if it never changes, you're needlessly refetching once the TTL lapses anyway. - **Choosing the TTL value** is itself a trade-off exercise — too long and staleness windows grow uncomfortable; too short and you lose most of the caching benefit and increase load on the origin. ## Explicit invalidation (purge-on-write) Explicit invalidation (purge-on-write) flips the model: the moment the source of truth changes, the writer (or a change-data-capture pipeline listening to the database's write-ahead log) actively deletes or updates the corresponding cache entry so the next read is forced to miss and refetch fresh data. This gives much tighter freshness than TTL alone, because staleness is bounded by 'how fast can the purge propagate' rather than 'how long is the TTL.' The catch is **coordination**: - The writer has to know every cache entry (potentially across multiple keys, if the same underlying data is cached under several derived views — e.g., a product cached both by ID and as part of a category listing) that could hold a stale copy. - In a distributed cache with multiple nodes or a globally distributed CDN, the purge message has to actually reach every node before it's considered done. In practice this means either a fan-out invalidation broadcast (pub/sub to every cache node, or a CDN purge API call per edge region) or accepting eventual consistency — a short window where some nodes have processed the purge and others haven't, so two users hitting different nodes at the same instant can see different answers. - Missed or dropped invalidation messages (a network partition during the broadcast, a cache node that was down when the purge fired) leave permanently stale entries behind unless there's also a TTL as a backstop — which is why most production systems combine purge-on-write with a TTL safety net rather than relying on purge alone. ## Versioned or keyed cache-busting Versioned or keyed cache-busting takes a structurally different approach: instead of trying to delete or update an existing cache entry in place, it makes every distinct version of the data live under its own distinct cache key — typically by embedding a content hash or a monotonically increasing version number into the key (e.g., `product:42:v17`, or a filename like `app.a3f9c1.js` for a static asset). When the data changes, nothing needs to be purged at all: the new version is simply written under a new key that no reader has ever cached, so every read of the new key is guaranteed to be a fresh miss the first time, and any reader still referencing the old key continues to get the old (now-orphaned) value until it naturally ages out via TTL or normal LRU/LFU eviction. This sidesteps the fan-out purge problem entirely — there's no race to propagate a delete to every node, because nodes that haven't heard about the new version simply don't have it cached yet, and correctness comes from readers looking up the current version pointer (often itself cached with a very short TTL, or resolved via a manifest file) rather than from actively evicting stale data. The cost is that old versions linger in the cache consuming memory until they're evicted by the normal policy, and the system needs a reliable, low-latency way to resolve 'what is the current version' everywhere reads happen — that pointer itself can become a smaller-scoped invalidation problem, though a much cheaper one since it's typically a single small key rather than fanning out purges for every changed data item. ## How they layer in production In production, these strategies are typically layered rather than chosen exclusively: 1. **Static assets** use content-hashed filenames (cache-busting) with effectively infinite TTLs since a changed file gets a new name. 2. **Frequently-changing dynamic data** uses purge-on-write for tight freshness with a TTL backstop in case a purge is dropped. 3. **CDNs** specifically often expose an explicit purge API precisely because pure TTL-based expiration is too coarse for content that needs to go live immediately (a breaking news correction, a price change), while pure purge-on-write is too fragile to trust as the only mechanism across globally distributed edge nodes.

  • Why do most production systems combine purge-on-write invalidation with a TTL, rather than relying on purge-on-write alone?
    Purge-on-write depends on the invalidation message actually reaching every cache node, which can fail during a network partition, a dropped message, or a node that was temporarily down. A TTL acts as a backstop guarantee: even if a purge is silently lost, the stale entry is bounded to expire within the TTL window instead of living forever, trading a small worst-case staleness window for much stronger reliability.
  • How does cache-busting via content-hashed filenames avoid the multi-node fan-out problem that purge-on-write has?
    Because the new content lives under a brand-new key/filename that no cache node has ever seen, there's nothing to actively delete or synchronize — every node simply treats a request for the new filename as an ordinary cache miss and fetches it fresh. The old filename's cached copies become orphaned but harmless, since nothing references them anymore, and they age out through normal eviction rather than requiring a coordinated purge.
  • In a multi-region CDN, what specifically makes purge-on-write invalidation slower or riskier than in a single-node application cache?
    A purge has to be broadcast to every edge location around the world, each potentially in a different network and provider, so propagation isn't instantaneous and can partially fail for some regions while succeeding for others. This creates a real window — often seconds to tens of seconds in practice — where users in different regions see different (stale vs. fresh) content for the same URL, which is why CDN purges are usually paired with conservative TTLs and monitored purge-completion status.

TTL is like a 'best before' sticker you trust blindly until it expires. Purge-on-write is like personally calling everyone who has a copy of a document to tell them to shred it the moment it changes — reliable if the call gets through, broken if someone doesn't pick up. Cache-busting is like renaming the file every time you edit it (report_v2.docx, report_v3.docx) so nobody's old copy is ever confused for the new one.

saying these in an interview costs you the question

  • Thinks TTL guarantees data is never stale beyond its own expiry moment
  • Doesn't recognize that purge-on-write must fan out to every cache node/region
  • Assumes cache-busting requires deleting old entries
  • Can't explain why purge-on-write alone is fragile without a TTL backstop
  • Confuses invalidation strategy with eviction policy (LRU/LFU)

context