skip to content

In a cache-aside pattern, a service updates a row in its database and then deletes the corresponding key from the cache (instead of writing the new value directly into the cache). Why is 'invalidate on write' usually preferred over 'update the cache value on write', and what race condition can still leave the cache holding stale data afterward?

level: middleimportance: must knowfreq 75%

answer

  1. delete not update
  2. delayed double-delete
  3. race: slow reader vs writer
  4. stale until TTL
  5. versioned/CAS writes

basics

~20 s

Deleting the old cache entry is simpler and safer than trying to recompute and overwrite it. But there's a timing gap: another request could reload the old value into the cache right after the delete, right before the write finishes, leaving the cache wrong until it expires.

solid answer

~40 s

Invalidate-on-write (delete the key) is preferred over update-on-write because it avoids duplicating the write logic that produced the value — the cache doesn't need to know how to compute the new representation, and it sidesteps double-writing bugs where the cache diverges from what the DB write actually stored (e.g., DB-side defaults, triggers, or partial writes). The known race: Thread A does a read, misses the cache, and queries the DB — getting the pre-update value. Concurrently, Thread B updates the DB to the new value and deletes the cache key. If Thread A's cache SET (still holding the old value) lands after Thread B's DELETE, the cache is left with stale data until TTL expiry. Mitigations include short TTLs, delayed double-delete, or versioned/CAS writes.

go deeper

for a junior

Should know that on write, the app typically removes the old cache entry rather than trying to compute a new one, without necessarily knowing why.

for a middle

Should explain why invalidate-on-write avoids duplicating the DB's write logic, and recognize that write and cache-populate are two separate, non-atomic operations.

for a senior

Should be able to walk through the concrete race condition interleaving step by step and name at least one mitigation like short TTL or delayed double-delete.

for a principal

Should weigh whether the race is worth solving with real coordination (locks, CAS) versus accepting a bounded staleness window, based on the actual freshness requirements of the specific data.

## Two ways to handle a write In cache-aside, reads and writes follow separate, independently-timed paths, and the pattern doesn't dictate a single correct way to handle writes — but the two realistic options are: | Strategy | What the write path does | |---|---| | **update-on-write** | after writing the new value to the database, also write that same new value into the cache | | **invalidate-on-write** | after writing to the database, delete the corresponding cache key rather than trying to write a new value into it | Under invalidate-on-write, the very next read for that key becomes an ordinary cache miss, which triggers the standard load-from-database-then-populate flow described by the base pattern, and the cache ends up holding a value that was read fresh from the database rather than one assembled by the write path. ## Why invalidating is preferred Invalidate-on-write is generally preferred because it avoids a whole class of correctness bugs that come from computing the cache's new value in two different places. If a write handler tries to construct the "new" cached value itself and push it into the cache, it has to reproduce whatever the database actually computed: - default column values - triggers - computed/generated columns - timestamps set by the DB engine - or partial-update semantics Any divergence between what the app thinks it wrote and what the database actually stored leaves the cache silently wrong. **Deleting the key sidesteps all of that:** the next reader is guaranteed to get whatever the database says the current value is, because it goes through the exact same read path as any other miss. It also collapses multiple back-to-back writes to the same key into one eventual reload instead of several redundant cache writes. ## The race that still leaves stale data The problem is that invalidate-on-write is not atomic with respect to concurrent reads, and the classic race goes like this: 1. **Thread A** issues a read for key K, misses the cache, and starts a query against the database — at this point it has read the pre-update value of K but has not yet written it back to the cache. 2. Concurrently, **Thread B** performs a write: it updates K in the database to a new value and then deletes K from the cache. 3. Because Thread A's database read already happened before Thread B's write committed, Thread A is holding the old value. 4. If Thread A's cache `SET` (writing that old value) executes after Thread B's `DELETE`, the cache ends up holding the stale pre-update value — and because nothing else will invalidate it, it stays wrong until its TTL expires naturally. ## Why it is so hard to catch This race requires a fairly specific interleaving (a concurrent slow reader racing a writer on the same hot key), so it's rare per-request but not rare in aggregate on high-throughput hot keys, and it's insidious because it doesn't fail loudly — the system just quietly serves stale data for up to a full TTL window, which can look like flaky bugs that are hard to reproduce because they depend on timing. It's more likely to bite when the database read in the miss path is slow (e.g., a complex query or replica lag) relative to the write path, widening the race window. ## Mitigations, and their ceiling Common mitigations include: - keeping **TTLs short** enough that staleness windows are bounded to something acceptable for the domain; - using a **"delayed double delete"** — deleting the cache key again a few hundred milliseconds after the write, to catch any stale value a racing reader wrote in between; - using **compare-and-set/versioned writes** so the cache `SET` only lands if the version it read is still current; - or, for read replicas, making sure the miss-path read is not served by a lagging replica relative to the write. None of these eliminate the race outright without adding real coordination (e.g. a distributed lock around the key), which is usually not worth the complexity for the rarity and bounded blast radius of the issue.

  • Why not just skip cache invalidation entirely and rely only on the TTL to eventually expire stale writes?
    Because TTL-only invalidation means every write leaves the cache stale for up to the full TTL duration, not just for a brief race window — for hot keys with strict freshness requirements (e.g. account balances, inventory counts) that's often unacceptable. Explicit invalidation on write shrinks the staleness window from 'up to TTL' down to 'the rare race window', which is a much better bound.
  • What is the 'delayed double delete' technique and what race does it close?
    After deleting the cache key on write, the app schedules a second delete of the same key a short time later (e.g. 500ms-1s). This catches the case where a concurrent slow reader's stale SET landed after the first delete — the second delete clears that stale value out before the TTL would have. It doesn't guarantee correctness but shrinks the window in which the race can leave stale data.
  • Would using update-on-write instead of invalidate-on-write eliminate the race condition?
    No — the same fundamental race exists either way, because it's about the ordering of a concurrent read's cache SET relative to a write's cache operation, not about whether that operation is a DELETE or a SET. Update-on-write actually adds the extra risk of the write path computing an incorrect new value, on top of the same race.

Like erasing a whiteboard note instead of trying to rewrite it correctly from memory — anyone who looks next will write down the current truth themselves rather than trust your possibly-wrong rewrite.

saying these in an interview costs you the question

  • Thinks update-on-write is strictly safer than invalidate-on-write
  • Can't describe a concrete interleaving that produces stale cache data
  • Assumes cache and database writes are atomic together
  • Doesn't know that the staleness from this race is bounded by TTL

context