skip to content

When your application updates a record that is also cached in Redis, do you delete the cached entry or overwrite it with the new value — and in which order relative to the database write?

level: middleimportance: must knowfreq 64%

answer

  1. delete, don't overwrite
  2. delete is idempotent; SET is not
  3. commit first, invalidate second
  4. invalidate-first re-caches old row for full TTL
  5. extra miss beats silent wrong value

basics

~20 s

Delete the cached key rather than overwriting it, and delete only after the database transaction commits. Deletion is idempotent, so racing writers cannot leave a wrong value stuck until the TTL expires. The cost is one extra miss.

solid answer

~50 s

**Delete, don't update.** After a successful database write, remove the cached key and let the next reader repopulate it through the normal miss path. Overwriting means two concurrent writers race independently in the cache: the database serializes their commits and picks a true final row, but the loser's cache write can land last and stick for the whole TTL. Deletion is idempotent and order-insensitive — whoever deletes last, the next reader reloads whatever the database actually settled on. Overwriting also drags read-path work (joins, aggregates, serialization) onto the write path for an entry nobody may read. The price of deleting is one extra miss instead of a silent wrong value. **Order: commit first, invalidate second.** Invalidating before the commit lets a concurrent reader miss, load the still-old row, and re-cache it with a fresh full TTL — stale long after the write finished. For removal, prefer `UNLINK` over `DEL` on large aggregates; the lazy-free question covers why.

code

text · 10 lines
text
# 1. database write commits first
#    BEGIN; UPDATE products SET price=2499 WHERE id=915; COMMIT;

# 2. only after the commit succeeds, remove the cached entry
UNLINK product:v1:915
(integer) 1

# 3. the next reader misses and repopulates from the committed row
GET product:v1:915
(nil)

go deeper

for a junior

Know the rule and the order: delete the cached key rather than overwriting it, and do it after the database write succeeds. Being able to say 'the next read reloads it from the database' is the core recall.

for a middle

Explain why deletion beats overwriting — deletes are idempotent so any interleaving of two writers converges, while two SETs can land in the opposite order from the commits. Also explain the invalidate-first hole that re-caches the old row for a full TTL.

for a senior

Frame it as trading a predictable extra miss for the elimination of a silent wrong value, and note the residual write-back race that commit-then-invalidate does not close. Mention TTL as the backstop for lost invalidations and monitoring failed deletes as a correctness signal.

for a principal

Generalize to the invariant: on the write path prefer operations whose interleavings all converge to the same state, which is why 'absent' is a safer cache state than 'one of several candidate values'. Name the exceptions where in-place writes are legitimate (single-writer-per-key, write-through, atomic mutations) and be explicit about which guarantee you are relying on.

## The two options on the write path *Cache-aside* (lazy loading) means the application, not Redis, owns cache population: a read looks in Redis, and on a miss loads from the database and stores the result. When a record then changes, you must decide what the write path does to the cached copy. There are exactly two options: - **Update in place** — recompute the cached representation and write it over the old one. - **Invalidate** — remove the cached key and let the next reader repopulate it through the ordinary miss path. Almost always, invalidate. ## Why invalidation wins under concurrency This is the decisive argument. Suppose writer A sets `price=10` and writer B sets `price=20` at nearly the same moment, and each pushes its own value into the cache after committing. The database serializes the two transactions and produces one true final row — say B's. But the two cache writes are separate, unsynchronized operations over a different connection with different latency. Nothing forces them into the same order as the commits. A's `SET` can easily land *after* B's, leaving Redis holding `price=10` while the database holds `price=20`. That divergence is not self-healing: it persists for the full remaining TTL, and every reader in that window gets a confidently wrong answer. Now replace both cache writes with deletes. Deletes are **idempotent** and **order-insensitive**: whichever of them executes last, the outcome is identical — the key is gone. The next reader misses, loads from the database, and caches whatever the database actually settled on. The race still happens; it just no longer has a wrong outcome. That is the general shape of the rule: prefer operations whose interleavings all converge to the same state. "Absent" is such a state; "holding one of several candidate values" is not. ## The other two reasons **Cost.** The cached entry is frequently not the raw row — it is a join, an aggregate, a rendered projection, serialized into some payload format. Updating in place means paying that entire read-path cost on every write, including for entries that will never be read again. Invalidation defers the cost to the moment someone actually asks, which is the whole point of a lazily-populated cache. **Duplication.** A write path that reconstructs every cached projection of a record has copied the read path's logic into a second place. The two drift: someone changes the read query, forgets the writer, and the cache begins serving a subtly different shape than the loader would produce. ## The price you pay Invalidation guarantees a miss on the next read of that record — a database round trip that an in-place update would have avoided. On a hot key updated frequently, that means repeated reloads. This is a good trade: **one predictable extra miss instead of a silent wrong value.** A miss is visible in your hit-rate metric and self-correcting; a stale cache entry is neither. (If a single very hot key produces a stampede of simultaneous reloads after invalidation, that is a separate concern with its own guards, such as a short-lived lock so only one loader hits the database.) ## Ordering: commit first, then invalidate Invalidating *before* the database write opens a wide, easily-hit hole: 1. Writer removes the cached key. 2. Writer's transaction is still running (or hasn't started). 3. A reader misses, loads the **old** row, and caches it with a **fresh full TTL**. 4. Writer commits. The cache is now stale, and the invalidation that was supposed to fix it has already happened. Nothing corrects it until expiry — potentially minutes or hours after the write. Worse, the window is as long as the write transaction itself, which is exactly when concurrent readers are most likely to arrive. Commit-then-invalidate has a much smaller hole: only a reader that *already* fetched the pre-commit row and is slow to write it back can lose the race with the delete. That window is one reader's database round trip, not one writer's transaction, and it requires unlucky interleaving. A separate question in this area covers detecting and guarding that residual stale-write-back race. ## Commit boundaries and failures Invalidate **after the transaction commits**, never inside it. Deleting inside an open transaction lets a reader repopulate from the pre-commit state, and if the transaction rolls back you have thrown away a perfectly valid entry for nothing. If the invalidation itself fails — Redis unreachable, connection reset — the entry stays stale. This is precisely why every cache-aside entry should carry a TTL: it bounds the blast radius of a lost invalidation to the remaining lifetime rather than forever. Log and count failed invalidations; a rising rate is a correctness signal, not background noise. ## Which removal command Both `DEL` and `UNLINK` remove the key; prefer `UNLINK` for large aggregates and as a default habit — the lazy-free question in this area explains the mechanism. ## The narrow case for updating in place Write-through caches (where the cache sits in front of the store and takes the write itself) and single-writer-per-key designs can legitimately write the value. If you can genuinely guarantee one writer per key, or you're storing a counter you mutate with an atomic Redis operation rather than overwriting, the concurrency argument dissolves. Say so explicitly when you claim the exception.

  • Why invalidate after the database write rather than before it?
    Invalidating first opens a window in which a concurrent reader misses, loads the still-old row, and re-caches it with a fresh full TTL — so the cache is stale long after the write finished, and the invalidation that should have fixed it has already happened. Committing first shrinks the window to one reader's in-flight database round trip instead of the whole write transaction. It also avoids throwing away a valid entry if the transaction rolls back.
  • Deleting the key means the next read always pays a database round trip. When is that unacceptable, and what do you do instead?
    On very hot keys, invalidation can trigger many simultaneous reloads for the same key, hammering the database. The usual guard is to let only one loader repopulate — a short-lived per-key lock or single-flight in the application — while others wait briefly or serve the previous value. Switching to in-place updates to avoid the miss is the wrong fix, because it reintroduces the last-write-wins hazard.
  • Is there any case where writing the new value into the cache is correct?
    Yes, when concurrent conflicting writes to the same key are structurally impossible or harmless: a single-writer-per-key design, a write-through cache where the cache itself owns the write ordering, or a value mutated through an atomic Redis operation such as INCR rather than blind overwrite. In those cases the interleaving argument doesn't apply, but you should state the guarantee explicitly rather than assume it.
  • What happens if the invalidation call to Redis fails?
    The stale entry survives, and no later event will remove it, so every cache entry needs a TTL to bound the damage to its remaining lifetime. Failed invalidations should be logged and counted as a correctness metric, since a rising rate means users are being served old data. Some systems retry asynchronously through a queue or an outbox so the delete is not lost with the request.

Overwriting the cache on write is like two people rewriting the same whiteboard from memory — whoever writes last wins, regardless of who was right. Deleting is wiping the board, so the next person copies the current figure straight from the ledger.

saying these in an interview costs you the question

  • Claiming the cache should be updated with the new value on every write, without acknowledging that two concurrent writers can leave the loser's value cached until TTL.
  • Invalidating the cache before or inside the database transaction, which lets a reader re-cache the pre-commit row with a fresh full TTL.
  • Assuming that because the database serializes the two commits, the corresponding cache writes must land in the same order — they are independent operations with independent latency.
  • Treating a missing TTL as fine because 'we always invalidate on write' — a failed or lost invalidation then leaves the entry stale forever.
  • Saying invalidation and in-place update are equivalent because 'both end up correct eventually', ignoring that a wrong cached value is silent while a miss is visible and self-correcting.

context