skip to content

TTL Strategy

You will learn how to pick and manage TTLs so a cache stays fresh without synchronized mass expiry: jitter, refresh-on-read, and caching misses. Interviewers ask because a fleet of keys expiring in the same second is a classic self-inflicted outage.

part ofRedisoverview, primer and where to startread it →
on this pageshow

questions

5

A record that your service caches in Redis is updated by the write path. Should the write path delete the cache key, overwrite it with a fresh TTL, or overwrite it while keeping the key's remaining time-to-live? Explain how each choice changes the worst-case staleness a reader can see.

level: juniorimportance: must knowfreq 58%

answer

  1. delete = fail-safe, miss is always correct
  2. fresh TTL on every write = safety net never fires
  3. keep remaining TTL = fresh value + hard age ceiling
  4. touch cache after the DB commit, not before
  5. hot key + invalidate = stampede, use single-flight

basics

~20 s

Deleting is the safest default: the next reader reloads from the database. Overwriting with a fresh TTL restarts the clock, so a hot key's expiry never fires and a wrong value can live forever. Overwriting while keeping the remaining TTL keeps a hard age bound.

solid answer

~50 s

Three policies, three staleness bounds. **Invalidate-on-write (delete the key):** the cache can only be empty or correct-as-of-the-last-load. Worst-case staleness collapses to the read-repopulate window, at the cost of a database read after every write and a possible stampede on a hot key. **Refresh-on-write with a new TTL:** the cache usually holds the newest value, but each write restarts the clock. On a key written more often than its TTL, expiry *never* fires, so the TTL stops being a safety net — if a bad or out-of-order value ever lands in the cache, nothing ages it out. **Refresh-on-write keeping the remaining TTL:** you get the fresh value *and* keep a hard upper bound on how long any entry can live since it was first populated, so expiry still backstops drift. I default to invalidate, and use keep-the-remaining-TTL refresh for hot keys where the reload cost hurts. The command-level knobs (`EX` vs `KEEPTTL`) and what a plain overwrite does to an existing expiry are covered in db-redis-eviction-ttl-mechanics-ttl-loss-on-overwrite.

code

text · 11 lines
text
# policy A - invalidate on write (default)
#   DB commit, then drop the entry; next reader reloads truth
DEL user:42

# policy B - refresh, restart the clock
#   fresh value, but expiry never fires on a frequently-written key
SET user:42 <newjson> EX 300

# policy C - refresh, keep the existing deadline
#   fresh value AND a hard ceiling on entry age since first load
SET user:42 <newjson> KEEPTTL

go deeper

for a junior

Know the three options and the one-line consequence of each: delete is safest, restarting the TTL on every write means it never expires, keeping the remaining TTL gives you a fresh value with a hard age limit. Say that you act on the cache after the database commit.

for a middle

Reason about the staleness bound explicitly for each policy, and name the failure the TTL is supposed to catch — missed invalidations, batch jobs writing the database directly, lost Redis writes — to explain why restarting the clock is dangerous.

for a senior

Frame it as choosing an age ceiling per key pattern under a real read/write ratio: invalidate for correctness-sensitive data with single-flight to control the stampede, keep-the-remaining-TTL for hot read-dominated keys, and be explicit about the write-write race and commit ordering.

for a principal

Make it a platform invariant: every cache namespace declares an absolute maximum entry age that no write path may extend, enforced in the shared cache client, with invalidation as the default policy and refresh-on-write an opt-in that must justify its staleness budget.

## The question behind the question When a cached record changes, the cache is momentarily wrong. The write path has to do *something* about that, and the choice is a policy decision — not a command trivia question. What you are really choosing is **the worst-case age of the value a reader can be handed**, and how much load you pay to keep that number small. The mechanical detail of which Redis command preserves or discards an existing expiry belongs to db-redis-eviction-ttl-mechanics-ttl-loss-on-overwrite; here, treat `EX` ("start a new deadline") and `KEEPTTL` ("leave the existing deadline alone") purely as the knob that expresses the policy you picked. ## Policy 1 — invalidate on write (delete the key) The write path updates the database and removes the cache entry. The next reader misses, loads from the source of truth, and repopulates. *Staleness bound:* essentially the length of the window between "database committed" and "cache deleted", plus any in-flight read that already fetched the old value and is about to write it back. There is no long tail: the cache never holds a value the write path knows to be obsolete. *Costs:* every write throws away warm data, so a write-heavy record produces a database read per write. On a very hot key, deleting can release a crowd of concurrent misses onto the database at once — the classic stampede — which you handle with single-flight (one loader per key, others wait) rather than by abandoning invalidation. *Why it is the usual default:* it is the policy that fails safe. A missing entry is always correct; a present entry might not be. ## Policy 2 — refresh on write with a brand-new TTL The write path writes the new value straight into the cache and starts the TTL over. *Staleness bound in the happy path:* near zero — the cache tracks the database write for write. *The trap:* the TTL exists as a **safety net** for everything your invalidation logic misses — a write that bypassed the cache path, a batch job editing rows directly, a lost Redis write, two concurrent writers interleaving so the older value lands in the cache last. Restarting the clock on every write means that on any key written more often than its TTL, **the expiry never fires**. The safety net is nominally configured and practically dead. A value that got into the cache incorrectly can survive indefinitely, and the usual symptom is one user seeing wrong data "forever" while everyone else looks fine. It is also the policy most exposed to write-write races: with two updaters, the last one to reach *Redis* wins the cache even if it was the first to reach the database, so cache and database can disagree with no bound on how long. ## Policy 3 — refresh on write, keep the remaining TTL The write path replaces the value but leaves the original deadline in place. *Staleness bound:* the value is as fresh as the last successful write, **and** the entry is guaranteed to disappear within the TTL originally installed when it was populated. You keep both properties: warm data on the happy path, and a hard age ceiling that guarantees a full reload from the source of truth at a known interval. This is the compromise for hot, read-dominated keys where deleting on every write would hammer the database, but where you still want the expiry to mean something. In effect it says: "trust my write path, but not forever." ## How to choose - **Correctness-sensitive data** (permissions, balances, anything a user will notice being wrong): invalidate. Pay the reload. - **Hot key, cheap tolerance for a few seconds of drift**: refresh keeping the remaining TTL, so the expiry still forces a periodic re-read of truth. - **Refresh with a fresh TTL**: only where the cache is genuinely a materialization of the write path and you have accepted that expiry will never fire — and then set a short absolute lifetime some other way, or you have no age bound at all. ## Two details worth saying out loud **Ordering.** Whatever the policy, act on the cache *after* the database commit. Deleting or overwriting first leaves a window where a concurrent read repopulates from the pre-commit state and re-poisons the cache. **The write to Redis can fail.** Invalidation degrades gracefully — a failed delete leaves a stale entry that the TTL will still reap, provided the TTL is a real deadline and not one that later writes keep restarting. That is the strongest argument for never letting the write path push the expiry endlessly into the future.

  • If you delete the cache key on every write, what happens on a very hot key, and how do you handle it?
    Every write empties the entry, so the next burst of readers all miss at once and pile onto the database — a stampede on a single key. The standard fix is single-flight: the first miss takes a short-lived lock and loads, while other readers either wait briefly for the result or serve the previous value. If the write rate is high enough that this happens constantly, that is the signal to switch that key to refresh-on-write keeping the remaining TTL.
  • Why is "commit to the database, then write the new value into Redis" racy with concurrent writers, and does deleting instead fix it?
    Two writers can commit to the database in one order and reach Redis in the opposite order, leaving the cache holding the older value with nothing to correct it. Deleting narrows the window a lot — a delete carries no value, so it cannot lose a race to a stale value — but it does not close it, because a reader that already fetched an old row can still repopulate after the delete. Practical mitigations are a short TTL as a backstop, version or timestamp checks on the cached payload, or a delayed second invalidation.
  • Should the write path touch the cache before or after the database commit?
    After. If you delete or overwrite first, a concurrent reader can miss, read the pre-commit state, and write that stale value back before your transaction lands — leaving the cache wrong precisely when you thought you had fixed it. Doing it after the commit means the worst case is a brief window where the cache is known-stale, which the delete then closes.

A fresh TTL on every write is like resetting a milk carton's expiry date each time you top it up — the carton never gets thrown out, so whatever went wrong inside stays forever. Keeping the original date still forces a replacement on schedule.

saying these in an interview costs you the question

  • Claiming a TTL bounds staleness regardless of write policy, without noticing that restarting the TTL on every write means it never elapses on a hot key
  • Treating refresh-on-write with a new TTL as strictly better because 'the cache is always current' — it is only current when every write reaches Redis in the right order
  • Saying invalidate-on-write is wrong because it costs a database read, rather than treating that read as the price of a fail-safe policy
  • Concluding that if the write path updates the cache, a TTL is no longer needed at all
  • Confusing 'kept the remaining TTL' with 'kept a stale value' — the value is the new one; only the deadline is inherited
  • Touching the cache before the database commit and not seeing the repopulate-from-pre-commit-state race

context

open as a page

How do you choose the TTL for a value you are caching in Redis? Walk through what actually drives the number, rather than defaulting to five minutes.

level: middleimportance: must knowfreq 65%

basics

~20 s

Start from how stale the data may be for the business, then check it against origin load: with N cached keys and TTL T, expiries alone push roughly N/T refetches per second. Shorten for volatile or sensitive data, lengthen when you also invalidate on write.

open as a page

A nightly warm-up job writes tens of thousands of Redis cache keys, all with the same TTL. Explain the failure this sets up and how randomising the TTL per key changes it.

level: seniorimportance: must knowfreq 50%

basics

~20 s

All the keys expire in the same second, so the origin sees one huge burst of misses and Redis a spike of expiry work. Add jitter — a random offset of roughly 10-25% of the base TTL per key — so expiries spread out over a window wider than the time to refill.

open as a page

Compare a fixed expiry (TTL set once when the value is written to Redis) with a sliding TTL that is extended on every cache hit. When is each the right choice, and what does the sliding version cost?

level: middleimportance: should knowfreq 45%

basics

~20 s

Fixed TTL bounds staleness: the entry dies at a known time no matter how often it is read. Sliding TTL (EXPIRE or GETEX on read) keeps popular entries alive but lets a continuously read key stay stale forever. Use sliding for idle timeouts, fixed for data freshness.

open as a page

A service is flooded with lookups for identifiers that do not exist in the database, and every one of them falls through Redis to the origin. How would you handle this with cache entries and TTLs, and what risks does that introduce?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Cache the miss: store an explicit not-found marker under the same key with a much shorter TTL than positive entries. Risks: the marker must be distinguishable from 'nothing cached', it must be deleted when the entity is created, and attacker-generated keys can flood memory.

open as a page