skip to content

In the cache-aside pattern, every cached entry carries a TTL (time-to-live) after which it's treated as a miss even if nothing invalidated it explicitly. What trade-off are you making when you pick a short TTL versus a long TTL, and why can't you rely on TTL alone to guarantee freshness?

level: seniorimportance: must knowfreq 70%

answer

  1. TTL = max staleness bound, not correctness
  2. content-blind, time-based only
  3. short TTL = fresher but more DB load
  4. pair TTL with explicit invalidation
  5. synchronized expiry -> stampede

basics

~20 s

Short TTL means fresher data but more trips to the database because things expire fast. Long TTL means less database load but data can be wrong for longer. TTL alone doesn't fix staleness — it just puts a ceiling on how long stale data can survive.

solid answer

~40 s

TTL is a time-based, content-blind expiry: after it elapses, the cache drops the entry regardless of whether the underlying data actually changed, and the next read takes the normal miss path. Short TTLs bound worst-case staleness tightly but increase miss rate, database load, and average latency. Long TTLs maximize hit rate and minimize load but widen the staleness window to the full TTL. TTL alone can't guarantee freshness because between write and expiry the cache serves whatever it has with zero awareness of changes — it's a safety net for missed invalidations, not a substitute for explicit invalidation on write, and production systems pair both rather than relying on TTL alone.

go deeper

for a junior

Should know that TTL means an entry automatically expires after a fixed time and is then treated as a miss.

for a middle

Should articulate the basic trade-off: shorter TTL means fresher-but-more-load, longer TTL means less-load-but-staler.

for a senior

Should explain clearly why TTL is a bound, not a guarantee, and why production systems pair it with explicit invalidation rather than relying on it alone.

for a principal

Should reason about TTL choice per data class based on actual business cost of staleness, and proactively flag synchronized-expiry stampede risk when reviewing a TTL strategy.

## What a TTL actually is In cache-aside, a TTL is a lifetime attached to a cache entry at the moment the application writes it, after which the cache server itself expires the entry — either by deleting it proactively on a background sweep or lazily on the next access — without any involvement from the application's invalidation logic. Once expired, a read for that key is treated exactly like an ordinary cache miss: the cache returns "not found," and the application falls through to the standard load-from-database-then-repopulate path, refreshing both the value and the TTL clock. TTL is therefore a purely time-based, **content-blind** mechanism — it has no idea whether the underlying data actually changed, only how long ago the entry was written. ## Why a TTL exists at all TTL exists as a safety net for staleness that isn't caught by explicit invalidation: - data whose write path forgot to invalidate the cache; - data that changed via a process outside the application's control (a direct DB update, an out-of-band batch job, another service entirely); - or data where explicit invalidation is impractical to wire up for every code path that can mutate it. ## Short against long Picking the TTL value is a direct trade-off between staleness tolerance and load/latency, and a long TTL does the opposite of a short one: | Choice | What it buys | What it costs | |---|---|---| | **A short TTL (seconds)** | bounds how long the cache can ever be wrong to a tight window, which matters for data where staleness is costly (pricing, inventory, session state) | it also means keys expire and get evicted far more often, so the miss rate against the database goes up, database load rises, and average read latency rises because a larger fraction of reads pay the full miss-path cost | | **A long TTL (hours)** | maximizes cache hit rate and minimizes database load and latency | widens the worst-case staleness window to the full TTL duration, which is dangerous for data where an hour-old value could be actively wrong or harmful | ## Why TTL alone cannot guarantee freshness TTL alone cannot guarantee freshness because it's a maximum-staleness bound, not a correctness mechanism — between the moment an entry is written and the moment it expires, the cache will happily serve however-stale-it-is data to every single reader with zero regard for whether the underlying row changed one second or one hour ago. If the write path doesn't also perform explicit invalidation (delete-on-write) or the invalidation path itself has bugs or race conditions, TTL is the only thing standing between a stale write and every subsequent reader, for the entire TTL window. This is why production systems almost always pair TTL with explicit invalidation rather than using TTL as their sole freshness mechanism: TTL handles the "something slipped through" case, while invalidation handles the common case promptly. ## Failure modes - **Synchronized expiry causing a stampede.** The most visible production failure mode from TTL choices: if many related keys (or one very hot key) are all written around the same time and given the same TTL, they all expire together, and a burst of concurrent requests all miss simultaneously and hammer the database at once — this is worse the shorter and more uniform the TTL is. - **Silent staleness.** The opposite failure mode from an overly long TTL is silent staleness that nobody notices until a customer complains that a price, balance, or status they see doesn't match reality, and by the time it's noticed the root cause (a missed invalidation weeks ago plus a very long TTL) can be hard to trace. - **Jitter mismanagement.** A subtler failure is TTL jitter mismanagement: teams that fix the stampede problem by adding randomness to TTLs sometimes widen staleness variance unpredictably across keys, making "how stale can this be right now" harder to reason about. ## What this looks like in practice A concrete example: an e-commerce product page caches price and stock with a short TTL (e.g. 60 seconds) as insurance against missed invalidations, while the actual price-update and inventory-decrement code paths perform explicit cache deletes so that, in the normal case, a price change or a purchase is reflected within milliseconds rather than up to a minute later. That short TTL is deliberately chosen — worth the extra database load — because a customer buying a product at a price that's a minute stale, or seeing "in stock" for an item that just sold out, is a real business cost, whereas a less price-sensitive piece of data like a user's display name might reasonably use a TTL of many hours because staleness there is nearly harmless.

  • If a team relies purely on TTL with no explicit invalidation on write, what's the practical staleness guarantee they're giving users?
    The strongest guarantee they can offer is 'no cached value is ever older than the TTL' — but that means every single read during that window could be serving data that's already out of date the instant a write happens. For any data where correctness matters within that window, TTL-only is not a strong enough guarantee.
  • Why does a shorter TTL increase database load even when the underlying data rarely changes?
    Because TTL expiry is unconditional and content-blind — the entry disappears from the cache purely because time passed, regardless of whether the data actually changed. That forces a fresh database read to repopulate the cache far more often than the data's actual change rate would require, wasting reads on data that was still perfectly valid.
  • What's one way to soften the stampede risk of a short, uniform TTL without abandoning short TTLs entirely?
    Add random jitter to the TTL of each entry (e.g. 55-65 seconds instead of a flat 60), so keys written around the same time don't all expire in the same instant. This spreads the reload load over a window instead of concentrating it into a single spike.

Like a 'best by' date on a carton of milk — it forces you to throw it out eventually even if nobody checked whether it actually spoiled, but it does nothing to stop you drinking spoiled milk the day before that date.

saying these in an interview costs you the question

  • Believes TTL alone guarantees the cache is always correct
  • Can't explain why a shorter TTL increases database load
  • Doesn't recognize synchronized TTLs as a stampede risk
  • Treats TTL and explicit invalidation as redundant/interchangeable rather than complementary

context