skip to content

Once a Redis key's TTL elapses, no client can read its value any more — GET returns nil. Given that, how do you serve a slightly stale cached value to most callers while exactly one worker recomputes the fresh one, and how do you decide how much staleness to allow?

level: principalimportance: should knowfreq 30%

answer

  1. Physical TTL = logical freshness + grace
  2. Freshness deadline stored in the payload (or a fresh: sentinel key)
  3. Lock winner rebuilds; losers return stale instantly
  4. Cap it: never serve beyond an absolute max age
  5. Eviction can still take the stale copy — best effort

basics

~20 s

Separate physical from logical expiry: give the Redis key a TTL longer than its freshness deadline and store that deadline inside the value. GET then always returns something; when the value is past its logical deadline, one caller wins a SET NX PX lock and rebuilds while everyone else returns the stale copy immediately.

solid answer

~1 min

Keep two expiries. The **physical** one is the Redis TTL (`EX`), set to freshness window plus a grace period. The **logical** one is a timestamp stored inside the payload — the point after which the value counts as stale. Read path: `GET` returns the payload. If `now < logical_expiry`, return it. If not, try `SET lock:<key> <token> NX PX <ttl>`; the winner recomputes and writes a new payload with a new logical expiry, and everyone else returns the stale value immediately without waiting. A two-key variant achieves the same thing with a long-lived value key plus a short-lived `fresh:<key>` sentinel whose absence signals staleness. This is what makes the mutex non-blocking, and it doubles as an availability buffer: if the source of truth is down, you can keep serving stale rather than failing — but cap that with a hard "too old to serve" bound so a permanently broken source does not serve month-old data forever. Sizing the grace window: at least worst-case recompute plus a retry, and staleness must fit the product's tolerance — seconds for prices and inventory, minutes for recommendations. Under `maxmemory` pressure a stale copy can still be evicted, so treat it as best effort.

code

text · 10 lines
text
# write
SET product:42 '{"v":{...},"fresh_until":1723640400}' EX 360

# read: GET, parse fresh_until, compare to now
GET product:42

# stale -> try to become the one rebuilder
SET lock:product:42 a3f9c1 NX PX 5000
# reply OK  -> recompute + rewrite with a new fresh_until and a new EX
# reply nil -> return the stale payload immediately

go deeper

for a junior

Know that once a Redis key's TTL elapses GET returns nil, and that the trick is to keep the TTL longer than the freshness you promise and record the freshness deadline yourself inside the value.

for a middle

Be able to walk the read path end to end: GET, compare stored deadline to now, on stale try SET NX PX, winner rebuilds, everyone else returns the old value without waiting. Mention the two-key sentinel variant.

for a senior

Add operations: sizing the grace window from worst-case recompute plus a retry, background versus inline refresh, metrics on served-value age, refresh failures counted separately from misses, and the interaction with maxmemory eviction.

for a principal

Frame it as a policy decision — staleness budget per data class, stale-serving as a deliberate availability buffer against upstream failure, a hard maximum age that fails loudly, and the fact that stale windows compose when another cache sits in front.

## The constraint A Redis TTL is a hard visibility boundary: once it elapses, the key is unreadable — `GET` returns nil, `EXISTS` returns 0, and there is no way to ask Redis for "the value that just expired". There is no built-in notion of "expired but still readable", no HTTP-style `stale-while-revalidate` directive that Redis honours on your behalf. (How Redis actually reclaims the memory behind an expired key is a separate concern and does not change what a client can observe.) So if you want stale-serving, you must build it out of a TTL you control plus metadata you store. ## Physical versus logical expiry The technique is to decouple two different questions: - **Physical expiry** — when Redis stops returning the key. This is the memory-management decision. Set it to `freshness_window + grace`. - **Logical expiry** — when the application considers the value stale. Stored *inside* the payload as an absolute timestamp. Because physical > logical, there is a grace window during which `GET` still returns the value while the application knows it is past due. That window is where the refresh happens, and it is why nobody ever has to block. Read path: 1. `GET product:42`. 2. Miss (grace exhausted, or never cached) → this really is a miss: take the lock, recompute, and either wait or degrade. Rare if the grace window is sized right. 3. Hit and `now < fresh_until` → return it. Done. 4. Hit and `now >= fresh_until` → attempt `SET lock:product:42 <token> NX PX 5000`. - Won: recompute, write a new payload with a new `fresh_until` and a new physical TTL, release the lock with a token-checking Lua script, return the fresh value (or return stale immediately and rebuild in the background — see below). - Lost: **return the stale value immediately**. No sleeping, no polling, no thread parked. ## The two-key variant If you cannot change the payload format, use a sentinel: - `v:product:42` — the value, TTL = freshness + grace (or no TTL at all, with explicit deletion). - `fresh:product:42` — a tiny marker with TTL = the freshness window. `EXISTS fresh:product:42` (or an `MGET` of both keys in one round trip) answers "is it fresh?". Absence means "stale, someone should refresh". Same semantics, no serialization change, one extra key per entry. In cluster mode give both keys a shared hash tag (`v:{product:42}`, `fresh:{product:42}`) so you can `MGET` them together. ## Sync versus async refresh The lock winner has two options: - **Refresh inline**: it recomputes and returns the fresh value, paying the latency. Fine when recompute is fast and you would like *someone* to see fresh data promptly. - **Refresh in the background**: it returns the stale value like everyone else and hands the rebuild to a worker or task. Now *no* request ever pays recompute latency, at the cost of needing a background execution path with its own error handling and concurrency limit. For latency-sensitive services this is the better default. ## Choosing the grace window Grace must cover: worst-case recompute latency, plus at least one retry after a failed recompute, plus scheduling delay if refreshes are asynchronous. Too short and the value becomes unreadable mid-refresh, dropping you into the real miss path — the situation you built this to avoid. Too long and stale data lives on after a refresh loop has quietly been broken for hours. ## Choosing the staleness budget This is a product decision, not an engineering one, and it should be made per data class: - **Seconds or less**: prices, inventory counts, anything a user is about to transact on, security decisions. - **Tens of seconds**: dashboards, counts, feeds, search facets. - **Minutes**: recommendations, related-items, trending lists, expensive aggregates. - **Deploy-scoped**: compiled configuration, static reference data. Write the chosen numbers down alongside the cache, because the next person will otherwise "tidy" them. And remember that staleness composes: if an in-process cache also sits in front, the windows add. ## Stale as an availability feature — with a hard cap Once you can serve stale, a failing source of truth stops being an outage: the cache absorbs it. That is genuinely valuable and worth designing for — the refresh fails, the stale value is returned, and users see slightly old data instead of errors. The danger is serving stale *forever*. Add a second, absolute bound — "never serve data older than X" — after which you fail loudly or degrade explicitly. Otherwise a broken refresh path is invisible until someone notices the numbers stopped moving last Tuesday. Instrument it: export the age of served values, alert when p99 age exceeds the freshness window, and count refresh failures separately from cache misses. ## Interaction with eviction A long physical TTL is a request, not a guarantee. Under `maxmemory` with an `allkeys-*` policy Redis may evict your stale copy regardless of remaining TTL, and a restart without persistence loses everything. Stale-serving is therefore a best-effort optimisation layered on top of a miss path that must still be correct — bounded recompute concurrency and a degradation response remain mandatory. ## Why this is usually the right default Compared with the alternatives: a bare mutex forces losers to wait or fail; probabilistic early refresh needs metadata plus enough traffic and does nothing on a cold cache; async background refresh of a fixed key list does not generalise. Stale-serving degrades gracefully in every direction — it turns expiry from an event into a soft transition, keeps latency flat, and converts source-of-truth outages into freshness incidents. Its price is that you must decide, explicitly, how old is too old.

  • How would you make the same scheme work without changing the cached payload format?
    Use two keys: `v:<key>` holding the raw value with TTL = freshness + grace, and `fresh:<key>`, a tiny marker with TTL = the freshness window only. Read both in one `MGET`; if the value is present but the marker is gone, the value is stale and someone should refresh it. In Redis Cluster give both keys a shared hash tag so they land in the same slot and `MGET` stays legal.
  • You are serving stale because the upstream source has been failing for an hour. What should the system do?
    Keep serving, but not silently and not indefinitely. Emit the age of every served value as a metric and alert once p99 age exceeds the freshness window, and count refresh failures separately from ordinary cache misses so the failure is visible. Enforce a hard maximum age beyond which the request fails explicitly or returns a documented degraded response, so a permanently broken refresh path cannot serve arbitrarily old data forever.

A newspaper on a café table: it is yesterday's, but you read it immediately while someone fetches today's edition. Nobody stands at the door waiting — though the café should still throw out papers older than a week.

saying these in an interview costs you the question

  • Claiming Redis can return an expired key if you ask nicely, or that PERSIST/TTL manipulation lets you read a value whose TTL already elapsed.
  • Making the losers of the lock race sleep, poll, or retry — that reintroduces the pile-up the design exists to prevent.
  • Setting the physical TTL equal to the freshness window, which leaves no grace period and collapses the scheme back into a plain expiring cache.
  • Treating a long TTL as a guarantee the stale copy will be there — under maxmemory with an allkeys-* policy it can be evicted at any time.
  • Serving stale with no absolute maximum age, so a broken refresh loop quietly serves indefinitely old data.

context