With Redis-assisted client-side caching, what can leave an application serving a value that Redis has already changed, and what must the client do when its tracking connection drops and reconnects?
answer
- best-effort: no ack, no replay, no ordering
- delivery window T1→T4 = stale locally
- race: invalidation arrives before the GET reply → mark in-flight, discard
- reconnect → flush everything, re-issue HELLO 3 + CLIENT TRACKING
- local TTL + size cap are non-negotiable
basics
~20 sInvalidation is best-effort and asynchronous: there is a delivery window, a race where the value is cached after its invalidation was already sent, and total loss of messages while the connection is down. On reconnect the client must flush its entire local cache, since it cannot know what changed meanwhile. A local TTL and size cap are mandatory backstops.
solid answer
~1 minThree distinct staleness sources: 1. **Delivery window.** The push travels asynchronously; between the write on the server and the client processing the message, the local copy is wrong. Small (sub-millisecond to milliseconds) but non-zero, and it grows under client GC pauses or a saturated event loop. 2. **Cache-after-invalidate race.** The client issues `GET`, and the key is modified before the reply is stored locally — the invalidation can arrive *before* the value is cached, so the client caches a value it was already told to drop. The fix is client-side: mark the key as "in flight" before sending the read, and if an invalidation for it arrives before the reply is stored, discard the reply rather than caching it. 3. **Connection loss.** Invalidations are not queued, acknowledged, or replayed. Everything sent while the socket was down is gone. On reconnect the only safe action is to **flush the entire local cache**, re-issue `HELLO 3` and `CLIENT TRACKING` (with prefixes if BCAST), and warm up again. The same applies when the RESP2 redirect target dies — the server signals `tracking-redir-broken`. So: local TTL as a backstop, bounded size with local eviction, flush on reconnect, and never place data whose staleness is unacceptable in a locally cached copy.
code
text · 11 lines# WRONG client logic
send GET k # invalidation for k may already be in flight
receive invalidate[k] # local map has no entry -> ignored
receive reply V # cached -> stale forever
# RIGHT client logic
mark k as PENDING (before sending)
send GET k
receive invalidate[k] # k is PENDING -> mark POISONED
receive reply V # k is POISONED -> return V to the caller,
# but do NOT store it locallygo deeper
Know that invalidation is best-effort, that a local TTL is still required, and that reconnecting means throwing the local cache away.
Name the three staleness sources — delivery window, cache-after-invalidate race, connection loss — and the reconnect procedure including re-issuing HELLO 3 and CLIENT TRACKING.
Describe the in-flight guard precisely, handle null-key-list and tracking-redir-broken pushes, bound local memory, and instrument invalidations received versus applied.
Set the rule for which data classes may live in a process-local copy at all, quantify the acceptable staleness window per class, and design the degradation path when tracking connections are unstable.
## Tracking is an accelerator, not a coherence protocol Hardware caches maintain coherence with protocols that are synchronous with respect to the memory operation. Redis tracking is nothing like that. It is a **best-effort, fire-and-forget notification**: no acknowledgement, no sequence numbers, no replay, no ordering guarantee relative to your own reads. Understanding what that permits is the whole of this question. ## Source 1: the asynchronous delivery window The timeline is: ``` T0 client A caches value V locally T1 client B: SET key V' # server value changes T2 server pushes invalidate[key] to A T3 A's socket receives it T4 A's event loop processes it and drops the local entry ``` Between T1 and T4, A serves V — a value that no longer exists. Typically that is a fraction of a millisecond inside a datacenter, but the tail is what matters: a full GC pause, a blocked event loop, a client busy parsing a large reply, or a saturated NIC stretches T3→T4 to hundreds of milliseconds. Cross-region, T2→T3 alone can be tens of milliseconds. There is no way to eliminate this. The design question is whether your data tolerates a millisecond-to-second-scale staleness window — for feature flags and catalog metadata, yes; for balances, entitlements, or a permission revocation, usually no. ## Source 2: the cache-after-invalidate race This one is subtle and is the bug most implementations get wrong. ``` T0 client sends GET key T1 server reads key = V, starts sending the reply, records interest T2 another client writes key = V' T3 server pushes invalidate[key] T4 client receives invalidate[key] -> local cache has no entry, nothing to drop T5 client receives the GET reply V -> stores V locally ``` The invalidation for the value arrived **before** the value did, so the client processed it as a no-op and then happily cached a value it had already been told was dead. The entry is now stale **indefinitely** — no further invalidation is coming until the key changes again. The correct client implementation records, **before sending the read**, that this key is being fetched (a "pending"/in-flight marker). If an invalidation for that key arrives while the fetch is in flight, the client marks the pending entry poisoned and **discards the reply instead of caching it**. Redis's own documentation calls this out; most mature client libraries implement it, but a hand-rolled local cache almost never does. It is a good interview signal to be able to describe it. ## Source 3: connection loss Invalidation messages live on the connection. There is no server-side queue, no persistence, no resume-from-offset. If the socket drops — a failover, a network partition, a client-side timeout, a proxy restart — every invalidation that would have been delivered during the gap is simply lost. When the client reconnects, the server has no record of what that new connection previously cached. The only correct action is to **flush the entire local cache on reconnect**, then re-establish tracking: `HELLO 3`, `CLIENT TRACKING on` with the same mode and (for BCAST) the same prefixes, since tracking state is per connection and does not survive it. Warming up again from misses is the price. A related failure exists in RESP2 mode: when invalidations are redirected to a second connection with `CLIENT TRACKING on REDIRECT <id>`, that receiver connection can die independently of the data connection. Redis notifies the tracking client with a `tracking-redir-broken` push so it knows delivery is broken; on RESP2 without such a signal the client must monitor the receiver's health itself. Either way: flush and re-establish. Other whole-cache events to handle: a push carrying a **null key list** (sent on `FLUSHALL`/`FLUSHDB`) means drop everything, and a failover to a replica means the client is now talking to a different server whose tracking state is empty. ## The mandatory backstops Because all three sources exist, a tracked local cache must still behave like an ordinary cache: - **Local TTL.** Bounds staleness when an invalidation is lost, missed, or raced. Tracking shortens typical staleness from "the TTL" to "milliseconds"; the TTL remains the guarantee. Choosing it is the same exercise as any cache TTL. - **Bounded size with local eviction.** An in-process map with no cap is a memory leak that eventually becomes an OOM or a GC pathology. Use an LRU/LFU-bounded structure. - **Flush on reconnect**, on `tracking-redir-broken`, and on a null-key-list invalidation. - **Tolerate spurious invalidations.** The server sends invalidations for keys it evicts from its own tracking table when `tracking-table-max-keys` is exceeded, and BCAST delivers messages for keys you never cached. Extra drops are always safe; the client must never treat an unexpected key name as an error. - **Never cache what must not be stale.** Revocations, entitlements, and balances belong in the round trip. ## Observability Instrument the local cache: hit ratio, entry count, invalidations received versus applied, and reconnect-triggered flushes. A rising flush rate means connection instability that is silently costing you your local hit ratio; a very high received-versus-applied ratio means your BCAST prefixes are too broad. Without those numbers you cannot tell whether the local cache is helping at all.
- Why must the client flush its whole local cache after reconnecting?Invalidation messages are not queued, persisted, or replayed, so everything the server would have sent while the socket was down is lost, and the client cannot know which keys changed. Tracking state is also per connection, so the new connection starts with no registered interest at all. Flushing and re-registering is the only safe recovery; the cost is a warm-up period of misses.
- Describe the race where a client caches a value it was already told to invalidate.The client sends a read; the key is modified before the reply is stored locally, so the invalidation push arrives first and is processed as a no-op because there is nothing cached yet. The reply then arrives and gets cached, leaving a stale entry that no further invalidation will correct until the key changes again. The fix is to mark the key in-flight before sending the read and discard the reply if an invalidation for it arrives in the meantime.
- If the server sends invalidations, why keep a TTL on the local entries at all?Because invalidation is best-effort: messages can be lost on connection failure, raced against an in-flight read by a client that does not implement the in-flight guard, or simply delayed. The TTL converts an unbounded staleness failure into a bounded one. Tracking then improves the typical case from the full TTL down to milliseconds without being the guarantee.
- What should a client do when an invalidation push carries a null key list?Treat it as an instruction to drop the entire local cache. Redis sends it when the database is emptied with FLUSHALL or FLUSHDB, where enumerating every affected key would be pointless. Handling it as an unknown or malformed message and ignoring it leaves the whole local cache stale.
The errata notice can arrive by post while the page you ordered is still in transit — unless you note that a page is on its way and bin it on arrival, you file a page you were already told is wrong.
saying these in an interview costs you the question
- Treating server-sent invalidation as a strong consistency guarantee with no staleness window.
- Keeping the local cache across a reconnect because 'the client reconnected automatically'.
- Not implementing the in-flight guard, so a raced read caches a value that was already invalidated.
- Assuming tracking registration survives the connection, so prefixes are never re-applied after reconnect.
- Treating an unexpected or spurious invalidation as an error instead of a harmless extra drop.