You are deciding whether a fleet of application instances should hold process-local copies of Redis values kept fresh by Redis's CLIENT TRACKING, rather than reading Redis on every request. How do you make that call and bound the risk?
answer
- decide per data class, not fleet-wide
- read/write ratio × working set × instance count
- what does one second of stale cost?
- TTL + size cap + flush-on-reconnect + kill switch
- price the plain short local TTL first
basics
~20 sAdopt it only where the read/write ratio is very high, the working set fits comfortably in process memory, and the data tolerates a millisecond-scale staleness window plus rare loss. Bound the risk with a local TTL, a hard size cap, flush-on-reconnect, per-class allowlists, and metrics on local hit ratio and invalidation volume — and start with one prefix, not the whole keyspace.
solid answer
~1 minI decide per **data class**, never fleet-wide. **Adopt when:** read/write ratio is high (thousands of reads per write), the set is small and prefix-nameable (config, flags, entitlement tables, small reference data), the value is small, and the business tolerates being a few milliseconds — occasionally seconds — behind. Payoff: the Redis round trip disappears from the hot path, and read load on a single hot shard drops sharply. **Don't adopt when:** the data is authorization, balances, revocations, or anything where a stale read is a correctness or security incident; the working set is large (you multiply memory by instance count); or writes are frequent, where invalidation traffic outweighs savings. **Bounding the risk:** local TTL as the real guarantee; hard entry/byte cap with local eviction; flush on reconnect and on a null-key-list push; mode chosen deliberately (default per-key when the set is scattered, BCAST with narrow, disjoint prefixes when many instances would otherwise saturate `tracking-table-max-keys`); `NOLOOP` on writers; and a kill switch that reverts to plain Redis reads without a deploy. **Prove it with numbers:** local hit ratio, invalidations received vs applied, p99 before/after, Redis ops saved, process RSS delta. If the local hit ratio is not high, you added a staleness window for nothing.
code
text · 10 linesHELLO 3
CLIENT TRACKING on BCAST PREFIX cfg: NOLOOP
OK
# local cache policy alongside it:
# max entries : 10_000
# max bytes : 32 MB
# local TTL : 30s <- the real guarantee
# on reconnect : flush all, re-issue HELLO 3 + CLIENT TRACKING
# kill switch : LOCAL_CACHE_ENABLED=false -> pass through to Redisgo deeper
Know the shape of the trade: faster reads and less Redis load, paid for with process memory and a staleness window, so only for data that tolerates being briefly out of date.
Apply the filters — read/write ratio, working-set size, prefix nameability — and name the mandatory backstops of local TTL, size cap, and flush on reconnect.
Own the rollout: per-class allowlist, mode choice with narrow prefixes and NOLOOP, kill switch, and the metrics that prove the win and expose staleness.
Set the policy across the fleet: which data classes may ever be locally cached, the staleness budget per class, memory and tracking-table budgets as the fleet grows, and when the simpler short-local-TTL answer is the right one.
## Frame it as buying latency with staleness and memory A tracked local cache trades three things: you **gain** the elimination of a network round trip on hits and a large reduction in read traffic to a single Redis node; you **pay** in process memory (multiplied by instance count), in a staleness window measured in milliseconds with an unbounded tail on failure, and in operational complexity that lives inside the client library and is hard to observe. The decision is whether that trade is favourable for a specific data class — not for "the application". ## The quantitative filters **Read/write ratio under the cached set.** This is the dominant term. At 10,000 reads per write, nearly every local read is a saved round trip and invalidation traffic is negligible. At 10 reads per write, you are shipping invalidations almost as often as you are serving hits, and the local copy is nearly always cold. Compute it from real metrics, not intuition. **Working set size × instance count.** A 5 MB config set across 200 pods is 1 GB of aggregate RAM you are choosing to spend — usually fine. A 2 GB entity set across 200 pods is not a design, it is an outage. If the set does not fit comfortably within a fraction of each instance's heap, stop. **Prefix nameability.** If the cacheable keys share a clean prefix, BCAST becomes available and server-side tracking memory is O(prefixes). If they are scattered, you are in per-key default mode, and the fleet's distinct cached-key count must stay well under `tracking-table-max-keys` (default 1,000,000) or the server starts evicting table entries and emitting spurious invalidations, which quietly destroys your local hit ratio. **Value size.** Large values are exactly what you do not want duplicated in every process, and they also lengthen the reply-processing time that widens the staleness window. ## The qualitative filter: what does stale cost? Write down, per data class, what a wrong value for one second causes. - **Feature flags, catalog metadata, currency tables, rendering config:** a second of staleness is invisible. Good candidates. - **Session validity, permission grants, revocations, feature entitlements tied to payment:** staleness is a security or compliance event. Do not cache locally; a revoked token that stays valid for a second on 200 pods is a real incident and is very hard to reason about after the fact. - **Balances, inventory counts, anything decrementing:** the local copy invites read-then-act logic that is wrong under concurrency regardless of the cache. This filter overrides the quantitative ones. A perfect read/write ratio does not make it acceptable to serve a revoked permission. ## Bounding the blast radius Once you adopt, the mitigations are non-negotiable: 1. **Local TTL.** Tracking is best-effort; the TTL is the actual guarantee. Set it to the maximum staleness you would accept if invalidation stopped working entirely — because sometimes it does. 2. **Hard bounds on the local cache.** Entry count and byte size, with local eviction. An unbounded in-process map is a memory leak whose symptom is a GC death spiral in production, not a cache miss. 3. **Flush on every uncertainty event:** reconnect, `tracking-redir-broken`, an invalidation push with a null key list, and failover to a different node. 4. **The in-flight guard** so a read raced by a write is not cached after its own invalidation. Verify your client library implements it rather than assuming. 5. **A runtime kill switch.** One config flag that makes the local cache a pass-through to Redis, flippable without a deploy. Every subtle staleness incident is debugged faster if you can remove the variable in seconds. 6. **Mode discipline.** Never BCAST with an empty prefix on a busy instance. Keep prefixes narrow and disjoint. Set `NOLOOP` on connections that also write. 7. **Roll out one data class at a time**, starting with the one whose staleness cost is lowest, and keep it there long enough to see a failover and a network blip. ## Make it observable or do not ship it The failure mode of a tracked local cache is *silence*: it either quietly does nothing (low hit ratio, all cost, no benefit) or quietly serves stale data. Minimum instrumentation: - local hit ratio and entry count per instance; - invalidations received vs. actually applied (a large gap means BCAST prefixes are too broad, or the tracking table is over-invalidating); - reconnect-triggered flush rate (rising = connection instability silently eating your hit ratio); - p99 latency and Redis ops/sec before and after, to show the win is real; - a sampled correctness probe comparing local values against Redis, reporting a disagreement rate. That last one is what turns "we think it is fine" into a number. ## The alternatives you should have priced first Before adopting tracking, check the cheaper options: a plain **short local TTL** with no tracking at all (dramatically simpler, staleness bounded by the TTL — often entirely sufficient for config data); **fewer, fatter Redis reads** via batching so the round trip is amortized; or **replicas** to spread read load if the problem is node capacity rather than latency. Tracking is worth its complexity when you need *both* near-zero read latency *and* staleness materially shorter than a TTL you could otherwise accept. If a 5-second local TTL would do, take it and skip the machinery. ## The honest summary line Adopt tracked local caching narrowly, for small, hot, low-write, staleness-tolerant data; bound it with a TTL, a size cap, and a kill switch; measure the hit ratio and the disagreement rate; and keep everything security- or money-shaped out of it.
- Which kinds of data would you refuse to hold in a tracked process-local cache?Anything where a stale read is a correctness or security event: permission grants and revocations, session validity, entitlements tied to payment, balances, and inventory counts. A revoked credential that stays valid for a second across two hundred instances is an incident, and the invalidation window plus the possibility of message loss make it indefensible. Those reads stay on the round trip.
- What cheaper alternatives should be priced before adopting tracking?A plain short local TTL with no tracking at all — far simpler, with staleness bounded by the TTL and no invalidation machinery to get wrong — is often sufficient for configuration-shaped data. Batching reads amortizes the round trip if the problem is chattiness, and adding replicas helps if the problem is node read capacity rather than latency. Tracking earns its complexity only when you need both near-zero read latency and staleness materially shorter than an acceptable TTL.
- How would you tell, after rollout, whether it was worth it?Compare p99 latency and Redis ops per second before and after, and read the local hit ratio: if it is not high, you bought a staleness window for nothing. Watch invalidations received versus applied to detect prefixes that are too broad, and the reconnect-flush rate to detect instability quietly destroying the hit ratio. A sampled probe comparing local values against Redis gives you an actual disagreement rate rather than a hope.
- Why does instance count feature so heavily in the sizing decision?Every instance holds its own copy, so process memory for the working set is multiplied by the fleet size, and in default per-key mode the fleet's distinct cached-key count is what fills the server's tracking table. A set that is trivially small on one node can be a gigabyte of aggregate RAM and a saturated tracking table at two hundred nodes, at which point the server starts over-invalidating and the local hit ratio collapses.
It is like giving every branch office its own printed copy of the price list: enormously faster at the counter, fine for descriptions, and a scandal for prices unless you can prove the errata always arrive.
saying these in an interview costs you the question
- Enabling it fleet-wide for all keys rather than per data class with an allowlist.
- Ignoring that process memory and tracking-table pressure scale with instance count.
- Putting authorization or revocation data behind a local copy because the read/write ratio looks attractive.
- Shipping without a local TTL, a size cap, or a runtime kill switch.
- Declaring success without measuring local hit ratio or any local-versus-Redis disagreement rate.