skip to content

You decide to put a short-lived in-process cache in front of Redis for a handful of extremely hot cache entries. How do you choose the TTL and the scope, and what failure modes does that introduce?

level: principalimportance: should knowfreq 30%

answer

  1. 1s local TTL ≈ 1 GET/s per instance per key
  2. Staleness adds: local TTL + Redis TTL
  3. Bounded LRU over measured-hot keys only
  4. Jitter TTLs; deploys cause a cold herd
  5. Never near-cache kill switches beyond your incident budget

basics

~20 s

Cache only measured-hot, read-mostly entries in a bounded in-process map with a TTL of about a second. That collapses each instance's Redis traffic for the key to one GET per second, at the price of staleness up to local TTL plus Redis TTL, and instances briefly disagreeing with each other.

solid answer

~1 min

A near cache is the highest-leverage hot-key fix because it removes the network entirely: at 5,000 rps per instance, a 1-second local TTL turns 5,000 Redis GETs into one. Design choices: - **Scope:** only keys proven hot by measurement, in a **bounded** LRU (Caffeine, Guava, an LRU dict) with an explicit max size — otherwise you duplicate the keyspace into every process's heap. - **TTL:** derive it from the staleness budget, not from a habit. Total worst-case staleness is local TTL + remaining Redis TTL, and they stack. Sub-second to a few seconds is where the read-amplification win is already near-total. - **Jitter:** local TTLs across instances drift into sync under steady traffic; add randomness so they do not all miss simultaneously and hammer Redis together. Failure modes: instances serve different values for the same key at the same moment (visible as flapping on refresh behind a load balancer); invalidation is impossible without a channel (Pub/Sub message or RESP3 client-side tracking); a rolling deploy empties every local cache at once, producing a cold herd on Redis; and heap plus GC pressure grows with entry size. Never near-cache data that must be immediately consistent — kill switches, permissions, balances — unless the TTL is inside the tolerated staleness.

go deeper

for a junior

Know that keeping the value in memory for a second or two removes almost all Redis calls, and that the copy can be out of date.

for a middle

Explain the read-amplification arithmetic, the need for a bounded cache, and that invalidation requires TTL expiry or an explicit channel.

for a senior

Cover jitter, cold-start herds after deploys, per-instance divergence visible to users, GC cost, and which data classes must never be near-cached.

for a principal

Set the policy: define the staleness budget per data class, pick TTLs from it, require a kill switch and observability, and decide where near caching sits relative to splitting and replicas.

## What a near cache actually buys Every Redis hit is a network round trip plus work on a shared, single-threaded server. An in-process cache turns that into a hash lookup. The arithmetic is what makes it the strongest hot-key tool: - 40 instances × 5,000 rps on one key = 200,000 Redis GETs/s. - With a 1-second local TTL, each instance issues at most ~1 GET/s for that key → ~40 GETs/s total. That is a four-orders-of-magnitude reduction from a change that touches only the caching client. No resharding, no extra machines, no key renaming. ## Choosing the scope The first decision is *which* entries. Two failure patterns bracket the answer: - **Cache everything locally** and every process ends up holding a large slice of the keyspace: heap grows, GC pauses lengthen, and your consistency story gets worse across the board for entries that were never hot. - **Cache nothing locally** and you keep paying the network for keys that never change. The right scope is a **bounded LRU over measured-hot, read-mostly, small entries**. Bounded means a hard `maximumSize` (or byte-weight limit) — an unbounded map in front of a cache is a memory leak with extra steps. "Measured" means fed by the hot-key detection you already run, or a static allowlist of known globals (feature flags, config blobs, currency tables, top-N lists). Small matters because the entry is duplicated in every process. ## Choosing the TTL TTL is a **staleness budget**, and the budget is the product's, not the engineer's. Two rules: 1. **Staleness stacks.** Worst case, a value that was written to Redis just before you read it can be served from your local cache for the full local TTL, and the Redis copy itself may be up to its own TTL behind the source of truth. Total = local TTL + Redis TTL. Teams routinely forget the addition and are surprised by a 65-second stale window from a "5-second" local cache. 2. **The win saturates fast.** Going from no local cache to 1 second removes ~99.98% of the traffic in the example above; going from 1 second to 60 seconds removes almost nothing more but multiplies staleness sixty-fold. So start at the smallest TTL that solves the load problem — usually well under a few seconds — and only extend it if the numbers demand it. Some entries justify longer: an entry that changes on deploy only (a compiled ruleset) can sit locally for minutes, provided there is an invalidation channel or the deploy itself clears it. ## Failure modes **Per-instance divergence.** Instances refresh at different moments, so at any instant different instances hold different versions. Behind a load balancer, a user pressing refresh may see the new value, then the old, then the new — "flapping". If the UI shows a value that the user just changed, this looks like a bug. Mitigations: pin a user's session to an instance for a short window, bypass the local cache for read-after-write paths, or accept it explicitly for data where it does not matter. **Invalidation is hard.** Once a copy is in a process's heap, Redis cannot reach in and delete it. Options in increasing sophistication: rely on TTL expiry only (simplest, and correct if the TTL is inside the staleness budget); publish invalidation messages on a Pub/Sub channel that every instance subscribes to (delivery is best-effort — a disconnected subscriber misses the message, so keep the TTL as a backstop); or use the RESP3 client-side caching / tracking protocol, where the server itself notifies clients that a cached key changed (a distinct mechanism with its own semantics, and a topic in its own right). **Synchronized expiry.** Under steady traffic, entries populated at the same moment expire at the same moment. Across many instances that means a synchronized burst of Redis GETs, and if Redis has also expired the key, a synchronized burst of recomputes. Add jitter (±10–25%) to local TTLs. **Cold-start herd.** A rolling deploy, an autoscaling event, or a crash loop empties local caches. Every restarted instance repopulates from Redis at full request rate for its first TTL window. Usually survivable — Redis is the thing you are protecting and it is fast — but if the local cache is masking an already-marginal Redis, the deploy becomes the incident. Stagger rollouts; consider pre-warming the handful of global keys at startup. **Memory and GC.** Entries are duplicated per process, and in managed-runtime languages large object graphs promote to the old generation and lengthen pauses. Prefer caching the serialized form, or the parsed form if parsing is the actual cost — measure which. **Debuggability.** "Redis has the right value but production shows the old one" is a confusing incident. Expose the local cache's contents and age via a diagnostic endpoint, emit local hit-ratio and entry-age metrics, and offer a kill switch that disables near caching without a deploy. ## What not to near-cache Anything where a stale read is a correctness or safety failure inside your TTL: authorization decisions and permission sets, feature kill switches used to stop an incident (a 60-second local cache means your kill switch takes 60 seconds), quota and balance checks, and anything a user just wrote and will immediately read back. For kill switches specifically, choose a TTL you would accept as your incident-mitigation latency — often 1–5 seconds — and document it. ## How it composes A near cache stacks with everything else: it is layer 0 in front of Redis, and it reduces the load that key splitting or replica reads would otherwise have to absorb. In practice, try it *first* for a read-hot key, measure the remaining Redis traffic, and only then decide whether splitting or replicas are still needed.

  • How do you invalidate an in-process cache when the underlying value changes in Redis?
    Three levels. Simplest is to rely on a short TTL and accept bounded staleness. Next is a Pub/Sub invalidation channel that every instance subscribes to, publishing the changed key on write — but delivery is best-effort, so a disconnected subscriber silently misses it and the TTL must remain the backstop. Strongest is RESP3 client-side caching, where Redis itself tracks which client cached which key and pushes an invalidation.
  • Users report a value that flips back and forth when they refresh the page. How does a near cache explain that, and what would you do?
    Each application instance holds its own copy with its own expiry, so during the refresh window some instances have the new value and some the old; a load balancer sends successive refreshes to different instances. Fix by shortening the TTL, bypassing the local cache on read-after-write paths for the acting user, or pinning that user's session to one instance briefly. Adding an invalidation broadcast narrows the window but does not close it.

Keeping a photocopy on your desk instead of walking to the filing cabinet each time: instant, but your copy is as old as the last time you walked, and the person at the next desk has a different one.

saying these in an interview costs you the question

  • Treating local TTL and Redis TTL as alternatives rather than additive staleness
  • Using an unbounded map as the local cache
  • Assuming Redis can invalidate an entry already sitting in an application process's heap
  • Near-caching feature kill switches or authorization data with a TTL longer than the incident-response budget
  • Setting a long local TTL because it 'saves more calls' without checking that the win beyond ~1 second is negligible

context