You run several app instances with local caches plus a shared Redis cache. How do you keep them consistent and resilient to Redis failures?
answer
- L2 shared = mostly free; L1 near-cache needs broadcast
- pub/sub invalidation = best-effort + TTL backstop
- keyspace notifications (notify-keyspace-events)
- CacheErrorHandler -> degrade to DB, not error
- stampede: sync=true is per-JVM only
basics
~20 sUse Redis as the shared cache (RedisCacheManager) so all instances see the same entries, and broadcast invalidations via pub/sub (RedisMessageListenerContainer) when local near-caches must be evicted. For failures, wrap the cache with a CacheErrorHandler or circuit breaker so Redis outages degrade to the DB instead of erroring.
solid answer
~50 sA distributed RedisCacheManager already keeps instances consistent for the shared tier — a write on node A is visible to node B on next read, and @CacheEvict deletes the shared key for everyone. The hard part is *local/near caches* (an in-JVM L1 in front of Redis L2): those must be invalidated across nodes. Two options: (1) Redis pub/sub — publish the evicted key on a channel, a RedisMessageListenerContainer on each node evicts its local copy; simple but lossy (a node offline during the broadcast keeps stale data until TTL). (2) Redis keyspace notifications (CONFIG notify-keyspace-events) to react to expirations/deletes. For resilience, implement CacheErrorHandler so RedisConnectionFailureException on get/put doesn't propagate — you fall back to the underlying method (the DB). Add short TTLs as a safety net, and consider a circuit breaker so a Redis outage doesn't add latency to every call. Beware cache stampede on cold/expired keys.
code
java · 20 lines@Configuration
@EnableCaching
public class ResilientCacheConfig implements CachingConfigurer {
@Override
public CacheErrorHandler errorHandler() {
// Redis down? log and fall through to the real method (DB), don't blow up.
return new CacheErrorHandler() {
private final Logger log = LoggerFactory.getLogger("cache");
@Override public void handleCacheGetError(RuntimeException e, Cache c, Object k) {
log.warn("cache get failed, bypassing: {}", e.getMessage());
}
@Override public void handleCachePutError(RuntimeException e, Cache c, Object k, Object v) {
log.warn("cache put failed: {}", e.getMessage());
}
@Override public void handleCacheEvictError(RuntimeException e, Cache c, Object k) {}
@Override public void handleCacheClearError(RuntimeException e, Cache c) {}
};
}
}go deeper
Understand that a shared Redis cache is visible to all instances, unlike per-JVM caches.
Use @CacheEvict for the shared tier and know Redis outages surface as exceptions without a handler.
Design pub/sub invalidation for near-caches, add a CacheErrorHandler, and use TTL as a safety net.
Trade off L1+L2 topology, invalidation delivery guarantees (pub/sub vs Streams vs keyspace notifications), stampede control, and Redis-as-performance-vs-availability dependency with circuit breaking.
**Two consistency problems, not one.** 1. *Shared Redis cache only.* With a single `RedisCacheManager` L2 and no per-node cache, consistency is mostly free: every node reads/writes the same Redis keys, `@CachePut`/`@CacheEvict` mutate the shared entry, and there's one source of truth. The residual race is classic read-through interleaving (two nodes fill from a stale DB), mitigated by TTL and by evicting *after* the DB commit. 2. *Near-cache / L1+L2.* For latency you may add an in-JVM cache (Caffeine) in front of Redis. Now each node has a private copy that can go stale when another node writes. Redis doesn't push invalidations to Spring caches automatically, so you must broadcast. **Invalidation broadcast via pub/sub.** On every eviction, `convertAndSend("cache.evict", cacheName + "::" + key)`; each node's `RedisMessageListenerContainer` receives it and evicts its *local* L1 entry. This is the standard fan-out invalidation pattern. Caveat: pub/sub is fire-and-forget — a node that is restarting or network-partitioned during the broadcast **misses** the eviction and serves stale data until its L1 TTL expires. So local caches must have a bounded TTL as a backstop; treat the broadcast as best-effort acceleration, not a guarantee. **Keyspace notifications.** Redis can emit events on key changes/expiry if you enable `notify-keyspace-events` (e.g. `Egx`). You can subscribe with a `PatternTopic("__keyevent@0__:expired")` to react when Redis expires a key — useful to cascade-evict L1 without your app explicitly publishing. Costs CPU on Redis and still shares pub/sub's no-replay weakness. **Resilience to Redis outages.** By default a cache operation that throws (e.g. `RedisConnectionFailureException`) propagates out of the `@Cacheable` method — turning a cache *optimization* into an availability *dependency*. Fix with `org.springframework.cache.interceptor.CacheErrorHandler`: override `handleCacheGetError`/`handleCachePutError` to log-and-swallow, so Spring proceeds to invoke the real method (hits the DB). Register it by extending `CachingConfigurer` and returning your handler from `errorHandler()`. Layer a circuit breaker (Resilience4j) around Redis so, during a sustained outage, calls skip Redis entirely instead of paying connect-timeout latency on every request. **Cache stampede / thundering herd.** When a hot key expires, many concurrent requests miss simultaneously and all hit the DB. Spring's `@Cacheable(sync = true)` serializes concurrent computation of the *same* key *within one JVM* (per-cache lock), but does NOT coordinate across nodes — cross-node stampede needs a distributed lock or probabilistic early expiration. Note `sync=true` also disables some features (no `unless`, single cache) and, on `RedisCache`, uses a lock implemented via Redis, so understand its semantics before enabling. **Operational levers.** Set a Redis `maxmemory` + eviction policy (`allkeys-lru`) so the cache can't OOM the server; keep cache TTLs meaningfully shorter than data volatility windows; monitor hit ratio and Redis latency. For strong cross-node ordering of invalidations, pub/sub isn't enough — use Streams or accept eventual consistency bounded by TTL. **Decision summary.** Shared L2 only → rely on RedisCacheManager + TTL. Add L1 for latency → add pub/sub (or keyspace-notification) invalidation *plus* an L1 TTL backstop. Always add a CacheErrorHandler so Redis is a performance dependency, not an availability one.
- Why isn't pub/sub alone a safe invalidation mechanism for local near-caches?Because it's fire-and-forget: a node that's restarting or partitioned when the eviction is broadcast never receives it and serves stale data. You need a bounded L1 TTL (or a durable channel like Streams) as a backstop.
- Does @Cacheable(sync=true) prevent a cache stampede across your cluster?No. It only serializes concurrent computation of the same key within a single JVM. Cross-node stampede on a hot expired key still needs a distributed lock or probabilistic early recomputation.
- What's the risk of NOT configuring a CacheErrorHandler?A Redis outage makes cache get/put throw, and that exception propagates out of your business method — so caching turns a nice-to-have into a hard availability dependency.
saying these in an interview costs you the question
- Treating pub/sub invalidation as guaranteed delivery for near-cache consistency
- Assuming a Redis outage silently bypasses the cache by default (it throws unless you add a CacheErrorHandler)
- Believing sync=true coordinates cache filling across all nodes
- Relying only on invalidation broadcasts with no TTL backstop on local caches