skip to content

How does the client-side registry cache work in Eureka, and what staleness does it introduce?

level: middleimportance: should knowfreq 45%

answer

  1. local registry copy = fast, server-outage tolerant
  2. registry-fetch-interval 30s, delta + hash reconcile
  3. server readOnlyCacheMap ~30s refresh
  4. caches stack: server + client + load balancer
  5. route-to-dead-instance is expected -> retries/CB

basics

~20 s

Each Eureka client keeps a local copy of the registry and refreshes it periodically (every 30s by default), usually by fetching only the changes (deltas). Lookups hit this local cache, not the server, so discovery is fast but can be up to ~30s stale.

solid answer

~50 s

Every Eureka client caches the full registry locally so that resolving a service name is an in-memory lookup, not a network call to the server on each request — that keeps discovery fast and lets clients keep routing even if the server is temporarily down. The client refreshes on eureka.client.registry-fetch-interval-seconds (default 30s), normally pulling only a delta of changes since the last fetch and reconciling against the server's expected hash; periodically it does a full fetch. There are actually layered caches: the server itself serves reads from a read-only response cache (refreshed ~30s), the client caches the result, and Spring Cloud LoadBalancer may cache on top. Net effect: a newly registered or evicted instance can take tens of seconds to become visible everywhere, so clients must tolerate calling stale or missing instances via retries and circuit breakers.

code

java · 16 lines
java
// Client-side cache tuning (application.yml)
// eureka:
//   client:
//     fetch-registry: true                 # fetch the registry at all
//     registry-fetch-interval-seconds: 30  # local cache refresh cadence
//     disable-delta: false                 # keep cheap delta fetches (default)

// Server-side read cache (dominant propagation delay for new instances)
// eureka:
//   server:
//     use-read-only-response-cache: true       # default; serve reads from readOnlyCacheMap
//     response-cache-update-interval-ms: 30000 # how often that read cache refreshes

// Total worst-case propagation for a NEW instance to receive traffic:
//   server read-only cache refresh (~30s) + client fetch interval (~30s) + LB cache
//   => can be ~60s+; do NOT assume instant visibility.

go deeper

for a junior

Know the client keeps a local copy refreshed ~every 30s so lookups are local.

for a middle

Explain delta fetches + hash reconciliation and that the cache keeps clients working during a server outage.

for a senior

Enumerate the stacked caches (server read-only, client, load balancer) and where propagation delay actually comes from.

for a principal

Weigh staleness vs server load when tuning for autoscaling, and design consumers to be correct despite stale/missing instances.

## Why cache at all If every service call meant asking the Eureka server 'where is payment-service?', the server would be a hot bottleneck and single point of failure. Instead, **each client downloads the registry and keeps a local copy**, so name resolution is a local, in-memory operation. A bonus: if the server goes down, clients keep routing using their last-known registry. ## The fetch loop - **`eureka.client.registry-fetch-interval-seconds`** (**default 30s**) — how often the client refreshes its local registry. - **Delta fetches**: by default the client fetches only the **delta** (instances that registered/changed/were cancelled since last fetch), not the whole registry — cheaper on the wire. The client applies the delta and compares a **hashcode** against the server's; on mismatch it does a **full fetch** to reconcile. You can disable deltas with `eureka.client.disable-delta: true`. - **`eureka.client.fetch-registry`** — whether the client fetches at all (a standalone server sets this false). ## Layered caching (the real staleness source) Staleness isn't just the client interval. There are **multiple caches in series**: 1. **Server read-only response cache** — the Eureka server does not read from its live map on every request; it serves from a **readOnlyCacheMap** updated from the read-write map on `eureka.server.response-cache-update-interval-ms` (**default ~30s**). So even a fresh registration isn't visible to readers until this refreshes. (You can disable with `eureka.server.use-read-only-response-cache: false`.) 2. **Client registry cache** — refreshed every 30s as above. 3. **Load-balancer cache** — Spring Cloud LoadBalancer can add its own caching layer over the discovery client. Add these up and a change can take **tens of seconds** to propagate everywhere. ## Interaction with eviction Client staleness stacks on top of the server-side lease/eviction lag (up to ~90s + eviction sweep). A crashed instance may live in the server registry *and* in every client's cache well after it died. ## Consequences / gotchas - **You will sometimes route to an instance that's already gone.** This is inherent, not a bug. Mitigate with **Spring Cloud LoadBalancer retries** and a **circuit breaker (Resilience4j)**. - Lowering `registry-fetch-interval-seconds` and disabling the server read-only cache reduces staleness but raises load and CPU on the server. - For fast scale-up (new instance should receive traffic quickly), the read-only response cache interval is often the dominant delay, not the client interval. - Initial registration visibility can also be delayed by `eureka.instance.initial-status` and the client's own registration timing. ## When to tune Defaults favor low load and stability. Tighten only when your workload needs faster propagation (aggressive autoscaling) and you accept the extra server load — and always keep client-side resilience regardless.

  • Why does Eureka fetch deltas instead of the full registry each time?
    To cut bandwidth: the delta contains only instances that changed since the last fetch. The client applies it and compares a hashcode with the server's; if they disagree it falls back to a full fetch to re-sync. Deltas can be disabled with eureka.client.disable-delta if reconciliation problems occur.
  • A new instance registered but isn't getting traffic for ~30-60s. Where's the delay?
    Mostly the layered caches: the server's read-only response cache (refreshed ~30s) plus each client's registry-fetch-interval (30s), and possibly the load balancer's cache. It's not the heartbeat mechanism. Reduce the response-cache and fetch intervals if faster visibility is needed.

saying these in an interview costs you the question

  • Thinking every service lookup calls the Eureka server in real time.
  • Believing a newly registered instance is instantly discoverable everywhere.
  • Not knowing about the server-side read-only response cache as a propagation delay.
  • Assuming the client cache is useless when the server is down (it's actually what keeps routing working).

context