A service instance registered with Eureka is killed abruptly, yet callers keep sending it requests for well over a minute. Which Eureka timers stack up to produce that window, and how would you shorten it?
answer
- nothing detects it — absence is inferred
- four timers, each additive
- eviction sweeps, it does not fire
- server read cache, then client fetch
- deregister on shutdown skips two
basics
~20 sFour delays compound: the 90-second lease must expire, the eviction task runs only every 60 seconds, the server's read-only response cache holds a stale answer for 30 seconds, and each client refetches the registry every 30 seconds.
solid answer
~50 sNothing in Eureka detects the death — it is inferred by absence, and each inference step is on its own timer. The lease has to expire (default 90 seconds of missed heartbeats); the server's eviction task only sweeps expired leases every 60 seconds; the server answers reads from a response cache that refreshes about every 30 seconds; and the caller refetches the registry every 30 seconds and load-balances from that in-memory copy. Worst case those stack to several minutes, and a long-standing quirk in Eureka's lease code effectively doubles the expiry window, which is why people measure three minutes rather than ninety seconds. You shorten it by **deregistering on graceful shutdown** so a cancel is explicit rather than inferred, and by trimming the lease and fetch intervals at the cost of registry chatter. But you cannot tune it to zero: the durable fix is that callers must treat the cached list as a hint — fail fast on connection refused and try the next instance.
go deeper
Know that Eureka infers death from missed heartbeats rather than detecting it, and that both the server and every caller cache the registry, so removal is never instant.
Be able to name the four timers and their defaults in order and add them up, then explain why deregistering on shutdown removes the two largest terms.
Demonstrate the operational instinct: drain before exit, make sure the shutdown path survives the container's grace period, and design callers to fail fast on connection refused and move to the next instance.
Frame the lag as inherent to caching discovery data at the caller, then set a platform policy — standard shutdown behaviour, standard connect timeouts and instance-level failover — instead of letting each team tune intervals to chase a window that cannot reach zero.
## Why the delay exists at all Eureka has no way to learn that an instance died. It infers death from *absence of heartbeats*, and then that inference has to travel through a chain of caches to reach the caller who actually opens the socket. Every hop in that chain is periodic, so the delays add rather than overlap. ## The four clocks, in order **1. The lease must expire.** `leaseExpirationDurationInSeconds` defaults to **90 seconds** against a 30-second renewal interval — three missed beats. Until then the lease is simply considered valid. **2. The eviction sweep must run.** Expired leases are not removed the instant they expire. A background eviction task runs on `evictionIntervalTimerInMs`, default **60 seconds**, and removes whatever has expired since the last pass. On average that adds another half-interval, in the worst case a full one. **3. The server's read cache must turn over.** A Eureka server does not serialise the registry per request; it serves reads from a cached response, whose read-only layer refreshes on `responseCacheUpdateIntervalMs`, default **30 seconds**. So even after eviction, the server can keep handing out the old answer for up to another half-minute. This cache can be disabled to trade CPU for freshness. **4. The client must refetch.** Each caller pulls the registry on `registryFetchIntervalSeconds`, default **30 seconds**. Until it does, the dead address is still in its in-memory list, and whatever load balancer it uses will happily select it. Some client-side load balancers keep their own derived server list with its own refresh, adding a fifth clock. ## The quirk that surprises people Measured windows are usually longer than the arithmetic suggests, because of a well-known bug in Eureka 1.x's lease implementation: renewing a lease stamps the last-update timestamp with *now plus the duration* rather than now, which effectively makes a lease last about **twice** its configured duration. The Eureka source itself carries a comment acknowledging it, and it is left in place because changing it would silently halve every deployment's tolerance for missed heartbeats. Assume roughly 180 seconds of lease, not 90, when you estimate the worst case. ## What to do about it **Deregister explicitly.** The single biggest win is turning an inferred death into a declared one. On graceful shutdown the client should send a cancel to the registry, which removes the lease immediately and skips clocks 1 and 2 entirely — leaving only the caches, a much smaller window. That means your shutdown path must actually run: a container killed with SIGKILL, or a shutdown hook that never gets scheduled because the grace period is too short, silently loses this. Better still, mark the instance out of service first, then let in-flight work drain, then exit. **Trim the intervals, knowingly.** You can shorten the renewal interval and lease duration, shrink the eviction interval, and disable the read-only response cache. Every one of those trades registry load and CPU for freshness, and shortening the lease without shortening the heartbeat makes false evictions likely on any GC pause or network blip. Also note that shortening the lease lowers the volume of heartbeats the server expects, which interacts with self-preservation. **Stop treating the registry as the truth.** This is the answer an interviewer is listening for. In client-side discovery the cached list is *always* a lagging approximation, and no interval setting changes that — it only changes how stale. The caller must survive a stale entry: connect with a short connect timeout, treat connection-refused as "this instance is gone", drop it from the local list, and try the next one. A retry to a *different* instance on a connection-level failure is safe in a way that a retry after a request was already accepted is not. ## The diagnosis in practice When someone reports "we scaled down and got errors for two minutes", check in this order: did the instance deregister (is there a cancel in the server log, or did the pod get SIGKILLed)? Is the entry still visible in `/eureka/apps`? Is the server in self-preservation, in which case it will not evict *at all*? And finally, is the caller's copy stale even though the server is clean — the case that points at fetch intervals rather than at the registry. ## The short version Lease expiry plus eviction sweep plus server response cache plus client fetch, each on its own timer, with a lease-doubling quirk on top. Deregister on shutdown to skip the first two, tune the rest a little, and make the caller resilient because you cannot tune the lag away.
- Why is shortening the lease duration alone a risky tuning change?Because the lease is your tolerance for missed heartbeats. Cut it to 30 seconds while renewals stay at 30 and a single dropped packet, GC pause, or brief registry hiccup expires a perfectly healthy instance. Shorten the renewal interval first, keep the lease at roughly three intervals, and accept the extra heartbeat traffic.
- Your instance deregisters cleanly on shutdown. Can a caller still hit it?Yes, briefly. Cancelling removes the lease at once, but the server's read cache and every caller's registry cache still hold the old address for up to their refresh intervals. That is why you drain first — mark out of service, wait out roughly one fetch interval, then stop accepting and exit.
- How would you confirm whether the stale entry is on the server or only in the caller?Read the registry directly with a GET on the application's path and see whether the instance is still listed. If the server is clean but the caller still dials the dead address, the lag is in that client's cached copy or its load balancer's derived list; if the server still lists it, look at eviction and at whether self-preservation is suppressing eviction entirely.
saying these in an interview costs you the question
- Says Eureka removes an instance the moment it stops responding
- Quotes only the 90-second lease and ignores the caches
- Thinks the server actively probes instances to detect death
- Proposes shortening the lease without shortening the heartbeat
- Assumes clean deregistration makes the address invisible immediately