What does self-preservation mode do on a Eureka server, what arithmetic triggers it, and why does it fire so readily in small or non-production deployments?
answer
- it stops evicting, nothing else
- expected renewals versus received
- two per instance per minute, 85%
- a partition and mass death look identical
- ratios are unstable with few instances
basics
~20 sWhen heartbeats received drop below about 85% of what the server expects, Eureka stops evicting expired leases and keeps serving the entries it already has. It is a deliberate choice of stale data over an empty registry under partition.
solid answer
~60 sA Eureka server computes how many renewals per minute it should be receiving — two per registered instance, given 30-second heartbeats — and multiplies that by a threshold, 85% by default. If actual renewals in the last minute fall below that number, the server concludes it is more likely that *the network broke* than that most of the fleet died simultaneously, so it stops evicting expired leases entirely and puts a red warning banner on its dashboard. The reasoning is availability over correctness: on a partition, a mass eviction would empty the registry and take down every service that was still perfectly healthy, so Eureka would rather hand out some addresses that are wrong. The catch is that the trigger is a *ratio*, so it scales badly downward — with three instances, killing one drops renewals by a third, far past the 15% margin, and the server locks up eviction and never removes the dead entry. That is why self-preservation is routinely turned off in development and left on in a real fleet.
go deeper
Recall the headline: when too few heartbeats arrive, the Eureka server stops deleting entries and warns on its dashboard, so the registry may list instances that are gone.
Explain the trigger concretely — expected renewals of two per instance per minute, an 85% threshold — and why a three-instance environment crosses it the moment one instance stops.
Show you have diagnosed it: check whether the server is in self-preservation before chasing stale-entry complaints, alarm on the state, and fix orderly shutdowns so scale-downs never resemble a partition.
Own the tradeoff as a stated position. Argue why a registry that is sometimes wrong beats one that is confidently empty during a partition, and where that argument stops holding — small fleets, or callers that cannot fail over on a refused connection.
## What the mode does Self-preservation is a single, blunt behaviour: **while it is active, the server does not evict expired leases.** Registrations, renewals, and cancellations still work; reads still work. What stops is the automatic removal of instances that have gone quiet. The dashboard shows an unmissable red banner warning that Eureka may be claiming instances are up when they are not. ## The arithmetic of the trigger The server maintains an *expected* renewal rate. With the default 30-second renewal interval, each registered instance should produce **two renewals per minute**, so the expected rate is `2 × instanceCount`. That is multiplied by `renewalPercentThreshold`, default **0.85**, to produce the minimum acceptable rate. Each minute the server compares renewals actually received against that number; falling below it turns self-preservation on, and recovering above it turns it off. Two details matter operationally. The expected count is recomputed on a slow timer — every 15 minutes by default — as well as when instances register or cancel, so after a large planned scale-down the server can carry a stale expectation for a while. And each server in a cluster computes this from **its own** received renewals, so a peer that is partitioned away from most clients enters self-preservation independently. ## Why it exists: the AP choice, made explicit From inside the server, two situations look **identical**: half the fleet crashed, and the network between the fleet and the server broke. In both cases heartbeats stop arriving. Eureka's designers argued that at scale, the second is far more likely than the first — and that the consequences are wildly asymmetric. If the server assumes death and evicts, and it was wrong, it deletes healthy instances from the registry. Callers refresh, find nothing behind those services, and the outage becomes total — caused entirely by the registry, at the exact moment the network was already unhealthy. If instead the server keeps the entries, and it was wrong, callers hold some addresses that no longer answer, and each one fails fast on a refused connection and moves to the next instance. Partial wrongness that the caller can route around beats an authoritative empty answer. That is the whole design in one sentence: **prefer stale data over no data**, and push the correction to the caller, which is the only party that finds out the truth by actually connecting. ## Why it misfires when the fleet is small The trigger is a ratio, and ratios are unstable at small denominators. With 300 instances, losing 20 barely moves the rate and self-preservation stays off, exactly as intended. With three instances, killing one removes 33% of the renewals — twice the 15% margin — and the server flips into self-preservation. Now it will not evict *anything*, including the instance you deliberately stopped, so callers keep getting a dead address indefinitely. Restart-heavy development environments and small staging clusters can sit in this state permanently, which is why developers so often conclude that "Eureka never removes anything". The same effect appears with a genuine planned scale-down in a modest production fleet: shrink from eight instances to four and you have removed half the renewals in one step. ## Operating it - **Turn it off in development and small non-production environments**, via `enableSelfPreservation`. The behaviour it protects against does not matter there, and the confusion it causes does. - **Leave it on in a real fleet**, where the ratio behaves and the partition scenario is real. - **Deregister cleanly on shutdown.** A cancelled lease lowers the expected renewal count as it goes, so orderly scale-down does not look like a partition. Hard kills are what make it look like one. - **Alarm on the mode itself.** "Eureka is in self-preservation" is a first-class operational signal: it means eviction is suspended and the registry's contents are no longer trustworthy. Do not diagnose stale-instance complaints without checking it. - **Do not treat it as health checking.** Self-preservation only decides whether to delete entries; it never verifies that a surviving entry is any good. ## The line to say in an interview Self-preservation is Eureka choosing availability over consistency at the moment the two conflict. It cannot distinguish a dead fleet from a broken network, so it assumes the network — and it degrades gracefully at scale precisely because it degrades badly at small scale, where the ratio it depends on stops being meaningful.
- While a Eureka server is in self-preservation, can new instances still register?Yes. The mode suppresses only the eviction of expired leases; registration, renewal, cancellation and reads all continue normally. A newly started instance appears in the registry as usual — the risk is that it appears alongside stale entries the server has stopped removing.
- Would you disable self-preservation in production to avoid stale entries?Generally no. Disabling it means the first real network partition triggers a mass eviction and empties the registry for services that never failed, converting a network problem into a full outage. Fix stale entries at the source — deregister on shutdown — and make callers fail over on connection errors.
- How does a planned scale-down differ from a partition, from the server's point of view?Only if the instances deregister. A clean cancel lowers the expected renewal count at the same moment it lowers the actual one, so the ratio holds. Instances killed without cancelling leave the expectation high while renewals drop, which is indistinguishable from a partition and can trip the mode.
saying these in an interview costs you the question
- Thinks self-preservation actively checks whether instances are alive
- Says it blocks new registrations while active
- Calls it a bug rather than a deliberate availability choice
- Recommends disabling it fleet-wide in production
- Cannot explain why small deployments trip it constantly