skip to content

What is Eureka's self-preservation mode, why does it exist, and what problems can it cause?

level: seniorimportance: should knowfreq 50%

answer

  1. guards against network partition -> don't evict everyone
  2. renewals/min: 2 per instance, threshold 0.85
  3. below threshold -> stop evicting leases
  4. cost: keeps dead instances during real outage
  5. enable-self-preservation default true; off in dev

basics

~20 s

If the server suddenly stops receiving enough heartbeats (below ~85% of expected), it assumes a network problem rather than mass instance death and stops evicting instances — keeping them registered to avoid wiping out a healthy registry. The downside: it can keep dead instances listed during a real outage.

solid answer

~40 s

Self-preservation is a safety net against network partitions. Eureka computes the number of heartbeat renewals it expects per minute (from the instance count and 30s heartbeat interval) and a threshold, eureka.server.renewal-percent-threshold (default 0.85). If actual renewals in the last window drop below that threshold, the server assumes the instances are probably still alive but unreachable (a partition) rather than genuinely dead, and it stops evicting expired leases — protecting the registry from being emptied by a transient network blip. The cost: during a genuine mass failure it keeps stale instances registered, so clients keep trying dead endpoints. It's controlled by eureka.server.enable-self-preservation (default true). In dev/small clusters, where instances come and go, it fires spuriously, so people often disable it; in production it's usually left on but sized carefully.

code

java · 16 lines
java
// Eureka SERVER config (application.yml)
// eureka:
//   server:
//     enable-self-preservation: true        # master switch (default true)
//     renewal-percent-threshold: 0.85       # trip when actual renewals < 85% of expected
//     eviction-interval-timer-in-ms: 60000  # eviction cadence when NOT self-preserving
//     renewal-threshold-update-interval-ms: 900000 # recompute expected-renewals baseline

// Expected renewals per minute (heartbeat interval = 30s -> 2/min per instance):
//   expectedRenewals   = 2 * numberOfRegisteredInstances
//   minRenewalsAllowed = expectedRenewals * renewalPercentThreshold  // 0.85
// If renewals-received-in-last-minute < minRenewalsAllowed -> self-preservation ON
//   -> server STOPS evicting expired leases (dead instances stay listed)

// Common dev override to avoid spurious tripping in tiny clusters:
// eureka.server.enable-self-preservation: false

go deeper

for a junior

Know it stops the server from deleting instances when too few heartbeats arrive.

for a middle

State the 85% renewal threshold and that it exists to survive network blips, with dev-disable as a common workaround.

for a senior

Explain the renewals-per-minute math, the availability-vs-correctness trade-off, and that it blocks eviction only.

for a principal

Reason about the AP posture, HA per-node evaluation, threshold recomputation under autoscaling, and why disabling in prod risks cascading outages.

## The failure it guards against Eureka is an **AP** system (availability over consistency). Imagine a network partition cuts the Eureka server off from many clients: heartbeats stop arriving, leases expire, and the naive behavior — **evict everything** — would empty the registry and take down the whole mesh, even though the services are actually fine. **Self-preservation** exists to prevent that catastrophic false-positive. ## How it decides Eureka tracks **renewals per minute**: - Each instance sends a heartbeat every `lease-renewal-interval-in-seconds` (default 30s) → **2 renewals/minute** per instance. - **Expected renewals** ≈ `2 * numberOfInstances`. - **`eureka.server.renewal-percent-threshold`** (**default 0.85**) sets the minimum acceptable fraction. The server derives an **expected minimum renewals** count. If the **actual renewals received in the last minute fall below that threshold**, the server enters **self-preservation mode**: it **stops expiring/evicting leases** until renewals recover. The dashboard shows the red banner: *'EMERGENCY! EUREKA MAY BE INCORRECTLY CLAIMING INSTANCES ARE UP WHEN THEY'RE NOT.'* ## The trade-off - **Benefit:** a transient network glitch or a Eureka server hiccup won't wipe out a valid registry — instances survive the blip. - **Cost:** during a **real** widespread outage, self-preservation **keeps dead instances registered**, so consumers keep routing to endpoints that are actually down. It trades correctness for availability. ## Why it misbehaves in small/dev clusters The threshold math assumes a large, stable population. With only a couple of instances, a single legitimate shutdown can drop renewals below 85%, tripping self-preservation and **pinning dead instances in the registry**. Hence in development people frequently set `eureka.server.enable-self-preservation: false`. ## Config levers - **`eureka.server.enable-self-preservation`** (default `true`) — master switch. - **`eureka.server.renewal-percent-threshold`** (default `0.85`) — raise to make it trip more easily, lower to make it more tolerant. - **`eureka.server.eviction-interval-timer-in-ms`** — how often eviction runs when *not* self-preserving. - **`eureka.server.renewal-threshold-update-interval-ms`** — how often the expected-renewals baseline is recomputed (matters when instances scale up/down). ## Gotchas - **Disabling it in production is dangerous** — a brief network partition can then empty your registry and cascade an outage. Prefer leaving it on and relying on client-side resilience for the stale-instance cost. - Self-preservation blocks **eviction**, not **registration or cancellation** — new instances still register, graceful cancels still remove. - It's a **per-server** state; in an HA peer cluster each node evaluates independently. - It interacts with heartbeat tuning: if you shorten `lease-renewal-interval`, the expected-renewals math changes and the `renewal-threshold-update-interval` must recompute for the threshold to stay meaningful. ## When to disable vs keep Keep it **on** in real multi-instance production; the stale-instance cost is handled by client retries + circuit breakers. Turn it **off** only in dev/CI or tiny clusters where spurious tripping is more harmful than a transient wipe.

  • Why is self-preservation often disabled in development but kept on in production?
    In dev/tiny clusters, one legitimate instance shutdown can drop renewals below 85% and trip it, pinning dead instances in the registry. In production with many instances, it protects against a real network partition wiping the whole registry — the stale-instance cost is absorbed by client-side retries and circuit breakers.
  • Does self-preservation stop new instances from registering?
    No. It only stops the eviction of expired leases. Registration and graceful cancellation still work; the server just won't proactively remove instances whose heartbeats stopped, on the assumption they're partitioned rather than dead.

saying these in an interview costs you the question

  • Saying self-preservation deletes instances (it does the opposite — it stops eviction).
  • Claiming it improves consistency (it sacrifices consistency for availability).
  • Recommending disabling it in production without acknowledging the partition-wipe risk.
  • Not knowing the 0.85 threshold or that heartbeats count as 2/min per instance.

context