How do service registries like Consul and Eureka determine that an instance is unhealthy and stop routing traffic to it? Contrast active health checks with heartbeat/lease-based detection.
answer
- Consul = active probes (HTTP/TCP/TTL/script)
- Eureka = heartbeat/lease renewal ~30s/90s
- consecutive-failure threshold avoids flapping
- Eureka self-preservation = stop evicting on mass failure
- heartbeat thread alive != app healthy
basics
~20 sConsul actively pings each instance, and Eureka's instances periodically say 'I'm alive.' If pings fail or the 'I'm alive' messages stop arriving within a set time window, the registry marks that instance as down and stops handing out its address.
solid answer
~50 sThere are two complementary mechanisms. Active health checks (Consul's default model) have the registry or its agent proactively probe each instance - HTTP GET, TCP connect, TTL check, or a custom script - on an interval, and mark it critical after a configured number of consecutive failures. Heartbeat/lease-based detection (Eureka's default model) flips the direction: each instance sends a periodic renewal ('I'm still here') to the registry, and if the registry doesn't receive a renewal within a lease-expiry window (Eureka's default is roughly 90 seconds, tunable), it evicts the instance. Consul supports both styles via its TTL check type. The key operational difference is who bears the cost and blast radius of failure: active checks scale probe volume with fleet size and detect crashes fast but add load; heartbeats are cheap for the registry but a client that hangs without crashing may still renew its lease, and a widescale network blip can trigger mass evictions - which is exactly what Eureka's self-preservation mode is designed to guard against.
go deeper
Understands that unhealthy instances get removed from the pool somehow, even if hazy on the exact mechanism.
Can explain active checks vs heartbeats and name Consul/Eureka as examples of each default style.
Discusses the failure-threshold tuning trade-off and the gap between 'heartbeat alive' and 'app actually healthy.'
Connects mass-failure detection to CAP trade-offs (Eureka self-preservation) and reasons about registry-level blast radius during infra-wide incidents.
## What a registry has to guarantee A service registry's entire value proposition rests on one guarantee: that when a caller asks 'give me healthy instances of service X,' the answer doesn't include instances that are actually down. Getting that guarantee right requires a failure-detection mechanism, and the two dominant designs - **active health checking** and **passive heartbeat/lease renewal** - trade off differently on cost, speed, and false-positive risk. ## Active health checking - the Consul default Active health checking, the model Consul defaults to, puts the registry (or a local agent) in the driver's seat. A Consul agent running alongside each service is configured with one or more checks, each run on an interval, e.g. every 10 seconds: - an **HTTP check** that expects a 200 from a health endpoint; - a **TCP check** that just verifies a socket connects; - a **script check** that runs an arbitrary command; - a **TTL check** (a hybrid, described below). If a check fails some configured number of times in a row (to avoid flapping on a single blip), the agent marks the service instance's check as critical, and Consul's catalog stops returning that instance in DNS or API responses for service lookups. The strength of this approach is that it directly tests what actually matters - can something reach this instance and get a good response - rather than just whether the process is technically running; it also detects failures proactively, without waiting on the failing instance to notice anything itself, so hard crashes and hung processes are caught by the same mechanism. The cost is that the registry (or its agents) must actively fan out probes to every instance of every service, load that grows with fleet size, and a health check endpoint itself must be lightweight or it becomes a liability. ## Heartbeat and lease renewal - the Eureka default Heartbeat (lease) based detection, the model Eureka defaults to, inverts the direction: instances are responsible for proving their own liveness by sending a periodic renewal request to the registry - by default every 30 seconds in Eureka - and the registry runs a background eviction task that removes any instance whose lease hasn't been renewed within an expiry window (default roughly 90 seconds, i.e. about 3 missed heartbeats). This is cheap for the registry - it just needs to track a last-renewal timestamp per instance and sweep periodically - but it has a subtler failure mode: a heartbeat only proves the instance's heartbeat thread is alive, not that the instance can actually serve real requests, so an app that's deadlocked on its main request path but whose heartbeat loop runs on a separate thread can keep renewing its lease while being functionally dead to callers. It's common to pair Eureka-style heartbeats with an actual liveness/readiness endpoint checked separately (e.g. by a load balancer) to close that gap. ## Consul's TTL check - a deliberate hybrid Consul's `TTL` check type is a deliberate hybrid: instead of the agent probing the instance, the instance itself calls back into Consul on an interval to say 'I'm healthy,' and if that callback doesn't arrive before the TTL expires, Consul marks it critical - this is functionally a heartbeat but implemented on top of Consul's active-check framework, useful when an instance wants to self-report application-level health (e.g. 'I'm technically up but my DB connection pool is exhausted') rather than exposing an HTTP endpoint for an agent to poll. ## The shared failure mode - eviction of a healthy fleet The production failure mode both approaches share, and that a good answer should flag, is the false-mass-eviction scenario: a network partition, a GC pause epidemic, or an overloaded registry can make many healthy instances simultaneously appear unhealthy - all their heartbeats or health checks fail at once, not because the instances are down but because the detection channel itself is degraded. - **Eureka's well-known self-preservation mode** exists precisely for this: if the rate of expiring leases crosses a threshold suggesting a systemic problem rather than genuine instance failure, Eureka stops evicting and keeps serving its last-known-good registry (favoring availability over strict correctness, a concrete CAP-theorem trade-off), accepting that some stale/dead entries may be returned rather than risk evicting the entire healthy fleet because the network hiccupped. - **Consul**, being Raft-based and consistency-oriented, doesn't have an equivalent stop-evicting mode by default, so a bad health-check storm can genuinely drain a service's registered instance count to zero even if the instances are fine - which is why check design (sane failure thresholds, checks that don't depend on the same infra that's flaking) matters as much as picking a tool.
- Why might a process that's actually deadlocked or unable to serve requests still pass an Eureka heartbeat check?Because the heartbeat renewal is often driven by a separate, lightweight background thread or client library loop that's independent of the request-handling path. If that thread keeps running while the main application logic is stuck (e.g. on a database deadlock), the instance keeps renewing its lease and looks alive to the registry even though it can't actually serve traffic - which is why teams pair heartbeats with real readiness/liveness probes rather than relying on them alone.
- What is Eureka's self-preservation mode and what trade-off does it represent?It's a safety mechanism where Eureka stops evicting instances once the rate of expiring leases crosses a threshold, on the theory that a network-wide problem (not mass instance death) is the likely cause. It trades correctness (the registry may keep serving addresses of genuinely dead instances) for availability (the registry doesn't empty itself out and cause total service unavailability during a network blip) - a direct, textbook CAP-theorem availability-over-consistency choice.
Active health checks are like a manager walking the floor every 10 minutes checking each employee is at their desk. Heartbeat/lease checks are like each employee having to badge in every 30 minutes to prove they're still on-site - the office assumes you've left if you stop badging in.
saying these in an interview costs you the question
- thinks heartbeats and health checks are the same thing with no distinction
- doesn't know any concrete tool defaults or that they're configurable
- assumes a passing heartbeat guarantees the app can serve real traffic
- unaware that mass health-check failures can be a symptom of network problems, not instance death