How do Consul health checks work with a Spring Boot app, and what happens to an instance that becomes unhealthy?
answer
- agent polls /actuator/health
- passing vs warning vs critical
- unhealthy = excluded, not deregistered
- critical-timeout = auto-reap
- TTL heartbeat = app pushes
basics
~10 sBy default Spring registers an HTTP check hitting the app's Actuator health endpoint. The agent polls it periodically; if it fails, Consul marks the instance unhealthy and discovery stops routing traffic to it.
solid answer
~40 sWhen a Spring app registers, Spring Cloud Consul also registers a **health check**. The default is an HTTP check that the local agent periodically calls against the Actuator `/actuator/health` endpoint at a configurable interval. Consul aggregates check status into passing/warning/critical; only *passing* (healthy) instances are returned by discovery queries, so an unhealthy instance is automatically excluded from load balancing without being deregistered. Properties like `spring.cloud.consul.discovery.health-check-path`, `health-check-interval`, `health-check-timeout`, and `health-check-critical-timeout` tune this — the last controls how long an instance may stay critical before Consul deregisters it entirely. Alternatively you can use a **TTL check** (`heartbeat.enabled=true`), where the app must actively ping Consul within a window or be marked critical. HTTP checks are agent-driven (pull); TTL checks are app-driven (push).
code
yaml · 20 linesspring:
cloud:
consul:
discovery:
health-check-path: /actuator/health
health-check-interval: 10s
health-check-timeout: 5s
health-check-critical-timeout: 30s # auto-deregister after 30s critical
query-passing: true
# Alternative push model:
# heartbeat:
# enabled: true
management:
endpoint:
health:
show-details: never
endpoints:
web:
exposure:
include: healthgo deeper
Know the app registers a health check that Consul polls and unhealthy instances stop receiving traffic.
Explain HTTP vs TTL checks, the /actuator/health default, and critical-timeout reaping.
Discuss health-group scoping to avoid cascading removals and the pull-vs-push trade-offs.
Design health semantics for a fleet: liveness vs readiness exposure, agent reachability, reaping policy, and blast-radius control.
**Why health checks matter.** Service discovery is only useful if it routes to *live* instances. Consul separates *registration* (I exist) from *health* (I am currently serving). Discovery queries return only healthy instances, so a crashed or overloaded instance is silently removed from the pool. **How Spring registers the check.** With `spring-cloud-starter-consul-discovery`, `ConsulAutoServiceRegistration` builds an `NewService` that includes an `NewService.Check`. By default this is an **HTTP check**: the *local Consul agent* issues an HTTP GET against your app, at the path given by `spring.cloud.consul.discovery.health-check-path` (default `/actuator/health`, assuming Spring Boot Actuator is on the classpath and the endpoint is exposed). **Check states.** A Consul check is `passing`, `warning`, or `critical`. Discovery (`ConsulDiscoveryClient` and the catalog `/health/service` API used under the hood) returns only instances whose checks are passing (you can opt to include others, but the default is healthy-only via `query-passing`). So an unhealthy app *stays registered* but is *excluded from results* — no traffic is routed to it. **Key properties** (`spring.cloud.consul.discovery.*`): - `health-check-path` — the endpoint the agent polls (default `/actuator/health`). - `health-check-interval` — how often (e.g. `10s`). - `health-check-timeout` — per-request timeout. - `health-check-tls-skip-verify` — for HTTPS endpoints with self-signed certs. - `health-check-critical-timeout` — how long an instance may remain *critical* before Consul **automatically deregisters** it (a.k.a. reaping). Without it, dead instances linger as critical forever. - `query-passing` — whether discovery filters to passing instances only (typically true). **Actuator interaction.** The `/actuator/health` endpoint aggregates health indicators (DB, disk, custom). If a downstream dependency is down, the endpoint can return `DOWN` (HTTP 503), which Consul reads as critical — potentially removing the instance. This is powerful but dangerous: an overly strict health group (e.g. failing because a non-critical cache is down) can cascade an entire fleet out of discovery. Use Actuator health groups to expose a *liveness*-style check to Consul distinct from full readiness. **TTL (heartbeat) checks.** Setting `spring.cloud.consul.discovery.heartbeat.enabled=true` switches to a **TTL check**: instead of the agent pulling, the app *pushes* a heartbeat to Consul (via the agent's check-pass API) on a schedule. If Consul doesn't hear within the TTL, the check goes critical. TTL suits apps behind NAT/firewalls where the agent can't reach the app's HTTP port, or apps that want to self-report readiness. The trade-off: a hung app that can still send heartbeats may look healthy falsely. **Gotchas.** - Forgetting Actuator on the classpath (or not exposing `health`) makes the default HTTP check 404/500 → instance marked critical → invisible to discovery. - Not setting `health-check-critical-timeout` leaves zombie critical entries. - Tying the Consul-facing check to a strict readiness group can amplify partial outages. - Health checks run from the *agent*, so agent-to-app network reachability matters, not client-to-app. **When to use which.** Prefer HTTP checks against Actuator for normal HTTP services (simple, agent-driven). Use TTL/heartbeat when the agent cannot reach the app or you want the app to self-assert liveness.
- An instance's /actuator/health returns DOWN because a non-critical cache is unreachable. What happens and how do you prevent a fleet-wide outage?Consul marks every instance critical and removes them from discovery, breaking the whole service. Fix by exposing a narrower health group (e.g. liveness) to Consul via health-check-path, so only truly fatal conditions fail the check.
- Difference between an unhealthy instance and a deregistered one?Unhealthy (critical) instances remain in the catalog but are filtered out of discovery queries. Deregistration removes them entirely — done on graceful shutdown, or automatically after health-check-critical-timeout.
saying these in an interview costs you the question
- Thinking an unhealthy instance is immediately removed from the catalog
- Believing the client (not the agent) performs the HTTP health check
- Assuming health checks work without Actuator on the classpath