skip to content

How do Consul health checks work with a Spring Boot app, and what happens to an instance that becomes unhealthy?

level: middleimportance: should knowfreq 55%

answer

  1. agent polls /actuator/health
  2. passing vs warning vs critical
  3. unhealthy = excluded, not deregistered
  4. critical-timeout = auto-reap
  5. TTL heartbeat = app pushes

basics

~10 s

By default Spring registers an HTTP check hitting the app's Actuator health endpoint. The agent polls it periodically; if it fails, Consul marks the instance unhealthy and discovery stops routing traffic to it.

solid answer

~40 s

When a Spring app registers, Spring Cloud Consul also registers a **health check**. The default is an HTTP check that the local agent periodically calls against the Actuator `/actuator/health` endpoint at a configurable interval. Consul aggregates check status into passing/warning/critical; only *passing* (healthy) instances are returned by discovery queries, so an unhealthy instance is automatically excluded from load balancing without being deregistered. Properties like `spring.cloud.consul.discovery.health-check-path`, `health-check-interval`, `health-check-timeout`, and `health-check-critical-timeout` tune this — the last controls how long an instance may stay critical before Consul deregisters it entirely. Alternatively you can use a **TTL check** (`heartbeat.enabled=true`), where the app must actively ping Consul within a window or be marked critical. HTTP checks are agent-driven (pull); TTL checks are app-driven (push).

code

yaml · 20 lines
yaml
spring:
  cloud:
    consul:
      discovery:
        health-check-path: /actuator/health
        health-check-interval: 10s
        health-check-timeout: 5s
        health-check-critical-timeout: 30s  # auto-deregister after 30s critical
        query-passing: true
        # Alternative push model:
        # heartbeat:
        #   enabled: true
management:
  endpoint:
    health:
      show-details: never
  endpoints:
    web:
      exposure:
        include: health

go deeper

for a junior

Know the app registers a health check that Consul polls and unhealthy instances stop receiving traffic.

for a middle

Explain HTTP vs TTL checks, the /actuator/health default, and critical-timeout reaping.

for a senior

Discuss health-group scoping to avoid cascading removals and the pull-vs-push trade-offs.

for a principal

Design health semantics for a fleet: liveness vs readiness exposure, agent reachability, reaping policy, and blast-radius control.

**Why health checks matter.** Service discovery is only useful if it routes to *live* instances. Consul separates *registration* (I exist) from *health* (I am currently serving). Discovery queries return only healthy instances, so a crashed or overloaded instance is silently removed from the pool. **How Spring registers the check.** With `spring-cloud-starter-consul-discovery`, `ConsulAutoServiceRegistration` builds an `NewService` that includes an `NewService.Check`. By default this is an **HTTP check**: the *local Consul agent* issues an HTTP GET against your app, at the path given by `spring.cloud.consul.discovery.health-check-path` (default `/actuator/health`, assuming Spring Boot Actuator is on the classpath and the endpoint is exposed). **Check states.** A Consul check is `passing`, `warning`, or `critical`. Discovery (`ConsulDiscoveryClient` and the catalog `/health/service` API used under the hood) returns only instances whose checks are passing (you can opt to include others, but the default is healthy-only via `query-passing`). So an unhealthy app *stays registered* but is *excluded from results* — no traffic is routed to it. **Key properties** (`spring.cloud.consul.discovery.*`): - `health-check-path` — the endpoint the agent polls (default `/actuator/health`). - `health-check-interval` — how often (e.g. `10s`). - `health-check-timeout` — per-request timeout. - `health-check-tls-skip-verify` — for HTTPS endpoints with self-signed certs. - `health-check-critical-timeout` — how long an instance may remain *critical* before Consul **automatically deregisters** it (a.k.a. reaping). Without it, dead instances linger as critical forever. - `query-passing` — whether discovery filters to passing instances only (typically true). **Actuator interaction.** The `/actuator/health` endpoint aggregates health indicators (DB, disk, custom). If a downstream dependency is down, the endpoint can return `DOWN` (HTTP 503), which Consul reads as critical — potentially removing the instance. This is powerful but dangerous: an overly strict health group (e.g. failing because a non-critical cache is down) can cascade an entire fleet out of discovery. Use Actuator health groups to expose a *liveness*-style check to Consul distinct from full readiness. **TTL (heartbeat) checks.** Setting `spring.cloud.consul.discovery.heartbeat.enabled=true` switches to a **TTL check**: instead of the agent pulling, the app *pushes* a heartbeat to Consul (via the agent's check-pass API) on a schedule. If Consul doesn't hear within the TTL, the check goes critical. TTL suits apps behind NAT/firewalls where the agent can't reach the app's HTTP port, or apps that want to self-report readiness. The trade-off: a hung app that can still send heartbeats may look healthy falsely. **Gotchas.** - Forgetting Actuator on the classpath (or not exposing `health`) makes the default HTTP check 404/500 → instance marked critical → invisible to discovery. - Not setting `health-check-critical-timeout` leaves zombie critical entries. - Tying the Consul-facing check to a strict readiness group can amplify partial outages. - Health checks run from the *agent*, so agent-to-app network reachability matters, not client-to-app. **When to use which.** Prefer HTTP checks against Actuator for normal HTTP services (simple, agent-driven). Use TTL/heartbeat when the agent cannot reach the app or you want the app to self-assert liveness.

  • An instance's /actuator/health returns DOWN because a non-critical cache is unreachable. What happens and how do you prevent a fleet-wide outage?
    Consul marks every instance critical and removes them from discovery, breaking the whole service. Fix by exposing a narrower health group (e.g. liveness) to Consul via health-check-path, so only truly fatal conditions fail the check.
  • Difference between an unhealthy instance and a deregistered one?
    Unhealthy (critical) instances remain in the catalog but are filtered out of discovery queries. Deregistration removes them entirely — done on graceful shutdown, or automatically after health-check-critical-timeout.

saying these in an interview costs you the question

  • Thinking an unhealthy instance is immediately removed from the catalog
  • Believing the client (not the agent) performs the HTTP health check
  • Assuming health checks work without Actuator on the classpath

context