Explain the composite structure behind /actuator/health: how the overall status is aggregated, how health groups and readiness/liveness probes fit in, and what edge cases you must handle.
answer
- HealthContributor tree: indicators + composites
- StatusAggregator: worst wins (DOWN>OUT_OF_SERVICE>UP>UNKNOWN)
- groups -> /health/{group}, own aggregator+mapper
- liveness=restart / readiness=de-route (K8s)
- never put external deps in liveness; bound timeouts
basics
~20 sThe health endpoint composes many HealthIndicator/HealthContributor beans into a tree. A StatusAggregator reduces their statuses to one overall status (by default the worst wins, e.g. any DOWN → overall DOWN). Health groups (like readiness, liveness) expose subsets at /actuator/health/{group} for Kubernetes probes.
solid answer
~40 s`HealthEndpoint` builds a composite: every `HealthContributor` bean (a `HealthIndicator`, or a nested `CompositeHealthContributor`) forms a tree of named components. Each indicator's `health()` yields a `Status`; a `StatusAggregator` reduces the children to the overall status using a severity order — by default `DOWN` > `OUT_OF_SERVICE` > `UP` > `UNKNOWN`, so the worst status wins. A `HttpCodeStatusMapper` then maps that to an HTTP code (UP→200, DOWN→503). **Health groups** let you expose a named subset at `/actuator/health/{group}` with its own included indicators, aggregator, and HTTP mapping. Boot pre-defines **`liveness`** and **`readiness`** groups backed by `LivenessStateHealthIndicator` and `ReadinessStateHealthIndicator` (`ApplicationAvailability`), auto-registered on Kubernetes (or via property) at `/actuator/health/liveness` and `/actuator/health/readiness`. Edge cases: a slow/blocking indicator can stall the whole check; a flapping dependency shouldn't fail liveness (only readiness); and details/component visibility must be controlled to avoid leaking internals.
code
java · 33 lines// Drive readiness from application code via AvailabilityChangeEvent.
// (Liveness should reflect ONLY internal correctness; readiness gates traffic.)
import org.springframework.boot.availability.AvailabilityChangeEvent;
import org.springframework.boot.availability.ReadinessState;
import org.springframework.context.ApplicationEventPublisher;
import org.springframework.stereotype.Component;
@Component
public class WarmupGate {
private final ApplicationEventPublisher events;
public WarmupGate(ApplicationEventPublisher events) { this.events = events; }
// While caches warm, refuse traffic -> /actuator/health/readiness = 503,
// Kubernetes stops routing (does NOT restart the pod).
public void beginWarmup() {
AvailabilityChangeEvent.publish(events, this, ReadinessState.REFUSING_TRAFFIC);
}
public void warmupComplete() {
AvailabilityChangeEvent.publish(events, this, ReadinessState.ACCEPTING_TRAFFIC);
}
}
// A custom StatusAggregator making a custom FATAL status outrank DOWN:
import org.springframework.boot.actuate.health.SimpleStatusAggregator;
import org.springframework.boot.actuate.health.StatusAggregator;
import org.springframework.context.annotation.Bean;
class HealthConfig {
@Bean
StatusAggregator statusAggregator() {
return new SimpleStatusAggregator("FATAL", "DOWN", "OUT_OF_SERVICE", "UP", "UNKNOWN");
}
}go deeper
Know health combines several checks into one status and that any failing check can make it DOWN.
Explain HealthIndicator aggregation and that groups expose subsets like readiness/liveness.
Detail StatusAggregator severity ordering, HttpCodeStatusMapper, and the readiness/liveness state model.
Design probe strategy: liveness=internal-only(restart) vs readiness=deps(de-route), bounded-timeout indicators, criticality via groups/aggregators, graceful-shutdown draining, and info-leak control.
The health endpoint is not a single check — it's a **composite tree** with a reduction function, and understanding that is what separates a principal-level answer. **The contributor model.** The building block is `HealthContributor`. There are two kinds: - **`HealthIndicator`** — a leaf: `Health health()` returns a `Status` plus a details map. Boot auto-configures many (`DataSourceHealthIndicator`, `DiskSpaceHealthIndicator`, `RedisHealthIndicator`, `MongoHealthIndicator`, `RabbitHealthIndicator`, `PingHealthIndicator`, `MailHealthIndicator`, …), one per detected dependency. - **`CompositeHealthContributor`** — an internal node grouping named children, so health is a *tree*, not a flat list. Each bean's name (indicator bean id minus the `HealthIndicator` suffix) becomes a key under `components`. (There are reactive analogues `ReactiveHealthIndicator`/`ReactiveHealthContributor` for WebFlux.) **Status and aggregation.** `Status` is an open type; the four standard values are `UP`, `DOWN`, `OUT_OF_SERVICE`, `UNKNOWN` (you may add custom ones). The overall status is computed by a **`StatusAggregator`** (default `SimpleStatusAggregator`) which orders statuses by severity — default order `DOWN`, `OUT_OF_SERVICE`, `UP`, `UNKNOWN` — and picks the most severe present. So **one `DOWN` child drives the whole endpoint `DOWN`**. You can customize the order (e.g. treat a custom `FATAL` as most severe) or provide your own aggregator. **HTTP mapping.** A **`HttpCodeStatusMapper`** (default `SimpleHttpCodeStatusMapper`) maps the final status to an HTTP code: `UP`→200, `DOWN`/`OUT_OF_SERVICE`→503, others→200. Probes and load balancers rely on the code, not the JSON. **Detail visibility.** `show-details` (never/when-authorized/always) and `show-components` control whether the `components` breakdown is rendered. Default `never` → `{"status":"UP"}`. This is a deliberate info-leak guard: the component list reveals your infrastructure. **Health groups.** A **group** is a named, independently-configured view. You define which indicators it `include`s/`exclude`s, and it can have its own `StatusAggregator`, `HttpCodeStatusMapper`, and `show-details`. It's served at **`/actuator/health/{groupName}`**. Example: a `db` group with only the datasource indicator, or a `custom` group for a specific dependency an LB should check. **Liveness & readiness (Kubernetes).** Boot models **application availability** via `ApplicationAvailability`, with two states: - **Liveness** (`LivenessState`: `CORRECT`/`BROKEN`) — 'is the app's internal state healthy, i.e. should it be restarted?' Backed by `LivenessStateHealthIndicator`. A `BROKEN` liveness should trigger a **restart**. - **Readiness** (`ReadinessState`: `ACCEPTING_TRAFFIC`/`REFUSING_TRAFFIC`) — 'can it serve requests right now?' Backed by `ReadinessStateHealthIndicator`. `REFUSING_TRAFFIC` should make the platform **stop routing** to the pod (but not restart it). Boot pre-registers **`liveness`** and **`readiness`** groups exposing these at `/actuator/health/liveness` and `/actuator/health/readiness`, **auto-enabled when running on Kubernetes** (detected) or when `management.endpoint.health.probes.enabled=true`. State transitions are driven by `AvailabilityChangeEvent`s (Boot publishes readiness `ACCEPTING_TRAFFIC` when the context is fully started, and `REFUSING_TRAFFIC` during graceful shutdown), and your code can publish them too via `ApplicationEventPublisher`. **Critical edge cases / gotchas (principal focus).** 1. **Liveness must not depend on external systems.** A common fatal mistake is including a DB/Redis check in the *liveness* group: when that dependency blips, Kubernetes **restarts** healthy pods, amplifying an outage into a crash loop. External dependencies belong in **readiness** (stop routing) — never liveness. 2. **Slow/blocking indicators stall the whole check.** Indicators run and a hung one (e.g. a socket with no timeout) can make `/actuator/health` time out, itself tripping probes. Give every indicator a bounded timeout; consider caching or async for expensive checks. 3. **Worst-wins can be too aggressive.** A non-critical optional dependency being `DOWN` shouldn't necessarily fail the whole app's readiness — model it as its own group or a custom aggregator/status so criticality is expressed correctly. 4. **Info leakage.** Full component details expose infra topology; keep `show-details`/`show-components` restricted in prod. 5. **Probe vs LB semantics differ.** Liveness = restart, readiness = de-route; conflating them causes either missed restarts or unnecessary restarts. 6. **Graceful shutdown ordering.** With graceful shutdown, readiness flips to `REFUSING_TRAFFIC` before the server stops, letting the LB drain — relying on liveness for this would restart instead of drain. **When to use groups/probes.** On Kubernetes, always split liveness (internal only) from readiness (dependencies + startup gating). Use custom groups when a specific downstream (or an LB) needs a tailored health view distinct from the aggregate.
- Why is it a mistake to include a database HealthIndicator in the liveness probe group?Liveness failure tells Kubernetes to RESTART the pod. If a transient DB outage flips liveness to BROKEN, K8s restarts otherwise-healthy pods, turning a dependency blip into a restart storm/crash loop. DB and other external checks belong in the readiness group, which only stops routing traffic — the correct response to a downstream being unavailable.
- How does the overall health status get computed when indicators disagree?A StatusAggregator reduces the child statuses using a severity ordering (default DOWN > OUT_OF_SERVICE > UP > UNKNOWN) and returns the most severe present — so any single DOWN makes the aggregate DOWN. The order (and thus which status 'wins') is customizable, including for custom statuses.
- What could make /actuator/health itself time out, and how do you prevent it?A blocking indicator with no timeout (e.g. a hung socket to a downstream) — the endpoint runs indicators and waits on them. Prevent it by giving every check a bounded timeout, caching expensive results, or using reactive/async indicators, so one slow dependency can't stall the whole probe and trip Kubernetes.
saying these in an interview costs you the question
- Putting external-dependency checks in the liveness group
- Thinking liveness and readiness are interchangeable (both just 'is it up')
- Assuming overall status is UP unless every indicator is DOWN (it's worst-wins: any DOWN → DOWN)
- Ignoring that a blocking indicator can stall the whole health endpoint
- Exposing full component details in production