Design a production health endpoint strategy covering aggregation severity, HTTP mapping for probes, and safe detail exposure. What are the trade-offs and pitfalls?
answer
- health = machine contract (probes/LB)
- never/when-authorized to avoid info disclosure
- 503 fails readiness (stop traffic) vs liveness (restart)
- split liveness/readiness via groups
- keep flaky deps out of liveness
basics
~20 sKeep show-details at never or when-authorized to avoid leaking infra data; let DOWN/OUT_OF_SERVICE map to 503 so load balancers and probes react; configure status.order and http-mapping consistently for any custom status; and separate liveness from readiness so transient dependency failures don't kill pods.
solid answer
~40 sA robust strategy treats the health endpoint as a machine contract. Security: default show-details to never (or when-authorized behind roles) so component details — DB URLs, disk paths, exception text — aren't disclosed anonymously. Probe integration: rely on the HttpCodeStatusMapper so infrastructure acts on status codes (DOWN/OUT_OF_SERVICE → 503) rather than parsing bodies. Aggregation: understand worst-wins via SimpleStatusAggregator and that a single flaky dependency turning DOWN fails the whole endpoint — which is why you split liveness (is the JVM alive) from readiness (are dependencies OK) using health groups, so a transient DB blip pulls the pod out of rotation without triggering a restart. Custom statuses must be wired in status.order and status.http-mapping consistently or you get silent under-reaction. Also decide whether external dependencies belong in liveness at all.
code
yaml · 16 linesmanagement:
endpoint:
health:
show-details: when-authorized
roles: ADMIN
probes:
enabled: true # exposes liveness & readiness groups
group:
liveness:
include: livenessState # JVM/app only — restart can fix it
readiness:
include: readinessState,db,redis # deps — failure stops traffic
status:
order: FATAL,DOWN,OUT_OF_SERVICE,UP,UNKNOWN
http-mapping:
fatal: 503go deeper
Understand health is consumed by probes/load balancers.
Know to keep details restricted and let DOWN map to 503.
Split liveness/readiness with groups and reason about worst-wins aggregation.
Own the end-to-end contract: security, probe semantics, restart-storm avoidance, and disciplined use of custom statuses.
**Health endpoint is a machine contract.** The primary consumers are load balancers and orchestrators (Kubernetes probes), not humans. Design decisions should optimize for how they behave. **1. Security of detail exposure.** - Default `management.endpoint.health.show-details` to `never`, or `when-authorized` gated by `management.endpoint.health.roles` for operator dashboards. The risk with `always` is information disclosure: component detail maps routinely contain DB connection info, disk paths, broker hosts, and exception messages that aid an attacker. - Consider a separate, unauthenticated **liveness/readiness** surface that reveals only the overall status, and a richer authenticated surface for operators. `show-components` (inherits `show-details`) lets you reveal component *names/statuses* without full detail maps if you want a middle ground. **2. HTTP mapping drives infrastructure.** `SimpleHttpCodeStatusMapper` maps DOWN/OUT_OF_SERVICE → 503, UP/UNKNOWN → 200. This is deliberately the contract probes rely on: a `readinessProbe` failing on 503 stops traffic; a `livenessProbe` failing on 503 **restarts** the pod. So the *meaning* you assign to statuses has operational teeth. If you introduce a custom `FATAL`, you must add `status.http-mapping.fatal=503` — otherwise it silently returns 200 and no probe reacts. **3. Aggregation is worst-wins — and that's a double-edged sword.** `SimpleStatusAggregator` returns the most severe status. If you put a downstream dependency (a third-party API) into the default health group, one transient outage of that API makes the whole endpoint DOWN → 503. On a **liveness** probe that means Kubernetes *restarts* your perfectly healthy pod, which won't fix the external outage and can cause a restart storm. Mitigations: - **Split liveness and readiness via health groups** (`management.endpoint.health.group.liveness.include=...`, `...group.readiness.include=...`). Liveness should include only checks that a restart can fix (the JVM/app itself). Readiness includes dependencies whose failure should merely stop routing traffic. - Give each group its own `StatusAggregator`/order if the severity semantics differ. - Keep flaky external dependencies out of liveness entirely; possibly out of readiness too if their failure shouldn't stop serving cached/degraded responses. **4. Custom statuses: consistency and blast radius.** Adding a status like `DEGRADED` or `FATAL` means wiring both `status.order` (severity) and `status.http-mapping` (code), per group if grouped. Every downstream consumer (dashboards, alerting rules, probes) must understand the new code. Prefer the built-in four unless a distinct operational action is warranted (e.g. FATAL pages on-call and returns 503; DEGRADED stays 200 but fires a warning alert). **5. Pitfalls checklist.** - `show-details: always` on an internet-facing endpoint — information disclosure. - Custom status not in `status.order` (loses aggregation) or not in `http-mapping` (returns 200 silently). - External dependency in the liveness group causing restart storms. - Assuming UNKNOWN is severe — it's least-severe and returns 200. - Forgetting probes act on the **status code**, so body-only changes (show-details) have zero effect on probe behavior. **6. When to keep it simple.** For most services: `show-details: when-authorized`, default aggregator/mapper, dependencies in readiness, JVM-only liveness. Reach for custom statuses or custom aggregator/mapper beans only when the operational model genuinely needs them.
- Why is putting a third-party API check in the liveness group dangerous?Liveness failure (503) makes Kubernetes restart the pod. A restart can't fix an external outage, so you get pointless restart storms. External deps belong in readiness (stop traffic) or nowhere, not liveness.
- How do readiness and liveness differ in how a 503 is treated?Readiness 503 removes the pod from service endpoints (stops routing traffic) but leaves it running; liveness 503 causes the kubelet to restart the container.
- What's the middle-ground option between hiding and fully exposing component details?show-components (which inherits show-details) can reveal component names and statuses while show-details keeps full detail maps restricted — or gate everything behind when-authorized + roles.
saying these in an interview costs you the question
- Exposing show-details=always on a public endpoint
- Putting external dependencies in liveness, causing restart storms
- Adding a custom status without both order and http-mapping
- Believing show-details changes probe behavior (probes use the status code)