skip to content

What makes a container health probe command useful rather than misleading, and what should it deliberately avoid testing?

level: seniorimportance: should knowfreq 50%

answer

  1. Fail only for what a restart of THIS instance could fix
  2. Don't probe downstream deps → correlated fleet-wide unhealthy
  3. Traverse the real serving path, not a special-case handler
  4. Probe runs in-container, charged to its own limits
  5. Distroless has no shell/curl → exec form + tiny health binary

basics

~20 s

A good probe tests that this instance can serve — a cheap, dedicated endpoint that exercises the serving path — with a tight timeout and a binary that actually exists in the image. It should not test downstream dependencies, do real work, or require credentials, because that turns one dependency outage into every container reporting unhealthy.

solid answer

~60 s

Principles I apply: 1. **Test the local instance, not the world.** Probing a database from the app's health check means a database blip marks every replica unhealthy — correlated failure, and whatever consumes the status then acts on all of them at once. Dependency state belongs in metrics, not in the probe. 2. **Exercise the real serving path.** A check that returns a static 200 from a handler outside the normal stack proves only that the port is open. Route it through the same server, thread pool, and framework so a wedged pool actually shows up. 3. **Keep it cheap and bounded.** It runs every interval forever, inside the container, against its own CPU and memory limits. No heavy queries, no allocations that push a tight limit. 4. **Use something that exists in the image.** `curl` is absent from most minimal bases and distroless has no shell — ship a tiny health binary or use the app's own subcommand, in exec form. 5. **Set a timeout below the interval** so a hang registers quickly, and print a one-line reason so `.State.Health.Log` is diagnostic. 6. **No secrets on the probe command line**, and don't let the endpoint leak internals.

code

dockerfile · 5 lines
dockerfile
FROM gcr.io/distroless/static-debian12
COPY server /server
HEALTHCHECK --interval=10s --timeout=2s --start-period=20s --retries=3 \
  CMD ["/server", "healthcheck", "--addr=127.0.0.1:8080"]
ENTRYPOINT ["/server"]

go deeper

for a junior

Say the probe must be quick, must use a command that exists in the image, and should check that this container can serve.

for a middle

Add the depth tradeoff — not a static 200, not a downstream dependency check — and the exec-vs-shell form issue in minimal images.

for a senior

Lead with correlated-failure reasoning, insist the probe traverse the real request path with a tight timeout, and account for its cost inside the container's own limits.

for a principal

Define the platform-wide contract for what health means, so consumers can act on it uniformly, and separate instance liveness from dependency observability by design rather than per service.

## The question the probe should answer A health probe answers exactly one question: *should this instance keep being treated as working?* Everything about probe design follows from taking that narrowly. The two ways to get it wrong are symmetrical. A probe that is **too shallow** always says healthy — a static handler that answers before any real machinery is touched will happily report success while the thread pool is deadlocked and every real request times out. A probe that is **too deep** says unhealthy for things this instance cannot fix — most commonly by checking a downstream dependency. ## Why dependency checking is the classic mistake Suppose every replica of a service probes `SELECT 1` against a shared database. The database has a 30-second failover. Instantly, *every* replica reports unhealthy — simultaneously, because they share the cause. Whatever consumes health status now acts on the entire fleet: restarts, replacements, or a rolling update that can never make progress. The service was fine; the probe manufactured a fleet-wide incident out of a brief dependency blip, and the restarts likely made recovery slower by dumping caches and stampeding reconnections. The rule: **a health check should fail only for reasons a restart or replacement of this instance could plausibly fix.** A downed dependency is not one of those. Track dependency state through metrics and alerting instead, where a human or a dependency-aware policy decides what to do. There is a narrow exception: a *permanently* broken local relationship to a dependency — a connection pool that has poisoned itself and will never recover without a process restart — is legitimately an instance problem. But that is a statement about internal state, not about whether the dependency is currently reachable, and it should be implemented as "my pool has been unusable for N minutes", not as "the database did not answer this instant". ## Making the probe meaningful Route the health endpoint through the normal serving stack: the same HTTP server, the same middleware, the same worker pool that handles user traffic. Then "the probe answered in 20ms" is evidence about the path that matters. If the endpoint is served by a separate listener or a special-cased early return, the check degrades to a port scan. Good things for the endpoint to reflect: the process has finished initialisation; the request path executes end to end; local resources it owns (worker pool, disk it writes to, its own cache) are functional. Keep the response tiny — the probe only needs the status. ## Cost and blast radius The probe runs **inside the container**, and its process is charged to the container's own cgroup. Three consequences people discover the hard way: - A memory-hungry probe on a tightly limited container can contribute to hitting the memory limit — the health check itself becomes the cause of an out-of-memory kill. - On a host with many containers, `interval=1s` with a shell-spawning probe is measurable CPU across the fleet. - A probe that touches real workload resources (opening a database connection every 5 seconds from every replica) consumes pool capacity that user traffic wanted. ## Making it work in minimal images An image built `FROM scratch` or a distroless base has no shell and no `curl`. Options: - Ship a tiny static health binary and call it in exec form: `HEALTHCHECK CMD ["/healthcheck"]`. - Give the application a self-check subcommand: `HEALTHCHECK CMD ["/app/server", "healthcheck"]` — no extra binary, and the check can inspect internal state directly. - For Go services, a few lines making one HTTP request compiles to a tiny binary that adds nothing meaningful to the image. Remember shell form (`CMD cmd || exit 1`) requires `/bin/sh`; in a shell-less image only exec form works, and exec form cannot use `||`, so the command itself must return the right status. ## Timeouts, output, and security Set `--timeout` from measured latency of the endpoint (a couple of seconds for something that normally answers in milliseconds), always below the interval — the hang is the failure you most need to catch, and the timeout is what catches it. Have the probe print a one-line reason on failure. Docker captures the output into `.State.Health.Log` (truncated), so `connection refused` or `pool exhausted` appears right where the next engineer looks, without an extra round trip into the container. On security: don't put credentials or tokens in the probe command — it is visible in image metadata and `docker inspect`. Don't expose an unauthenticated health endpoint that dumps versions, config, dependency hostnames, or stack traces; keep the public surface to a status, and if you want a richer diagnostic endpoint, protect it separately. ## A checklist - Fails only for instance-local, restart-fixable reasons. - Traverses the real serving path. - Cheap, bounded, no side effects. - Binary exists in the image; exec form where there is no shell. - Timeout well under interval; one-line failure reason on stdout. - No secrets in the command; no sensitive detail in the response.

  • Should a health endpoint verify the database connection?
    Generally no. If every replica checks the same database, one blip marks the whole fleet unhealthy at once and any consumer of that status then acts on all of them simultaneously — the probe converts a brief dependency issue into a fleet-wide event. Report dependency reachability through metrics and alerts instead. The narrow exception is a locally poisoned connection pool that only a process restart can fix, expressed as sustained local unusability rather than a single failed query.
  • Why can a health check contribute to an out-of-memory kill?
    The probe executes inside the container and its process is accounted to the same cgroup as the application, so its memory counts against the container's limit. A probe that starts a shell and a heavyweight client every few seconds on a container with a tight limit adds real pressure, and on a container already near its ceiling that can be the increment that triggers the kernel OOM killer. Keeping probes to a small static binary avoids the whole class.

A health check is a pulse check on one patient, not a diagnosis of the hospital. If checking the pulse also requires the pharmacy to be open, every patient flatlines the moment the pharmacy closes.

saying these in an interview costs you the question

  • Probing downstream dependencies so a shared outage marks every replica unhealthy
  • A health endpoint that short-circuits before the real serving path, proving only that the port is open
  • Calling `curl` in an image that never installed it, or shell form in a shell-less image
  • Passing tokens or passwords on the probe command line, where `docker inspect` exposes them
  • Ignoring the probe's own CPU/memory cost inside a limited container

context