skip to content

How does the Dockerfile HEALTHCHECK instruction work, and what do its --interval, --timeout, --retries and --start-period options control?

level: middleimportance: must knowfreq 52%

answer

  1. exit 0 healthy, 1 unhealthy, 2 reserved
  2. starting -> healthy/unhealthy after N consecutive fails
  3. start-period: failures do not count, success ends it
  4. worst-case detect = interval x retries + timeout
  5. Docker does not restart on unhealthy; k8s ignores it

basics

~20 s

HEALTHCHECK CMD runs a command inside the container on a schedule: exit 0 means healthy, 1 means unhealthy. --interval is the gap between checks, --timeout kills a hung check, --retries is how many consecutive failures flip the status, --start-period is a startup grace window where failures do not count.

solid answer

~50 s

`HEALTHCHECK [OPTIONS] CMD <command>` makes the daemon run that command inside the container periodically. Exit code 0 marks the container healthy, non-zero (use 1) unhealthy. The reported status starts as `starting`, becomes `healthy` on the first success, and `unhealthy` only after --retries consecutive failures (default 3). Options: --interval (default 30s) between checks, --timeout (default 30s) after which a hung check is killed and counted as a failure, --retries, and --start-period (default 0s), a grace window for slow-booting apps in which failures do not count toward retries and do not mark the container unhealthy; a success ends it early. Newer engines add --start-interval for faster probing during that window. Key caveat: Docker itself does not restart an unhealthy container. Swarm reschedules, and Compose's `depends_on: condition: service_healthy` gates startup. Kubernetes ignores the image's HEALTHCHECK entirely and uses its own probes. The check binary must exist inside the image, which is the classic problem with distroless bases.

code

dockerfile · 2 lines
dockerfile
HEALTHCHECK --interval=15s --timeout=3s --retries=3 --start-period=30s \
  CMD curl -fsS http://localhost:8080/healthz || exit 1

go deeper

for a junior

Know the syntax, exit-code meaning, and what each of the four options does.

for a middle

Compute worst-case detection time, explain the starting state and the start-period grace, and know that the tool must exist in the image.

for a senior

Discuss who consumes the status (Swarm, Compose gating, nothing on plain docker run, probes in Kubernetes), check cost, and why probing dependencies causes correlated failure.

for a principal

Set fleet-wide conventions: what health means versus readiness, detection-time budgets tied to SLOs, and avoiding correlated restart storms.

## Mechanics The instruction has two forms: `HEALTHCHECK [OPTIONS] CMD <command>` and `HEALTHCHECK NONE`, the latter used to cancel a check inherited from a base image. The command runs *inside* the container, using the container's own filesystem and network namespace, on the daemon's schedule. Its exit status is the verdict: 0 healthy, 1 unhealthy. Status 2 is reserved and should not be used. Both shell form (`CMD curl -f http://localhost:8080/health || exit 1`) and exec form (`CMD ["curl", "-f", "http://localhost:8080/health"]`) work; exec form skips the shell. The daemon records the last few results, including exit code and truncated output, in the container state. Read them with `docker inspect --format '{{json .State.Health}}' <container>`; `docker ps` shows the status next to the container. ## The four options **--interval** (default 30s): time between the end of one check and the start of the next. Shorter means faster detection but more load; a check is a full process spawn inside the container. **--timeout** (default 30s): a check exceeding this is killed and counted as a failure. A timeout longer than the interval is a smell - detection lags and checks queue up conceptually. **--retries** (default 3): the number of *consecutive* failures needed to move from healthy to unhealthy. A single success resets the counter. Worst-case detection time is roughly interval x retries plus timeout, which is the number to quote when someone asks how fast a failure is noticed. **--start-period** (default 0s): a startup grace window measured from container start. Failures during it do not increment the retry counter and cannot mark the container unhealthy; the status stays `starting`. The first success ends the window immediately. This exists for apps with slow boots - JVM warmups, migrations - so they are not declared dead before they ever came up. Modern engines also support **--start-interval**, a shorter interval used only during the start period, so a fast-starting container is marked healthy quickly while still tolerating a long worst case. ## Who acts on the result This is where candidates go wrong. Plain `docker run` does nothing with an unhealthy status: the container keeps running, and even `--restart` policies react to exit codes, not health. Docker Swarm reschedules unhealthy tasks. Docker Compose uses it for ordering via `depends_on: <svc>: condition: service_healthy`, which is the most common reason to add a HEALTHCHECK to an image used in local or CI stacks. Kubernetes ignores the image's HEALTHCHECK completely and relies on its own probes defined in the pod spec, so an image destined only for Kubernetes gains nothing operationally from the instruction. Load balancers and reverse proxies in front of the container use their own probes as well. ## Writing a good check Probe what proves the process can serve traffic - an in-process HTTP endpoint that does a cheap internal check. Avoid probing downstream dependencies: if the database is briefly unavailable, marking every application container unhealthy converts a dependency blip into a cascading restart storm. Keep the check cheap: it runs forever, on every container. The tool must exist in the image. Minimal and distroless images ship no curl or wget, so either add a tiny static health binary, use a language-runtime one-liner already present, or have the application expose a `--healthcheck` subcommand that dials itself. Also remember the check runs as the image's configured user, so it must have permission to run and to reach the local port. ## Practical defaults A reasonable starting point for a web service: `--interval=15s --timeout=3s --retries=3 --start-period=30s`. Detection then takes at most about 45 seconds, slow starts are tolerated, and a hung request cannot stall the checker.

  • A container is reported unhealthy but keeps running and serving nothing. Why didn't Docker restart it?
    The Docker Engine only reports health status; it takes no remedial action, and restart policies key off the process exit code, not health. Something else must act: Swarm reschedules unhealthy tasks, Compose uses health for dependency gating, and in Kubernetes a failing liveness probe - not the image HEALTHCHECK - restarts the container. If you need self-healing on plain Docker, an external supervisor must watch the status and act.
  • How do you write a health check for a distroless image with no shell or curl?
    Ship a check that needs nothing extra: give the application a subcommand that dials its own endpoint and exits 0 or 1, or copy a tiny statically linked health binary from a builder stage. Then use exec form, HEALTHCHECK CMD ["/app", "healthcheck"], since there is no shell to interpret the shell form.
  • Should the health check verify the database connection?
    Usually not for a liveness-style check. If a shared dependency hiccups, every instance flips unhealthy at once and the platform restarts or removes them all, turning a transient outage into an outage of your service. Keep the check local and cheap, and handle dependency state with readiness semantics, retries and circuit breakers instead.

saying these in an interview costs you the question

  • Thinking Docker restarts a container when it becomes unhealthy
  • Confusing --start-period with --interval, or believing failures during the start period count
  • Using exit code 2 to signal unhealthy
  • Writing a check that calls downstream services, so one dependency blip marks the whole fleet unhealthy
  • Assuming Kubernetes honors the image's HEALTHCHECK instead of pod probes

context