skip to content

What does the Dockerfile `HEALTHCHECK` instruction do at runtime, and where do you see its result?

level: juniorimportance: must knowfreq 66%

answer

  1. Probe runs inside the container on a schedule
  2. Exit 0 = healthy, 1 = unhealthy, 2 = reserved
  3. `docker ps` → `(healthy)` / `(unhealthy)` / `(health: starting)`
  4. `docker inspect` → `.State.Health` with FailingStreak + Log
  5. `HEALTHCHECK NONE` disables an inherited check

basics

~20 s

It tells Docker a command to run periodically inside the running container to test whether the app actually works. Exit 0 means healthy, 1 means unhealthy. The result shows in docker ps as (healthy), (unhealthy) or (starting) next to the container's uptime, with details in docker inspect.

solid answer

~50 s

`HEALTHCHECK` attaches a probe to the image: a command Docker executes inside the container on a schedule to answer "is this thing actually serving?" — as opposed to "is the process still alive", which is all the container state tells you. The contract is an exit status: **0 = healthy, 1 = unhealthy** (2 is reserved, don't use it). Results surface in three places: - `docker ps` STATUS column: `Up 5 minutes (healthy)`, `(unhealthy)`, or `(health: starting)` before the first verdict. - `docker inspect --format '{{json .State.Health}}' <ctr>` — current status, `FailingStreak`, and a `Log` of the last few probe runs with each one's exit code and captured output. - A `health_status` event on the Docker event stream, which is what other tooling reacts to. The classic case it catches: a web server process that is still running but wedged, returning errors or not accepting connections. Process-alive says fine; the health check says unhealthy.

code

dockerfile · 5 lines
dockerfile
FROM alpine:3.20
RUN apk add --no-cache wget
COPY server /usr/local/bin/server
HEALTHCHECK CMD wget -qO- http://localhost:8080/healthz >/dev/null || exit 1
ENTRYPOINT ["/usr/local/bin/server"]

go deeper

for a junior

Say what it is (a periodic command run inside the container), the 0/1 exit contract, and that docker ps shows healthy/unhealthy/starting.

for a middle

Add where the detail lives (.State.Health with FailingStreak and Log), exec vs shell form, and the missing-binary pitfall in minimal images.

for a senior

Frame it as producing a verdict that something else must consume, and be explicit that Docker Engine itself neither restarts nor reroutes on unhealthy.

for a principal

Treat the probe as a contract the image publishes to every orchestration layer above it, and standardise its semantics so platform tooling can rely on the state meaning the same thing everywhere.

## The problem it solves A container is "running" as long as its main process exists. That is a very weak statement about usefulness. A Java service can be alive with a deadlocked thread pool, a web server can hold its port open while every request 500s, a worker can be connected to nothing after its database credentials expire. Process liveness says "up"; users say "down". `HEALTHCHECK` closes that gap by letting the image declare an *application-level* test that the runtime executes periodically inside the container. ## Syntax In a Dockerfile: ```dockerfile HEALTHCHECK [OPTIONS] CMD <command> HEALTHCHECK NONE ``` `HEALTHCHECK NONE` disables a check inherited from a base image — useful when you build on an image that ships one you don't want. The `CMD` part follows the same two forms as everywhere else in a Dockerfile: - **shell form** — `HEALTHCHECK CMD curl -f http://localhost:8080/healthz || exit 1` runs through `/bin/sh -c`, so shell features like `||` work. - **exec form** — `HEALTHCHECK CMD ["/bin/healthcheck"]` runs the binary directly with no shell involved, which is what you need in an image that has no shell. In Compose the same thing is spelled with a `healthcheck:` block, where `test: ["CMD", ...]` is exec form and `test: ["CMD-SHELL", "..."]` is shell form. ## The exit-code contract The probe communicates only through its exit status: - **0** — the container is healthy. - **1** — the container is unhealthy. - **2** — reserved; do not use it (it does not mean anything useful and may be treated as a failure with different semantics in future). This is why so many examples end with `|| exit 1`: `curl -f` already returns non-zero on an HTTP error status, but normalising it to exactly 1 keeps the contract explicit. Anything the probe writes to stdout/stderr is captured and stored (truncated) in the health log, so a check that prints a useful reason gives you free diagnostics. ## Where the probe runs Crucially, the command runs **inside the container**, in its namespaces — so `localhost` is the container itself, and the probe can reach a port that was never published to the host. It also means the binary you call must exist in the image: `curl` is absent from many minimal bases, and distroless images have no shell at all, which is why a health check that works locally can fail with "not found" in a slimmed production image. ## Reading the result **`docker ps`** appends the health state to the status: ``` STATUS Up 2 minutes (health: starting) Up 5 minutes (healthy) Up 3 hours (unhealthy) ``` The three states are: - **starting** — no verdict yet; the container has just begun and failures are not yet being counted against it. - **healthy** — the most recent probe (or the first successful one) passed. - **unhealthy** — the probe has failed enough consecutive times to flip the verdict. **`docker inspect`** gives the detail: ``` docker inspect --format '{{json .State.Health}}' web ``` That returns `Status`, `FailingStreak` (how many consecutive failures so far), and `Log` — an array of recent probe runs, each with `Start`, `End`, `ExitCode`, and `Output`. When a container is flapping, the log is usually where the real error message is hiding. **Events**: `docker events --filter event=health_status` emits a line each time the status *changes*, which is the hook other tools use to react. ## What it does not do Two limits are worth stating up front because they surprise people: 1. **Docker Engine does not restart an unhealthy container.** Restart policies react to the main process exiting, and an unhealthy container has not exited. Something else has to act on the status. 2. **It does not affect traffic routing on plain Docker.** There is no built-in load balancer to remove the container from. Orchestration layers built on top (Swarm services, Compose's `depends_on: condition: service_healthy`) are what consume the state. So the instruction's job is to *produce a trustworthy verdict*; deciding what to do with the verdict is a separate design step.

  • Your health check works when you test the command by hand but reports unhealthy inside the container. What are the usual causes?
    Most often the binary is missing from the image — `curl` and even a shell are absent from many minimal and distroless bases, so the probe fails with a not-found error visible in `.State.Health.Log`. The other frequent cause is addressing: the probe runs inside the container, so it must target `localhost` and the *container* port, not the host-published port or the host's address. Reading the captured `Output` in the health log usually names the problem immediately.
  • How do you disable a health check that comes from the base image?
    Put `HEALTHCHECK NONE` in your Dockerfile to clear the inherited definition, or start the container with `docker run --no-healthcheck`. This matters when a base image ships a probe that doesn't fit your workload — for instance a database image checking a socket your build doesn't expose — because otherwise your container reports unhealthy for reasons unrelated to your application.

Process liveness is checking that the shop's lights are on. A health check is walking in and asking for a coffee — the only way to learn that the machine is broken even though the shop is technically open.

saying these in an interview costs you the question

  • Thinking a health check runs on the host, and probing the published host port instead of localhost inside the container
  • Assuming a `HEALTHCHECK` is enough to make Docker restart a broken container
  • Using exit code 2 for "unhealthy" — it is reserved; unhealthy is 1
  • Writing a probe that calls `curl` in an image where curl was never installed

context