What does the Dockerfile `HEALTHCHECK` instruction do at runtime, and where do you see its result?
answer
- Probe runs inside the container on a schedule
- Exit 0 = healthy, 1 = unhealthy, 2 = reserved
- `docker ps` → `(healthy)` / `(unhealthy)` / `(health: starting)`
- `docker inspect` → `.State.Health` with FailingStreak + Log
- `HEALTHCHECK NONE` disables an inherited check
basics
~20 sIt tells Docker a command to run periodically inside the running container to test whether the app actually works. Exit 0 means healthy, 1 means unhealthy. The result shows in docker ps as (healthy), (unhealthy) or (starting) next to the container's uptime, with details in docker inspect.
solid answer
~50 s`HEALTHCHECK` attaches a probe to the image: a command Docker executes inside the container on a schedule to answer "is this thing actually serving?" — as opposed to "is the process still alive", which is all the container state tells you. The contract is an exit status: **0 = healthy, 1 = unhealthy** (2 is reserved, don't use it). Results surface in three places: - `docker ps` STATUS column: `Up 5 minutes (healthy)`, `(unhealthy)`, or `(health: starting)` before the first verdict. - `docker inspect --format '{{json .State.Health}}' <ctr>` — current status, `FailingStreak`, and a `Log` of the last few probe runs with each one's exit code and captured output. - A `health_status` event on the Docker event stream, which is what other tooling reacts to. The classic case it catches: a web server process that is still running but wedged, returning errors or not accepting connections. Process-alive says fine; the health check says unhealthy.
code
dockerfile · 5 linesFROM alpine:3.20
RUN apk add --no-cache wget
COPY server /usr/local/bin/server
HEALTHCHECK CMD wget -qO- http://localhost:8080/healthz >/dev/null || exit 1
ENTRYPOINT ["/usr/local/bin/server"]go deeper
Say what it is (a periodic command run inside the container), the 0/1 exit contract, and that docker ps shows healthy/unhealthy/starting.
Add where the detail lives (.State.Health with FailingStreak and Log), exec vs shell form, and the missing-binary pitfall in minimal images.
Frame it as producing a verdict that something else must consume, and be explicit that Docker Engine itself neither restarts nor reroutes on unhealthy.
Treat the probe as a contract the image publishes to every orchestration layer above it, and standardise its semantics so platform tooling can rely on the state meaning the same thing everywhere.
## The problem it solves A container is "running" as long as its main process exists. That is a very weak statement about usefulness. A Java service can be alive with a deadlocked thread pool, a web server can hold its port open while every request 500s, a worker can be connected to nothing after its database credentials expire. Process liveness says "up"; users say "down". `HEALTHCHECK` closes that gap by letting the image declare an *application-level* test that the runtime executes periodically inside the container. ## Syntax In a Dockerfile: ```dockerfile HEALTHCHECK [OPTIONS] CMD <command> HEALTHCHECK NONE ``` `HEALTHCHECK NONE` disables a check inherited from a base image — useful when you build on an image that ships one you don't want. The `CMD` part follows the same two forms as everywhere else in a Dockerfile: - **shell form** — `HEALTHCHECK CMD curl -f http://localhost:8080/healthz || exit 1` runs through `/bin/sh -c`, so shell features like `||` work. - **exec form** — `HEALTHCHECK CMD ["/bin/healthcheck"]` runs the binary directly with no shell involved, which is what you need in an image that has no shell. In Compose the same thing is spelled with a `healthcheck:` block, where `test: ["CMD", ...]` is exec form and `test: ["CMD-SHELL", "..."]` is shell form. ## The exit-code contract The probe communicates only through its exit status: - **0** — the container is healthy. - **1** — the container is unhealthy. - **2** — reserved; do not use it (it does not mean anything useful and may be treated as a failure with different semantics in future). This is why so many examples end with `|| exit 1`: `curl -f` already returns non-zero on an HTTP error status, but normalising it to exactly 1 keeps the contract explicit. Anything the probe writes to stdout/stderr is captured and stored (truncated) in the health log, so a check that prints a useful reason gives you free diagnostics. ## Where the probe runs Crucially, the command runs **inside the container**, in its namespaces — so `localhost` is the container itself, and the probe can reach a port that was never published to the host. It also means the binary you call must exist in the image: `curl` is absent from many minimal bases, and distroless images have no shell at all, which is why a health check that works locally can fail with "not found" in a slimmed production image. ## Reading the result **`docker ps`** appends the health state to the status: ``` STATUS Up 2 minutes (health: starting) Up 5 minutes (healthy) Up 3 hours (unhealthy) ``` The three states are: - **starting** — no verdict yet; the container has just begun and failures are not yet being counted against it. - **healthy** — the most recent probe (or the first successful one) passed. - **unhealthy** — the probe has failed enough consecutive times to flip the verdict. **`docker inspect`** gives the detail: ``` docker inspect --format '{{json .State.Health}}' web ``` That returns `Status`, `FailingStreak` (how many consecutive failures so far), and `Log` — an array of recent probe runs, each with `Start`, `End`, `ExitCode`, and `Output`. When a container is flapping, the log is usually where the real error message is hiding. **Events**: `docker events --filter event=health_status` emits a line each time the status *changes*, which is the hook other tools use to react. ## What it does not do Two limits are worth stating up front because they surprise people: 1. **Docker Engine does not restart an unhealthy container.** Restart policies react to the main process exiting, and an unhealthy container has not exited. Something else has to act on the status. 2. **It does not affect traffic routing on plain Docker.** There is no built-in load balancer to remove the container from. Orchestration layers built on top (Swarm services, Compose's `depends_on: condition: service_healthy`) are what consume the state. So the instruction's job is to *produce a trustworthy verdict*; deciding what to do with the verdict is a separate design step.
- Your health check works when you test the command by hand but reports unhealthy inside the container. What are the usual causes?Most often the binary is missing from the image — `curl` and even a shell are absent from many minimal and distroless bases, so the probe fails with a not-found error visible in `.State.Health.Log`. The other frequent cause is addressing: the probe runs inside the container, so it must target `localhost` and the *container* port, not the host-published port or the host's address. Reading the captured `Output` in the health log usually names the problem immediately.
- How do you disable a health check that comes from the base image?Put `HEALTHCHECK NONE` in your Dockerfile to clear the inherited definition, or start the container with `docker run --no-healthcheck`. This matters when a base image ships a probe that doesn't fit your workload — for instance a database image checking a socket your build doesn't expose — because otherwise your container reports unhealthy for reasons unrelated to your application.
Process liveness is checking that the shop's lights are on. A health check is walking in and asking for a coffee — the only way to learn that the machine is broken even though the shop is technically open.
saying these in an interview costs you the question
- Thinking a health check runs on the host, and probing the published host port instead of localhost inside the container
- Assuming a `HEALTHCHECK` is enough to make Docker restart a broken container
- Using exit code 2 for "unhealthy" — it is reserved; unhealthy is 1
- Writing a probe that calls `curl` in an image where curl was never installed