skip to content

questions

5

What does the Dockerfile `HEALTHCHECK` instruction do at runtime, and where do you see its result?

level: juniorimportance: must knowfreq 66%

answer

  1. Probe runs inside the container on a schedule
  2. Exit 0 = healthy, 1 = unhealthy, 2 = reserved
  3. `docker ps` → `(healthy)` / `(unhealthy)` / `(health: starting)`
  4. `docker inspect` → `.State.Health` with FailingStreak + Log
  5. `HEALTHCHECK NONE` disables an inherited check

basics

~20 s

It tells Docker a command to run periodically inside the running container to test whether the app actually works. Exit 0 means healthy, 1 means unhealthy. The result shows in docker ps as (healthy), (unhealthy) or (starting) next to the container's uptime, with details in docker inspect.

solid answer

~50 s

`HEALTHCHECK` attaches a probe to the image: a command Docker executes inside the container on a schedule to answer "is this thing actually serving?" — as opposed to "is the process still alive", which is all the container state tells you. The contract is an exit status: **0 = healthy, 1 = unhealthy** (2 is reserved, don't use it). Results surface in three places: - `docker ps` STATUS column: `Up 5 minutes (healthy)`, `(unhealthy)`, or `(health: starting)` before the first verdict. - `docker inspect --format '{{json .State.Health}}' <ctr>` — current status, `FailingStreak`, and a `Log` of the last few probe runs with each one's exit code and captured output. - A `health_status` event on the Docker event stream, which is what other tooling reacts to. The classic case it catches: a web server process that is still running but wedged, returning errors or not accepting connections. Process-alive says fine; the health check says unhealthy.

code

dockerfile · 5 lines
dockerfile
FROM alpine:3.20
RUN apk add --no-cache wget
COPY server /usr/local/bin/server
HEALTHCHECK CMD wget -qO- http://localhost:8080/healthz >/dev/null || exit 1
ENTRYPOINT ["/usr/local/bin/server"]

go deeper

for a junior

Say what it is (a periodic command run inside the container), the 0/1 exit contract, and that docker ps shows healthy/unhealthy/starting.

for a middle

Add where the detail lives (.State.Health with FailingStreak and Log), exec vs shell form, and the missing-binary pitfall in minimal images.

for a senior

Frame it as producing a verdict that something else must consume, and be explicit that Docker Engine itself neither restarts nor reroutes on unhealthy.

for a principal

Treat the probe as a contract the image publishes to every orchestration layer above it, and standardise its semantics so platform tooling can rely on the state meaning the same thing everywhere.

## The problem it solves A container is "running" as long as its main process exists. That is a very weak statement about usefulness. A Java service can be alive with a deadlocked thread pool, a web server can hold its port open while every request 500s, a worker can be connected to nothing after its database credentials expire. Process liveness says "up"; users say "down". `HEALTHCHECK` closes that gap by letting the image declare an *application-level* test that the runtime executes periodically inside the container. ## Syntax In a Dockerfile: ```dockerfile HEALTHCHECK [OPTIONS] CMD <command> HEALTHCHECK NONE ``` `HEALTHCHECK NONE` disables a check inherited from a base image — useful when you build on an image that ships one you don't want. The `CMD` part follows the same two forms as everywhere else in a Dockerfile: - **shell form** — `HEALTHCHECK CMD curl -f http://localhost:8080/healthz || exit 1` runs through `/bin/sh -c`, so shell features like `||` work. - **exec form** — `HEALTHCHECK CMD ["/bin/healthcheck"]` runs the binary directly with no shell involved, which is what you need in an image that has no shell. In Compose the same thing is spelled with a `healthcheck:` block, where `test: ["CMD", ...]` is exec form and `test: ["CMD-SHELL", "..."]` is shell form. ## The exit-code contract The probe communicates only through its exit status: - **0** — the container is healthy. - **1** — the container is unhealthy. - **2** — reserved; do not use it (it does not mean anything useful and may be treated as a failure with different semantics in future). This is why so many examples end with `|| exit 1`: `curl -f` already returns non-zero on an HTTP error status, but normalising it to exactly 1 keeps the contract explicit. Anything the probe writes to stdout/stderr is captured and stored (truncated) in the health log, so a check that prints a useful reason gives you free diagnostics. ## Where the probe runs Crucially, the command runs **inside the container**, in its namespaces — so `localhost` is the container itself, and the probe can reach a port that was never published to the host. It also means the binary you call must exist in the image: `curl` is absent from many minimal bases, and distroless images have no shell at all, which is why a health check that works locally can fail with "not found" in a slimmed production image. ## Reading the result **`docker ps`** appends the health state to the status: ``` STATUS Up 2 minutes (health: starting) Up 5 minutes (healthy) Up 3 hours (unhealthy) ``` The three states are: - **starting** — no verdict yet; the container has just begun and failures are not yet being counted against it. - **healthy** — the most recent probe (or the first successful one) passed. - **unhealthy** — the probe has failed enough consecutive times to flip the verdict. **`docker inspect`** gives the detail: ``` docker inspect --format '{{json .State.Health}}' web ``` That returns `Status`, `FailingStreak` (how many consecutive failures so far), and `Log` — an array of recent probe runs, each with `Start`, `End`, `ExitCode`, and `Output`. When a container is flapping, the log is usually where the real error message is hiding. **Events**: `docker events --filter event=health_status` emits a line each time the status *changes*, which is the hook other tools use to react. ## What it does not do Two limits are worth stating up front because they surprise people: 1. **Docker Engine does not restart an unhealthy container.** Restart policies react to the main process exiting, and an unhealthy container has not exited. Something else has to act on the status. 2. **It does not affect traffic routing on plain Docker.** There is no built-in load balancer to remove the container from. Orchestration layers built on top (Swarm services, Compose's `depends_on: condition: service_healthy`) are what consume the state. So the instruction's job is to *produce a trustworthy verdict*; deciding what to do with the verdict is a separate design step.

  • Your health check works when you test the command by hand but reports unhealthy inside the container. What are the usual causes?
    Most often the binary is missing from the image — `curl` and even a shell are absent from many minimal and distroless bases, so the probe fails with a not-found error visible in `.State.Health.Log`. The other frequent cause is addressing: the probe runs inside the container, so it must target `localhost` and the *container* port, not the host-published port or the host's address. Reading the captured `Output` in the health log usually names the problem immediately.
  • How do you disable a health check that comes from the base image?
    Put `HEALTHCHECK NONE` in your Dockerfile to clear the inherited definition, or start the container with `docker run --no-healthcheck`. This matters when a base image ships a probe that doesn't fit your workload — for instance a database image checking a socket your build doesn't expose — because otherwise your container reports unhealthy for reasons unrelated to your application.

Process liveness is checking that the shop's lights are on. A health check is walking in and asking for a coffee — the only way to learn that the machine is broken even though the shop is technically open.

saying these in an interview costs you the question

  • Thinking a health check runs on the host, and probing the published host port instead of localhost inside the container
  • Assuming a `HEALTHCHECK` is enough to make Docker restart a broken container
  • Using exit code 2 for "unhealthy" — it is reserved; unhealthy is 1
  • Writing a probe that calls `curl` in an image where curl was never installed

context

open as a page

Explain what `--interval`, `--timeout`, `--retries` and `--start-period` control on a container health probe, and how they determine when the reported state flips.

level: middleimportance: must knowfreq 58%

basics

~20 s

--interval is the wait between probe runs, --timeout is how long one probe may take before it counts as a failure, --retries is how many consecutive failures flip the state to unhealthy, and --start-period is an initial window where failures don't count — one success at any time marks the container healthy.

open as a page

A container reports `(unhealthy)` in `docker ps` but keeps running and still receives traffic. Why doesn't Docker Engine act on that status, and how do you make something act on it?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Docker Engine only reports health; it never restarts on it. Restart policies react to the main process exiting, and an unhealthy container hasn't exited. To act on it you need a consumer: an orchestrator, Compose dependency conditions, or your own watcher on the health_status event stream.

open as a page

What makes a container health probe command useful rather than misleading, and what should it deliberately avoid testing?

level: seniorimportance: should knowfreq 50%

basics

~20 s

A good probe tests that this instance can serve — a cheap, dedicated endpoint that exercises the serving path — with a tight timeout and a binary that actually exists in the image. It should not test downstream dependencies, do real work, or require credentials, because that turns one dependency outage into every container reporting unhealthy.

open as a page

How can you add, change, or disable a container's health probe at `docker run` time without rebuilding the image, and when is that the right move?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

docker run accepts --health-cmd, --health-interval, --health-timeout, --health-retries and --health-start-period to define or retune a probe, and --no-healthcheck to disable one entirely. Use them to tune against a real workload, or to override a base image's unsuitable check.

open as a page