skip to content

A container reports `(unhealthy)` in `docker ps` but keeps running and still receives traffic. Why doesn't Docker Engine act on that status, and how do you make something act on it?

level: seniorimportance: must knowfreq 56%

answer

  1. Restart policies fire on exit; unhealthy never exits
  2. Engine reports health, never acts on it
  3. No built-in routing → traffic keeps arriving
  4. Consumers: Swarm, Compose service_healthy, `docker events` watcher
  5. Cleanest single-host fix: app exits non-zero → --restart=on-failure

basics

~20 s

Docker Engine only reports health; it never restarts on it. Restart policies react to the main process exiting, and an unhealthy container hasn't exited. To act on it you need a consumer: an orchestrator, Compose dependency conditions, or your own watcher on the health_status event stream.

solid answer

~60 s

Health status on plain Docker Engine is **information, not control**. Two separate mechanisms are commonly conflated: - **Restart policies** (`--restart=on-failure|always|unless-stopped`) trigger when PID 1 **exits**. An unhealthy container is still running, so no policy fires — ever. - **Health checks** produce a status and emit `health_status` events. Nothing in the engine subscribes to them. There is also no built-in routing to remove the container from, so traffic keeps arriving. Ways to make the status actionable, in rough order of preference: 1. **Let the orchestration layer consume it.** Swarm reschedules unhealthy service tasks; Compose gates startup with `depends_on: condition: service_healthy`. 2. **Watch the event stream** — `docker events --filter event=health_status` — and have a small supervisor (the well-known "autoheal" sidecar pattern) restart containers that flip to unhealthy. 3. **Make the app self-terminate.** If the process detects an unrecoverable state and exits non-zero, `--restart=on-failure` handles it and the two mechanisms line up. Option 3 is often cleanest for single-host setups, because the app knows more than the probe does. Whatever you choose, a restart loop needs backoff and an alarm — silent restarts hide real bugs.

code

bash · 6 lines
bash
docker events --filter event=health_status --format '{{.Status}} {{.Actor.Attributes.name}}' |
while read -r status name; do
  case "$status" in
    *unhealthy) echo "restarting $name"; docker restart "$name" ;;
  esac
done

go deeper

for a junior

Say that health status is only reported, and restart policies react to the process exiting — an unhealthy container hasn't exited, so nothing happens.

for a middle

Name the consumers: Compose condition: service_healthy for startup ordering, Swarm for rescheduling, or a watcher on docker events.

for a senior

Argue why the separation is deliberate, include app self-termination as the cleanest single-host fix, and demand backoff, caps, and probes that don't test downstream dependencies.

for a principal

Frame health as a sensor whose consumer is an architectural choice; set the policy for which layer actuates, and account for correlated failure — one dependency outage must not restart the whole estate at once.

## Two mechanisms that look related and aren't The surprise comes from assuming health status feeds restart policy. It doesn't, on either side. **Restart policies** are attached at `docker run` time (`--restart=no|on-failure[:max]|always|unless-stopped`). The daemon watches the container's **main process**. When PID 1 exits, the daemon consults the policy and the exit code and decides whether to start the container again. The only input is *exit*. **Health checks** run a probe inside a *running* container and record a status. Nothing in Docker Engine consumes that status: it is written to `.State.Health`, rendered in `docker ps`, and emitted as an event. There is no code path from `unhealthy` to `restart`. So an unhealthy container sits there indefinitely, running and broken, and — since Docker Engine has no load balancer — still reachable through published ports, DNS on user-defined networks, and any client holding a connection. ## Why the design is defensible It is easy to read this as an omission, but the separation is reasonable: - **Restarting is a policy decision, not a runtime one.** Whether a broken instance should be restarted, drained, replaced elsewhere, or left up for debugging depends on the system around it. The engine deliberately doesn't guess. - **Automatic restart on unhealthy is genuinely dangerous in the wrong situation.** If a shared dependency (a database) goes down and every service's probe checks it, engine-level auto-restart would restart the entire estate simultaneously — turning a recoverable dependency outage into a thundering-herd cold start. - **A running-but-unhealthy container is often more useful than a restarted one**: you can `docker exec` into it, dump threads, and read state that a restart would destroy. ## Consuming the status **Event-driven supervision.** Every status *change* emits an event: ``` docker events --filter event=health_status ``` A small watcher subscribing to that stream and restarting the named container is the canonical single-host answer (packaged versions of this are widely known as "autoheal" containers; they typically opt in via a label and mount the Docker socket). Two cautions: mounting the Docker socket into a container grants effective root on the host, so the watcher is a privileged component to be treated as such; and the watcher must implement backoff and a restart cap, or a permanently broken service becomes a restart storm. **Startup ordering.** Compose's `depends_on` with `condition: service_healthy` holds a dependent service until its dependency reports healthy. That solves the classic "app starts before the database is accepting connections" race properly, replacing the wait-for-it/sleep hacks. Note it is a *startup* gate — it doesn't manage anything after boot. **Orchestrated services.** Swarm services treat an unhealthy task as failed and replace it according to the service's restart policy, and hold back a rolling update if new tasks don't become healthy. That is the built-in "something acts on it" story, and it is one reason health checks matter more once you're above plain `docker run`. **Self-termination.** Frequently the best design: give the application an internal watchdog for conditions it knows are unrecoverable (deadlocked pool, permanently lost dependency after N attempts, corrupted internal state) and have it exit non-zero. Now `--restart=on-failure` applies, restart accounting is uniform, exit codes tell the story, and there is no privileged sidecar. The application also has far more context than an HTTP probe about whether the condition is really unrecoverable. ## What to check when a container sits unhealthy 1. `docker inspect --format '{{json .State.Health}}' <ctr>` — the `Log` array holds the last several probe runs with exit codes and captured output. This is usually where the real error is. 2. Ask whether the *probe* is wrong rather than the app: a check that hits a downstream dependency will report unhealthy for an app that is perfectly fine but whose database is down. That is a probe-design defect, and it is exactly the case where auto-restart would make things worse. 3. Check `FailingStreak` and timing against `--interval`/`--timeout` to see whether it is flapping or solidly down. 4. Decide deliberately: restart, replace, or hold for investigation — and if it is a recurring class, wire the consumer so a human isn't the loop. ## The one-line summary for an interview Docker Engine's health check is a **sensor**, not an **actuator**. Restart policies actuate on exit. Connect the two yourself — via the orchestrator, an event-driven supervisor, or by making the app exit when it knows it is dead.

  • Why can automatically restarting every unhealthy container make an incident worse?
    If probes check a shared downstream dependency, one dependency outage marks every instance unhealthy at once, and blanket auto-restart then cold-starts the whole estate simultaneously — cache loss, connection stampede, and a slower recovery than if the containers had simply waited. It also destroys the in-memory evidence you would use to diagnose the failure. Restarts need backoff, a cap, and probes that test the local service rather than its dependencies.
  • What is the drawback of the sidecar watcher that restarts unhealthy containers?
    It needs the Docker socket mounted, which is equivalent to root on the host — so a compromise of that small container compromises the machine. It is also an extra moving part that must implement its own backoff and cap, and it sits outside your application's knowledge of whether the condition is recoverable. On a single host it is pragmatic; if the workload already runs under an orchestrator, use the orchestrator's own health handling instead.
  • How would you decide between the watcher approach and having the app exit on its own?
    Prefer app self-termination when the process can reliably distinguish unrecoverable states, because it needs no privileged component, unifies restart accounting under the normal restart policy, and leaves a meaningful exit code behind. Use an external watcher when you cannot change the application, or when the failure mode is exactly the kind the process cannot detect about itself — a hang, a deadlock, or a wedged event loop that would never reach its own watchdog.

A smoke detector wired to nothing. It reliably tells you there's a fire; it will not call the fire brigade until you connect something to it.

saying these in an interview costs you the question

  • Believing `--restart=always` or `on-failure` reacts to an unhealthy status
  • Expecting Docker to stop routing traffic to an unhealthy container
  • Writing a probe that checks a downstream database, so a dependency outage marks the app itself unhealthy
  • Adding an auto-restart loop with no backoff or cap
  • Ignoring `.State.Health.Log`, which holds the failing probes' actual output

context