Does a Dockerfile HEALTHCHECK instruction control whether a managed platform sends traffic to your container?
answer
- Who reads that instruction, and who ignores it
- The engine records a status, then does what
- A restart policy watches the process, not health
- The check runs with the image's own binaries
- The platform polls from outside, on its own path
basics
~20 sNo. HEALTHCHECK is image metadata that the Docker engine acts on: it runs the command inside the container and records a health status you can read with docker inspect. Managed platforms and Kubernetes ignore it and poll an endpoint they are configured with themselves.
solid answer
~40 s`HEALTHCHECK` writes a command plus its interval, timeout, retries and start period into the image configuration. When the Docker engine runs that image it executes the command inside the container on that schedule and sets `State.Health.Status` to `starting`, then `healthy` or `unhealthy` after enough consecutive results. That status is reporting only -- the engine does not restart an unhealthy container, and `--restart` reacts to the process exiting, not to health. Managed platforms and Kubernetes do not read the instruction at all; they poll a path and port you configure on their side, from outside the container. So the image's real obligation is to serve a cheap HTTP health endpoint on the injected port. Keep `HEALTHCHECK` for local runs and Compose dependency ordering, and do not rely on it in production.
code
dockerfile · 7 linesFROM eclipse-temurin:21-jre
WORKDIR /app
COPY build/libs/thumbnailer.jar /app/thumbnailer.jar
ENV PORT=8080
HEALTHCHECK --interval=20s --timeout=3s --start-period=40s --retries=3 \
CMD java -cp /app/thumbnailer.jar com.example.HealthProbe http://127.0.0.1:${PORT}/healthz || exit 1
ENTRYPOINT ["sh", "-c", "exec java -jar /app/thumbnailer.jar --server.port=${PORT}"]go deeper
Remember that HEALTHCHECK is a Dockerfile instruction the Docker engine runs inside the container, and that its result shows up as a status in docker ps and docker inspect. It is not what a managed platform looks at.
Explain the mechanics: interval, timeout, retries and start period; exit code 0 versus 1; the starting /healthy /unhealthy states; and the fact that the engine records the status without acting on it.
Show judgement about the endpoint itself -- keep it cheap, keep dependency checks out of it so one downstream blip cannot take the whole fleet out of rotation, and separate 'still starting' from 'broken' with an initial delay.
Own the convention across services: one health path, one meaning, dependency state reported through metrics rather than through the signal that controls traffic, and a clear position on whether images carry a HEALTHCHECK at all given that production ignores it.
### What the instruction does `HEALTHCHECK` is one of the few Dockerfile instructions that has a runtime effect rather than a build-time one. It stores in the image config a command and four options: `--interval` (how often to run it), `--timeout` (how long one run may take), `--retries` (how many consecutive failures before the container is declared unhealthy), and `--start-period` (a grace window during which failures do not count against the retry budget, for slow-starting applications). Newer engines add `--start-interval`, a shorter probing interval used only during that start period. `HEALTHCHECK NONE` disables a check inherited from a base image. When the Docker engine runs the container it executes the command **inside** the container namespaces on that schedule. Exit code 0 means healthy, exit code 1 means unhealthy; anything else is reserved. The result is recorded in the container's state: `starting` while the start period is in effect and no successful run has happened yet, then `healthy` or `unhealthy`. You can read it with `docker inspect --format '{{.State.Health.Status}}' <name>`, and the last few probe outputs are kept in `.State.Health.Log`, which is genuinely useful for debugging. ### What it does not do This is the part candidates get wrong. On a plain Docker engine the health status is **advisory**. The engine does not kill, restart or replace an unhealthy container. A `--restart` policy reacts to the main process exiting, not to a failing health check, so a container that has gone unhealthy but is still running stays exactly where it is, serving errors, marked `(unhealthy)` in `docker ps` and otherwise untouched. Something outside the engine has to act on the status: Swarm replaces unhealthy tasks of a service, and Compose can use it to gate startup ordering with `depends_on: condition: service_healthy`. Managed container platforms and Kubernetes take a different route entirely: they do not read the image's `HEALTHCHECK` at all. They poll a path and port that you configure in the platform, from outside the container, and they use the result to decide whether the instance receives traffic and whether it should be replaced. Shipping an image with a beautiful `HEALTHCHECK` and no HTTP health endpoint therefore satisfies nobody on the platform side. ### The image's actual obligation What the platform wants from the image is an HTTP endpoint on the injected port -- conventionally something like `/healthz` -- that answers quickly and without authentication. Three properties matter. It must be **cheap**. If the platform polls every few seconds across every replica, an endpoint that runs a database query is generating a constant background load in the exact proportion you least want. It must not **fan out**. A health endpoint that checks four downstream dependencies turns any one of their blips into a fleet-wide outage: the platform sees every replica fail, pulls them all from rotation or restarts them all, and a degraded dependency becomes a total one. Report on the process's own ability to serve; deal with dependencies through timeouts, retries and degraded responses. It must distinguish **not ready yet** from **broken**. A Spring Boot photo-thumbnail service on a JRE base takes several seconds to bring up its context; failing during that window is normal, and the platform's configuration should carry an initial delay, exactly as `--start-period` does for the Dockerfile instruction. ### A practical trap with the instruction itself The check command runs with the container's own binaries. `HEALTHCHECK CMD curl -f http://localhost:8080/healthz` works only if `curl` is in the image -- a slim JRE base may ship neither `curl` nor `wget`, and a distroless image certainly ships neither. The container then reports unhealthy forever for reasons that have nothing to do with the application. Either use something the image definitely has (the application's own runtime, or a tiny static healthcheck binary you copy in), or accept that adding `curl` to a production image widens its attack surface and its size for the sake of a status nothing in production reads. ### So should you write one? It has real value where the engine is the thing running your container: a single-host deployment, a Compose file where one service must not start before another is genuinely ready, and local development where `docker ps` telling you `(unhealthy)` beats discovering it through a failing test. It has essentially no value as a production traffic control on a managed platform, and treating it as one is the misconception the question is really probing. Write the HTTP endpoint first, because both worlds can use it; add the instruction as a convenience when the engine is the consumer.
- A container is marked unhealthy and has --restart always. Will the engine restart it?No. The restart policy reacts to the main process exiting, not to the health status. An unhealthy container whose process is still running is left alone, showing `(unhealthy)` in `docker ps`. Acting on health requires something above the engine -- Swarm replacing a service task, or a platform polling its own endpoint.
- Should a health endpoint check the database the service depends on?Generally no. If every replica reports unhealthy the moment a shared dependency blips, the platform pulls the whole fleet at once and turns a degraded dependency into an outage. Report on this process's own ability to serve, and handle dependency failure with timeouts, retries and degraded responses instead.
- Why does HEALTHCHECK CMD curl -f ... often fail on a slim base image?The command runs inside the container using the image's own binaries, and a slim or distroless base may ship neither curl nor wget. The check then fails permanently for reasons unrelated to the application. Use the application's own runtime or a small static probe binary rather than adding a network client to a production image.
saying these in an interview costs you the question
- Thinks HEALTHCHECK decides traffic on a managed platform
- Believes an unhealthy container is restarted automatically
- Assumes Kubernetes reads the image's HEALTHCHECK
- Writes a health endpoint that queries every dependency
- Uses curl in a HEALTHCHECK on a distroless base
- Confuses the restart policy with health-driven replacement