Explain what `--interval`, `--timeout`, `--retries` and `--start-period` control on a container health probe, and how they determine when the reported state flips.
answer
- interval = gap between probes (30s default)
- timeout = kill a hung probe, counts as failure (30s default)
- retries = consecutive failures to flip unhealthy (3 default)
- start-period = failures don't count; first success ends it
- Detection ≈ retries × (interval + timeout)
basics
~20 s--interval is the wait between probe runs, --timeout is how long one probe may take before it counts as a failure, --retries is how many consecutive failures flip the state to unhealthy, and --start-period is an initial window where failures don't count — one success at any time marks the container healthy.
solid answer
~50 sFour knobs, with Docker's defaults: - **`--interval=30s`** — gap between probe runs (measured from the end of one to the start of the next). - **`--timeout=30s`** — a probe exceeding this is killed and counted as a failure. Critical, because a hung probe is exactly the condition you're trying to detect. - **`--retries=3`** — number of *consecutive* failures required to flip to `unhealthy`. One success resets the streak. - **`--start-period=0s`** — an initial grace window for slow-starting apps: failures during it do **not** count toward `retries` and the state stays `health: starting`. A single success ends the start period immediately and marks it healthy. So worst-case detection time after the app breaks is roughly `retries × (interval + timeout)`. With defaults, up to about three minutes — usually too slow, hence typical production values like `--interval=10s --timeout=2s --retries=3`. The common mistake is using `--interval` as a startup delay. That only slows detection forever; `--start-period` is the correct tool because it costs nothing after the app is up.
code
dockerfile · 2 linesHEALTHCHECK --interval=10s --timeout=2s --start-period=60s --retries=3 \
CMD wget -qO- http://localhost:8080/healthz >/dev/null || exit 1go deeper
Name the four options and what each one means, plus the defaults 30s/30s/3/0s.
Explain the transitions — consecutive failures, streak reset on success, start-period ending on first success — and compute detection time as retries × (interval + timeout).
Derive values from measured latency and startup time, insist on a timeout well below the interval, and account for probe cost against the container's own resource limits.
Set organisation-wide defaults tied to what consumes the verdict (dependency gating vs alerting vs rescheduling), and treat detection latency as an explicit SLO input rather than a per-service accident.
## The state machine A container with a health check moves between three states: `starting`, `healthy`, `unhealthy`. The four options decide when transitions occur. ## `--interval` (default 30s) The delay between probe runs, timed from the end of one attempt to the beginning of the next. It sets your detection granularity: you cannot learn about a failure faster than one interval, and it also sets the recurring cost, since every probe spawns a process inside the container. Shorter is more responsive and more expensive. On a host packing dozens of containers, a 1-second interval with a heavyweight probe is a real CPU line item; 5–15 seconds is the usual production range for a service, longer for batch-ish workloads. ## `--timeout` (default 30s) The maximum duration of a single probe. Exceed it, and Docker kills the probe and records that attempt as a failure. This is the option people forget, and it is the one that makes health checking useful. The failure mode you most want to catch — a wedged server that accepts connections but never responds — makes the probe *hang*, not fail fast. With the 30-second default and a 30-second interval, a hung service is detected agonisingly slowly. A timeout must be well under the interval; typically it should be a small multiple of the endpoint's normal latency, e.g. 2s for a health endpoint that answers in 20ms. ## `--retries` (default 3) How many **consecutive** failures are needed to declare the container unhealthy. `.State.Health.FailingStreak` is the live counter, and any success resets it to zero. The purpose is noise suppression: a single dropped probe during a GC pause or a momentary scheduling hiccup should not flip an otherwise fine service. The cost is latency — a real outage takes `retries` probes to be believed. Retries of 2–3 is the normal balance; 1 makes the check twitchy, 10 makes it useless as an alarm. ## `--start-period` (default 0s) A startup grace window measured from container start. During it: - probe failures **do not** count toward `retries`; - the container stays in `health: starting` rather than going `unhealthy`; - the **first success ends the period immediately** and marks the container healthy — you are not forced to wait out the whole window. This exists for applications with genuinely slow boots: JVM services warming up, apps running migrations at start, anything loading a large model or cache. Without it, a service that takes 45 seconds to serve traffic gets marked unhealthy at ~30 seconds and (depending on what consumes that state) may be torn down before it ever finishes starting — a restart loop that looks like a crash but is pure misconfiguration. The key insight is why it beats simply raising `--interval` or `--retries`: those degrade detection *forever*, whereas `--start-period` costs nothing once the app has answered once. Set it above the p99 startup time you have actually measured, and don't be shy — a generous start period has no steady-state cost. Docker 25.0 added a companion, `--start-interval` (default 5s), which is the probing frequency *during* the start period. It lets you poll quickly at boot (so the container is marked healthy the moment it's ready) while keeping a relaxed steady-state `--interval`. ## Doing the arithmetic Worst-case time from "the app broke" to "state reads unhealthy" is roughly: ``` retries × (interval + timeout) ``` Defaults: 3 × (30s + 30s) = up to 3 minutes. A tuned service: 3 × (10s + 2s) = about 36 seconds. Whether that is acceptable depends entirely on who consumes the verdict — a dependency gate at startup can be slow, an alarm feeding an incident response cannot. ## Practical guidance - Set `timeout` from measured endpoint latency, not by guessing, and always keep it below `interval`. - Set `start-period` from measured startup time, with headroom. - Keep `retries` low enough that the check is a useful alarm, high enough that a single hiccup doesn't flip it. - Remember the probe runs inside the container and consumes its resources — its process counts against the container's own limits, so a memory-hungry probe on a tightly limited container is a self-inflicted problem. - All four are overridable at run time (`--health-interval` and friends) without rebuilding the image, which is handy for tuning against a real workload before baking values into the Dockerfile.
- Why is `--start-period` better than just raising `--retries` for a slow-starting application?Raising `retries` degrades detection permanently: once the app is up, every real outage now takes that many more probes to be believed. `--start-period` applies only until the first successful probe, so it buys startup tolerance at zero steady-state cost. It also keeps the container in an explicit `starting` state rather than pretending failures are healthy, which is more honest to anything consuming the status.
- What happens if a probe hangs and never returns?Docker kills it once `--timeout` elapses and records the attempt as a failure, incrementing the failing streak. That is exactly why the timeout must be set deliberately: the wedged-but-listening server is the failure mode health checks exist for, and with the 30-second default alongside a 30-second interval, detection is far slower than most people expect.
saying these in an interview costs you the question
- Using `--interval` or extra `--retries` as a startup delay instead of `--start-period`
- Leaving `--timeout` at the 30s default so a hung probe takes minutes to register
- Believing the start period always runs to completion — a single success ends it immediately
- Thinking `--retries` counts total failures rather than consecutive ones (any success resets the streak)
- Setting a 1s interval on a heavy probe and treating the CPU cost as free