When a container keeps crashing seconds after it starts, how does the Docker daemon pace its restart attempts, what limits the number of attempts, and what resets that pacing?
answer
- 100 ms, doubling, capped near a minute
- only `on-failure` takes `:N`
- N counts *consecutive* attempts
- ~10 s uptime = successful start → reset
- RestartCount + exit code in docker inspect
basics
~20 sThe daemon waits an exponentially growing delay between attempts — roughly 100 ms, doubling each time, up to a cap of about a minute — so a crash loop doesn't hammer the host. on-failure:N caps attempts; the delay and the counter reset once the container runs successfully for a short while.
solid answer
~60 sDocker does not retry in a tight loop. After each failed start the daemon waits, and the wait grows exponentially — starting around 100 ms and doubling per attempt, capped (in current engines) at roughly one minute. So a container that dies instantly is retried after ~0.1 s, ~0.2 s, ~0.4 s and so on, quickly settling into slow retries. Two things bound it: - **`on-failure:N`** caps the number of consecutive restart attempts; when the budget is exhausted the container is left exited. `always` and `unless-stopped` have no cap — they retry forever, just slowly. - **A successful run resets state.** Once the container stays up for a short period (about ten seconds), the daemon treats it as having started successfully: the backoff delay drops back to its base and the retry budget is refilled. That's why a flapping container can burn far more than N restarts over a day. Observe it with `docker inspect -f '{{.RestartCount}} {{.State.Status}}'` and `docker events --filter event=restart`. The policy also only engages after the container has started successfully once, so a container that can never start doesn't spin at full speed.
code
bash · 8 linesdocker run -d --name flap --restart=on-failure:5 alpine sh -c 'sleep 1; exit 1'
docker ps -a --filter name=flap --format '{{.Status}}'
# Restarting (1) 3 seconds ago
docker inspect -f '{{.RestartCount}} {{.State.ExitCode}}' flap
docker events --filter container=flap --filter event=start --filter event=diego deeper
Know that Docker waits longer and longer between restarts and that only on-failure takes a retry count.
Give the shape of the curve (about 100 ms, doubling, capped) and explain that a successful run of roughly ten seconds resets both the delay and the counter.
Turn it into a diagnostic routine — RestartCount, exit code, docker logs, docker events — and point out that a retry budget bounds tight loops, not slow flaps.
Argue about where failure-rate policy belongs: Docker can only express consecutive-attempt caps, so anything needing failure budgets over a window, alerting, or dependency-aware startup ordering belongs to a supervision layer above the daemon.
## Why pacing exists A container whose process dies immediately — bad config, missing env var, a database that isn't up yet — would, under naive supervision, be restarted thousands of times a second. Each restart means creating a task in the runtime, setting up namespaces, mounting the filesystem, and writing log lines. Unpaced, one broken container can saturate a CPU, fill a disk with logs, and starve healthy containers on the same host. Every serious supervisor therefore backs off, and the Docker daemon is no exception. ## The backoff curve The daemon's restart manager keeps a per-container delay. It starts at about **100 milliseconds** and **doubles** after each restart, with a **cap in the region of one minute** in current engines. In practice: ~0.1 s, 0.2 s, 0.4 s, 0.8 s, 1.6 s … and within about ten attempts the container is being retried roughly once a minute. This is exponential backoff without jitter — jitter matters for many clients hitting one server, but here the retries are local to a single host. The practical consequence is a diagnostic one. If you `docker ps -a` on a crash-looping container you may catch it in `Restarting (1) 12 seconds ago`; the number in parentheses is the **last exit code**, and `docker inspect -f '{{.RestartCount}}'` gives the cumulative count of daemon-initiated restarts. A container sitting in `Restarting` for a long time between attempts is not stuck — it is deep into the backoff curve. ## The retry budget Only `on-failure` takes a retry limit: `--restart=on-failure:5`. It caps **consecutive** restart attempts. When the budget runs out the daemon stops trying and the container stays `exited` with its last exit code — no error is surfaced anywhere except the container state, which is why crash loops need alerting on container state rather than on the CLI. `always` and `unless-stopped` have **no** cap. They retry indefinitely; the only thing bounding the damage is the backoff. A common misconception is that `--restart=always:5` exists — it does not; the count is only meaningful for `on-failure`. ## What resets the state Both the growing delay and the retry counter are reset when the container **runs successfully**. The daemon's criterion is duration: if the container stays up for roughly ten seconds it is considered to have started properly, and the restart manager resets its delay to the base value and clears the accumulated count. This same rule is why the documentation says a restart policy "only takes effect after a container starts successfully" — a container that never manages to run for that window doesn't get treated as a healthy workload being supervised. This reset behaviour has a sharp operational edge: a container that runs for thirty seconds and *then* dies never accumulates consecutive failures. With `on-failure:3` it will be restarted three times, run for thirty seconds, reset, be restarted three more times, and so on — forever. The retry budget bounds a *tight* crash loop, not a slow flap. If you need "give up permanently after N failures across time", Docker's restart policy cannot express it; you need an orchestrator or external supervision with a failure window. Manual intervention also resets things: `docker stop` clears the supervision state for that container, and `docker start` begins from a clean slate. ## Diagnosing a crash loop The workflow is: `docker ps -a` to see the `Restarting`/`Exited (code)` status; `docker inspect -f '{{.State.ExitCode}} {{.RestartCount}} {{.State.Error}}'` for the numbers; `docker logs --tail 50 <id>` for the process's own last words — logs persist across restarts, so you do not need to catch it live; and `docker events --filter container=<id>` to watch the die/restart pairs in real time and confirm the widening gaps. The fix is almost never to tune the policy. A crash loop means the process cannot run: a missing mount, an unreachable dependency, a permission problem after switching to a non-root user, or an OOM kill (exit 137 with `OOMKilled: true`). The restart policy is only the thing making the failure repeat visibly. ## Contrast with orchestrators Orchestrators implement the same idea with different constants and much better visibility — a Kubernetes kubelet, for instance, backs off restarts of a failing container and surfaces the state explicitly to the user rather than leaving it in a container's inspect output. If you are choosing between Docker's local backoff and an orchestrator's, the difference that matters is observability and policy expressiveness, not the mathematics of doubling a delay.
- With `--restart=on-failure:3`, can a container end up being restarted more than three times over its life?Yes, easily. The limit is on consecutive attempts, and the counter is reset once the container runs successfully for about ten seconds. A service that starts, works for a minute, then crashes will get a fresh budget of three every cycle, so it can flap indefinitely. Bounding failures over a time window requires supervision above Docker.
- A container has been sitting in the `Restarting` state for over a minute. Is it stuck?Almost certainly not — it is deep in the exponential backoff, where the delay has grown to roughly the one-minute cap, so the daemon is simply waiting before the next attempt. Confirm with `docker events` to see the die/start pairs and with `docker logs` to read why the process keeps exiting. Treat it as a crash loop to diagnose, not as a hung daemon.
saying these in an interview costs you the question
- Thinking Docker retries instantly in a tight loop with no delay.
- Believing `--restart=always:5` is valid — only `on-failure` accepts a count.
- Assuming the retry limit is a lifetime total rather than a consecutive-attempt count.
- Saying the backoff grows without bound instead of hitting a cap.
- Reading `Restarting (137)` as "137 restarts" rather than the last exit code.