skip to content

A workload already runs under an outer supervisor — a host init system such as systemd, or a cluster orchestrator that reschedules failed workloads. How do you decide what the container-level `--restart` policy should be, and where should responsibility for restarting live?

level: principalimportance: should knowfreq 30%

answer

  1. one control loop per workload
  2. systemd Restart= vs docker --restart → pick one
  3. orchestrator-managed → `no`
  4. docker sees exit codes only, not health/capacity/placement
  5. consecutive-attempt cap ≠ failure budget over a window

basics

~20 s

Exactly one layer should own restarts. If an outer supervisor starts and stops the container, set the container policy to no — otherwise two supervisors fight, and a workload the outer layer stopped keeps coming back. Docker's policy is for plain single-host deployments.

solid answer

~50 s

The rule is **one owner per workload**. Docker's restart policy is host-local, exit-code-driven supervision with no notion of desired replica count, dependencies, health or capacity. An outer supervisor has richer intent, so it should win. - **Under systemd** (a unit that runs the container): use `--restart=no` and express restarts as `Restart=on-failure` in the unit. Two supervisors on one container produce a container that refuses to stay stopped and a unit whose reported state no longer matches reality. - **Under a cluster orchestrator**: the node agent supervises containers per the workload spec and the control plane handles rescheduling, backoff and health-driven restarts. A container-level `always` here is at best redundant and at worst masks failures the orchestrator should see. - **Plain single host, no outer supervisor**: Docker's own policy *is* the supervision — `unless-stopped` for services, `on-failure:N` for jobs. Also weigh what restart can't fix: restarting doesn't repair a bad config or a dependency outage, and unbounded local retries can hide a failure that should page someone.

code

bash · 8 lines
bash
# /etc/systemd/system/api.service
# [Service]
# Restart=on-failure
# RestartSec=5
# ExecStart=/usr/bin/docker run --rm --name api --restart=no myorg/api:1.4
# ExecStop=/usr/bin/docker stop api

sudo systemctl daemon-reload && sudo systemctl enable --now api

go deeper

for a junior

Know that if something else already restarts your container, Docker's policy should be no so the two don't conflict.

for a middle

Explain the concrete conflict with a systemd unit and pick the right policy for each of the three cases (systemd-owned, orchestrator-owned, plain host).

for a senior

Add what Docker's policy cannot see — health, capacity, placement, failure windows — and describe how you'd audit a fleet and migrate ownership without downtime.

for a principal

Lead with the one-control-loop principle, then argue the harder questions: whether restarting is the right action at all, where the failure budget lives, and how restart counts feed alerting so self-healing doesn't hide an SLO breach.

## The question behind the question "Which restart policy?" is really "who owns the desired state of this workload?". Restarting is a control loop: observe actual state, compare to desired state, act. Running two such loops over the same object is the classic distributed-systems mistake — they disagree, and the disagreement shows up as flapping, phantom restarts, or a workload that cannot be taken down. ## What Docker's restart policy can and cannot express The daemon's policy is deliberately small. It knows: did the main process exit, with what code, how many consecutive times, and was the stop deliberate. From that it can restart with exponential backoff and an optional consecutive-attempt cap. It does **not** know: how many instances should exist, whether the process is alive-but-wedged (health checks mark a container unhealthy but never restart it), whether a dependency is down, whether the host has capacity, or whether the failure has already exceeded a budget that should stop retrying and alert instead. It also cannot move a workload — if the host is the problem, restarting locally is the wrong action taken forever. That gap defines when you want something else on top. ## Case 1: a host init system as the outer supervisor A common pattern is a systemd unit that runs a container. If the unit has `Restart=on-failure` **and** the container has `--restart=always`, both loops act on the same container. Symptoms: you `systemctl stop` the unit, the daemon restarts the container anyway; or systemd restarts the unit while the daemon is already retrying, producing duplicate or name-conflicting containers. Docker's own guidance is not to combine restart policies with an external process manager. The resolution is to pick a side. If you want systemd to be the source of truth — because it gives you ordering (`After=`), dependency wiring, resource control, and journal integration — set `--restart=no` and let `Restart=` in the unit do the work. If you'd rather the daemon own it, make the unit a one-shot that only creates the container, and use `unless-stopped`. ## Case 2: a cluster orchestrator as the outer supervisor Under an orchestrator the node agent supervises containers according to the workload spec's own restart semantics, and the control plane layers on health-driven restarts, backoff, rescheduling to other nodes, and replica reconciliation. The container-level `--restart` flag is not part of that contract — the orchestrator creates containers itself and expresses restart intent in its own API. Setting `always` on containers it manages either has no effect or interferes with the agent's view of the world. The more interesting principal-level point is that the orchestrator's supervision is *different in kind*, not just in quality: it can decide that restarting on this node is futile and place the workload elsewhere, and it can surface a crash loop as a first-class, alertable state rather than as a growing counter in a container's inspect output. That observability difference is usually the real argument for moving supervision up. Swarm is the in-between case: a Swarm *service* uses its own restart policy (`condition: none | on-failure | any`, with delay, max attempts and a window), evaluated by the orchestrator; the container-level flag is ignored for service tasks. Note that a per-window failure budget is expressible there and *not* expressible with the container-level flag, whose cap only counts consecutive attempts. ## Case 3: no outer supervisor On a plain host with no orchestrator, Docker's policy is the whole supervision story, and it is a perfectly reasonable one. `unless-stopped` for long-running services; `on-failure:N` for jobs that should retry but be allowed to finish; `no` for one-shot tasks driven by a scheduler or CI. Pair it with `live-restore` if engine patching must not take workloads down. ## The judgement calls a principal is really being asked about **Should this restart at all?** Restarting is only correct when the failure is likely transient — a race at startup, a dependency that came up late, a rare memory corruption. Restarting a container with a bad config file forever converts a loud, diagnosable failure into a quiet loop. For unrecoverable configuration errors, failing fast and staying down (with alerting on the exited state) is better than infinite retries. **Where's the failure budget?** Docker can cap consecutive attempts but not failures over a window, so "restart up to five times an hour, then page" cannot be expressed at the container level. If that's the policy you need, it belongs one layer up. **Does restarting mask an SLO breach?** A workload that restarts every ninety seconds may look "up" on a naive check while dropping in-flight requests each cycle. Restart counts should be an emitted, alertable metric, not just a self-healing convenience. **Is restart hiding a dependency-ordering problem?** Teams often use restart loops as a crude wait-for-dependency mechanism. It works, but it converts a modelling problem into perpetual noise; explicit readiness handling or startup retry logic inside the application is cleaner and gives better error messages. A strong answer names the principle (one control loop per workload), maps it onto the three cases above, and then moves past mechanics to the failure-budget and observability questions.

  • What concretely goes wrong if both systemd and Docker are set to restart the same container?
    The two loops act on the same object with different views of intent. Stopping the systemd unit can leave the daemon restarting the container behind systemd's back, so the workload refuses to stay down; conversely systemd may relaunch the run command while the old container still exists, producing name conflicts or duplicate instances. The unit's reported state also stops reflecting reality, which breaks any alerting built on it.
  • When is choosing *not* to restart the better engineering decision?
    When the failure is deterministic — a malformed config, a missing secret, an incompatible schema — restarting cannot succeed and merely converts a clear failure into a noisy loop that delays diagnosis. Failing fast and leaving the container exited, with alerting on that state, gets a human involved sooner. Restart policies pay off for transient faults, not for broken inputs.
  • How would you express "retry up to five times per hour, then stop and alert"?
    Not with Docker's container-level policy: `on-failure:N` caps *consecutive* attempts and resets once the container runs successfully, so it cannot bound failures over a time window. You need supervision above the daemon — a Swarm service restart policy with a window, an orchestrator's crash-loop handling plus alerting, or an external watchdog that watches restart counts and stops the workload when the budget is spent.

Two supervisors on one container is like two thermostats wired to the same furnace: each is individually sensible, and together they cycle the system endlessly because neither can see the other's intent.

saying these in an interview costs you the question

  • Setting `--restart=always` on containers an orchestrator or systemd already manages.
  • Treating restart policy as a substitute for health checks or readiness/dependency handling.
  • Believing `on-failure:N` gives a failure budget over time rather than a consecutive-attempt cap.
  • Assuming more aggressive restarting is always safer, ignoring that it can mask a broken config or an SLO breach.
  • Claiming the container-level `--restart` flag governs Swarm service tasks.

context