What happens to running containers when the Docker daemon (`dockerd`) itself is stopped, restarted or upgraded, and what does the daemon's `live-restore` option change about that?
answer
- default: dockerd stop = containers stop
- daemon start reconciles via restart policy
- live-restore in daemon.json
- patch upgrades only, no Swarm
- no supervision / no health checks while daemon down
basics
~20 sBy default, stopping the daemon stops all running containers; on daemon start they come back only if their restart policy says so. Enabling live-restore in daemon.json leaves containers running while the daemon is down — but only for patch-level daemon upgrades, and not in Swarm mode.
solid answer
~50 sBy default the daemon owns the container lifecycle: when `dockerd` stops, it shuts running containers down, so a daemon restart or engine upgrade is a workload outage. On startup the daemon reconciles state — containers with `always` come back, `unless-stopped` come back only if they were running, `on-failure`/`no` follow their own rules. Setting `"live-restore": true` in `/etc/docker/daemon.json` decouples them: container processes (managed by containerd and the shims below it) keep running while the daemon is down, and the new daemon re-attaches to them. Caveats that matter in production: - Supported for **patch-level** daemon upgrades only; major/minor upgrades may not re-attach cleanly. - **Not supported in Swarm mode.** - While the daemon is down there is **no API, no health-check evaluation, no restart supervision** — a container that dies stays dead until the daemon returns and reconciles. - Container stdout/stderr goes into a FIFO buffer; a long outage can fill it and block the container's writes. So `live-restore` protects against daemon restarts, not against container failures.
code
bash · 9 linescat /etc/docker/daemon.json
# { "live-restore": true }
sudo systemctl reload docker
docker info --format '{{.LiveRestoreEnabled}}'
# containers stay up across a daemon restart
sudo systemctl restart docker
docker ps --format '{{.Names}}\t{{.Status}}'go deeper
Know that by default restarting the Docker service stops your containers, and that they come back only if they have a restart policy.
Add the reconciliation step at daemon start and the existence of the live-restore daemon option with its patch-upgrade and Swarm limitations.
Explain the containerd/shim layering that makes live-restore possible and be explicit about the supervision gap, log buffering, and how you'd sequence an engine upgrade.
Separate the failure domains — daemon availability versus container failure versus host failure — and argue for redundancy above the host (drain and reschedule) over squeezing uptime out of a single daemon.
## The default coupling On a stock Docker installation the daemon and the containers share a fate. `systemctl stop docker` (or an engine package upgrade that restarts the unit) causes the daemon to shut down its running containers — they receive the stop sequence and exit. For a single-host deployment that means every engine upgrade is a scheduled outage for every workload on the box. When the daemon starts again it performs a reconciliation pass over its container store and applies each container's restart policy: - `always` → started, even if it had been deliberately stopped before; - `unless-stopped` → started only if it was running when the daemon went away; - `on-failure` → containers that were running are brought back up; - `no` → left alone. This reconciliation is exactly why restart policies are persisted in the container's host config rather than held in memory, and it is the concrete reason `always` and `unless-stopped` differ at all. ## What the runtime stack actually looks like Understanding `live-restore` requires knowing that `dockerd` is not the container's parent process. The stack is: `dockerd` (API, images, networks, build) → `containerd` (container supervision) → a per-container shim → `runc`, which sets up namespaces and cgroups and then exits, leaving the container's PID 1 parented to the shim. Because the shim is what actually holds the container, the container process can survive `dockerd` exiting — the coupling in the default configuration is a *policy decision* by the daemon, not a technical necessity. ## Enabling live-restore Add to `/etc/docker/daemon.json`: ``` { "live-restore": true } ``` and reload the daemon. From then on, when `dockerd` stops, running containers stay running; when it comes back, it re-discovers them through containerd and re-attaches to their shims, log streams and state. Users see uninterrupted traffic through a daemon restart. ## The caveats that decide whether you can use it **Patch upgrades only.** The documented support boundary is patch-level daemon versions. Across a major or minor version change, on-disk state and shim interfaces can change enough that re-attachment isn't guaranteed, and the daemon may end up stopping the containers anyway. Plan minor/major engine upgrades as drains, not as live restores. **Not with Swarm mode.** A daemon participating in a swarm cannot use live-restore; swarm's own reconciliation loop assumes the daemon is authoritative and present. If a node's daemon disappears, the manager reschedules tasks elsewhere — which is the orchestrator solving the same problem a different way. **No supervision during the gap.** This is the point most candidates miss. While the daemon is down there is no API, no `docker` CLI, no health-check execution, no event stream, and crucially **no restart-policy enforcement**. If a container crashes during that window, nothing restarts it; it simply stays dead until the daemon returns and reconciles state. Live-restore keeps healthy containers alive; it does not extend supervision. **Log buffering.** Container stdout/stderr is written through a FIFO that the daemon normally drains into the logging driver. With the daemon absent, that buffer fills. A chatty container during a long daemon outage can block on writes — a subtle way an apparently "live-restored" workload stalls. **Config changes need care.** Some daemon-level changes (notably the default logging driver or bridge networking parameters) don't apply cleanly to containers restored from a previous daemon; you may find restored containers keeping old settings until they're recreated. ## Choosing the approach Single-host, uptime-sensitive services with routine patch upgrades: turn `live-restore` on and treat engine patching as a live operation, while still testing minor upgrades as drains. Multi-host with an orchestrator: don't bother — drain the node, let the orchestrator move workloads, upgrade, and bring the node back. Redundancy above the host is a stronger answer than surviving a daemon restart on one box. Either way, be explicit in an interview about the layering: `live-restore` addresses *daemon* availability; restart policies address *container* failure. They are different failure domains, and one does not cover the other. A container with `--restart=unless-stopped` on a `live-restore` host survives both a crash (policy restarts it) and a daemon patch upgrade (live-restore keeps it up) — but not a crash *during* the daemon upgrade, which is a small, real gap worth naming.
- With `live-restore` enabled, a container crashes while the daemon is being upgraded. Does its `--restart=always` policy bring it back?Not during the outage — restart policies are enforced by the daemon, and with the daemon down nothing is watching for exits. The container stays dead until `dockerd` comes back, reconciles state, and applies the policy, at which point it is started again. So live-restore narrows the outage for healthy containers but leaves a supervision gap for failing ones.
- Why can containers keep running at all while `dockerd` is stopped?Because `dockerd` is not the container's parent. It delegates to containerd, which starts a per-container shim; `runc` configures namespaces and cgroups then exits, leaving the container's PID 1 held by the shim. The shim and containerd survive the daemon exiting, so the default behaviour of stopping containers is a daemon policy choice that `live-restore` simply turns off.
saying these in an interview costs you the question
- Assuming containers always keep running through a daemon restart — without live-restore they don't.
- Believing live-restore keeps restart policies and health checks working while the daemon is down.
- Claiming live-restore covers major and minor engine upgrades, or works in Swarm mode.
- Confusing live-restore with checkpoint/restore or live migration of containers.
- Forgetting that log output buffers and can block a container during a long daemon outage.