skip to content

How would you roll out a daemon.json change that requires a dockerd restart across a production fleet?

level: principalimportance: nice to knowfreq 30%

answer

  1. The unit of risk is the whole host
  2. Never a hand edit; render and review it
  3. Do the free part before the expensive part
  4. Batch by service replicas, not by host count
  5. Rollback cannot rely on the engine

basics

~20 s

Restarting dockerd stops every container on that host unless live-restore is already on, and a bad key stops the daemon entirely. Render and validate the file from configuration management, canary one host, then roll in batches with a rollback that works without Docker.

solid answer

~50 s

The unit of risk is the host, not the container: one bad key stops `dockerd`, and a restart takes every workload on that machine with it. So the rollout is staged. The file is rendered by configuration management from one reviewed template, never hand-edited, and every host runs `dockerd --validate` before anything is applied. Enable `live-restore` first — it is reloadable, so it costs nothing — and let it protect the restarts that follow. Then roll: one canary host, verify with `docker info` and by watching the workloads for a full duty cycle, then batches sized so no service loses more than one replica at a time, draining or rescheduling first where the workload cannot be interrupted. Rollback must not depend on the engine: reverting the file and restarting the service is the plan, and the old file stays on disk. Schedule around anything that runs on a clock, because a missed cron window does not retry itself.

code

bash · 5 lines
bash
set -euo pipefail
sudo dockerd --validate
sudo systemctl restart docker
docker info --format '{{.DockerRootDir}} {{.LiveRestoreEnabled}}'
docker ps --format '{{.Names}}\t{{.Status}}'

go deeper

for a junior

Focus on the fact underneath this: restarting the Docker daemon is not free, because the containers it supervises stop with it. That is why engine configuration changes are planned rather than done ad hoc.

for a middle

Be able to split the change into the part a reload applies and the part that needs a restart, and to name what you would check afterwards to prove the daemon adopted the new settings.

for a senior

Show the mechanics of a safe roll: validate before applying, canary, batch, verify each host with real signals, and account for containers whose restart policy will not bring them back.

for a principal

Own the strategy — configuration as a reviewed artefact, batches sized by service availability rather than host count, live-restore as fleet policy, and a rollback path that assumes the engine on the host is dead.

## Frame the blast radius first A `daemon.json` change is not like a deployment. A deployment fails one workload; a bad engine configuration fails one *host* and everything on it, because `dockerd` refuses to start on a file it cannot parse. And even a perfectly good change costs a daemon restart, which without `live-restore` stops every running container on that machine. Take a concrete fleet: 37 engine hosts, and among the workloads a subscription-billing cron container built from a 1.7 GB Elixir release image that fires at seven minutes past each hour and does not retry a missed window. The rollout has to be designed around both facts — the risk that a host dies, and the certainty that each host has an interruption. ## Make the file an artefact, not an edit The first decision is that nobody edits `/etc/docker/daemon.json` on a host. Configuration management renders it from one reviewed template, so every host in a role has a byte-identical file and drift is detectable. That buys three things: the change is reviewed once rather than typed 37 times; the previous version is recoverable from version control rather than from memory; and the set of hosts that have the new file is a fact you can query rather than a belief. Pair it with a pre-flight gate. `dockerd --validate` parses a configuration file and exits without starting an engine, so it can run on a live host safely, and it catches the malformed JSON and unknown keys that cause most start-up failures. Wire it into the change so a host that fails validation is skipped rather than restarted. Also check for the duplicate-directive trap: if the systemd unit's `ExecStart` already passes a flag, setting the same option in the file makes the daemon exit. `systemctl cat docker` shows the effective unit including drop-ins. ## Sequence the change to shrink the interruption Split the change into what reloads and what does not. Reloadable keys — registry mirrors, insecure registries, debug level, live-restore, concurrency limits — go first with a SIGHUP, which touches nothing. In particular, turn `live-restore` on ahead of time. It is reloadable, and it governs how the *outgoing* daemon shuts down, so it only protects restarts performed by a daemon that already had it enabled. Getting it on early turns the later restarts from outages into pauses in the control plane. Be honest about its limits when you plan: it does not apply to hosts that have joined a swarm; it is meant for patch-level engine upgrades rather than arbitrary version jumps; and while the daemon is down nobody is draining container output, health checks are not evaluated and the event stream is dead, so a long absence can block a chatty container's writes. Live-restore shortens an outage, it does not abolish it. ## Roll in batches sized by service, not by host One canary host first, chosen to carry a representative workload but not a singleton. Verify with evidence rather than optimism: `docker info` for the effective Storage Driver, Docker Root Dir, Registry Mirrors and Live Restore Enabled; `docker ps` for the containers that were supposed to come back; and the service's own signals for a full duty cycle — for a cron-shaped workload that means letting a real scheduled run complete, not just watching the process start. Then batch by the constraint that matters: no service may lose more than one replica at a time. That is a property of the placement, not of the host list, so the batches are computed from where the replicas are. Workloads that cannot be interrupted at all — the billing cron mid-run — are drained, rescheduled, or the host is done in a window where they are idle. Containers created with the default `--restart no` policy do not come back after a restart; either fix the policy first or include an explicit start step, because "the daemon is up" is not the same as "the host is serving". ## Plan the rollback for a host you cannot reach through Docker The failure you are guarding against is an engine that will not start, which means every Docker-based tool is unavailable on that host. So the rollback must be an operating-system operation: restore the previous file and restart the service. Keep the prior version on disk, and know that the blunt recovery — move the file aside and start the daemon on its defaults — restores service on any host, at the cost of that host running unconfigured until you put a good file back. Halt the roll on the first failure rather than discovering it 37 times. ## What is really being asked This question is about owning a change whose failure mode is host-wide and whose success still costs an interruption. The strong answer separates the reloadable part from the restart part, gets live-restore in place before it is needed, sizes batches by service availability rather than by convenience, validates before applying, and has a rollback that does not assume the engine works. The weak answer is a configuration-management push to all hosts at once, discovered by the on-call pager.

  • Why enable live-restore in a separate, earlier step rather than in the same change?
    Because it governs the shutdown behaviour of the daemon that is already running. A daemon started without it kills its containers when it exits, no matter what the file on disk says. It is a reloadable key, so applying it with a SIGHUP costs nothing and interrupts nothing — and once every host's live daemon has it, the restarts that follow are far cheaper.
  • What is involved in moving the daemon's data-root to a different disk across the fleet?
    It is a stop-copy-start operation, not a reload. Stop the daemon, copy the existing tree preserving hard links, ownership and extended attributes, set `data-root`, then start and verify with `docker info`. Nothing migrates automatically: point the daemon at an empty directory and it starts clean, with every image and volume invisible though still on the old disk. Budget for the copy time and roll one host at a time.
  • How do you keep engine configuration from drifting across hosts over time?
    Render the file from one template per host role and let configuration management enforce it, so an out-of-band edit is reverted and reported. Verify from the daemon's side rather than the file's — periodically compare `docker info` output against expected values — because that catches the case where the file is correct but the running daemon never picked it up.

saying these in an interview costs you the question

  • Pushes the new daemon.json to every host at once
  • Ignores that a daemon restart stops the host's containers
  • Assumes live-restore protects the restart that enables it
  • Has no rollback that works without a working engine
  • Treats data-root as a value you can simply change
  • Edits daemon.json by hand on individual production hosts

context