skip to content

After an edit to daemon.json, dockerd will not start. How do you diagnose and recover?

level: seniorimportance: should knowfreq 54%

answer

  1. The CLI needs the process that died
  2. Ask the init system, not Docker
  3. The error names the offending directive
  4. Restore service before finishing the diagnosis
  5. Parse-check the file before applying it

basics

~20 s

Read the daemon's journal with journalctl -u docker: dockerd prints the exact configuration error before exiting - malformed JSON, an unrecognised key, or an option set both in daemon.json and as a unit flag. Move the file aside to recover.

solid answer

~40 s

The engine is down, so `docker` commands cannot help — go to the service. `systemctl status docker` shows it in a restart loop, and `journalctl -u docker --no-pager -n 50` carries the real message, which `dockerd` prints just before exiting. The recurring causes are all in the file: invalid JSON (a trailing comma, a comment), a key the daemon does not recognise, an option specified both in `daemon.json` and as a flag on the unit's `ExecStart` line, or a `storage-driver` the host cannot initialise. Recovery first, forensics second: move the file aside, `systemctl start docker`, and containers with a restart policy come back. Then fix the file offline and check it with `dockerd --validate` before putting it back. If the journal is thin, run `dockerd --debug` in the foreground and read the error directly.

code

bash · 5 lines
bash
systemctl status docker --no-pager
journalctl -u docker --no-pager -n 50
sudo mv /etc/docker/daemon.json /etc/docker/daemon.json.bad
sudo systemctl start docker
sudo dockerd --validate --config-file /etc/docker/daemon.json.bad

go deeper

for a junior

Know that a dead engine means the docker CLI cannot answer anything, and that the daemon's own error goes to the system journal. Being able to run systemctl status and journalctl for the docker service is the bar.

for a middle

Explain the common causes and what each error text means: invalid JSON, an unknown key, a directive set both in the file and as a unit flag. Know that dockerd --validate checks a file without starting an engine.

for a senior

Show the operator's ordering — restore service by moving the file aside, then diagnose from the journal, then validate before reapplying — and account for the workloads that did not come back on their own.

for a principal

Own the prevention: engine configuration rendered and validated by tooling, applied to a canary host first, and a rollback that does not depend on someone being able to log in to a broken box.

## The situation A host runs a subscription-billing cron in a container built from a 1.7 GB Elixir release image. Someone adds a registry mirror to `/etc/docker/daemon.json`, restarts the daemon, and now nothing on the host is running — `docker ps` answers `Cannot connect to the Docker daemon at unix:///var/run/docker.sock. Is the docker daemon running?`. That error is not a diagnosis; it is the CLI telling you it has no one to talk to. The diagnosis lives with the service. ## Step 1 — ask the init system, not the CLI ``` systemctl status docker journalctl -u docker --no-pager -n 50 ``` `systemctl status` shows the unit failed or looping (`Active: activating (auto-restart)` with a non-zero result). The journal holds the daemon's own stderr, and a configuration failure is always printed there before the process exits. This is the single highest-value habit for this failure mode: the engine is not silent, it is just not talking through the CLI. ## Step 2 — read the message, which names the fault The messages are specific and each points at a different mistake: - **Malformed JSON.** `unable to configure the Docker daemon with file /etc/docker/daemon.json: invalid character '}' looking for beginning of object key string`. A trailing comma after the last key, a `//` comment, or an unquoted value. JSON permits none of them. - **A key the daemon does not know.** The daemon reports that directives in the file do not match any configuration option, and names them. Usually a typo (`storage-drivers`, `registry-mirror`) or a key copied from a different tool's configuration. - **A directive set twice.** `the following directives are specified both as a flag and in the configuration file: hosts`. The shipped systemd unit passes flags on its `ExecStart` line — commonly `-H fd://` — and setting `"hosts"` in `daemon.json` collides with it. The daemon refuses to guess and exits. `systemctl cat docker` shows exactly which flags the unit passes. - **A driver that cannot initialise.** A `storage-driver` value the host's kernel or filesystem cannot support produces a graphdriver initialisation error at start-up. - **A path that does not work.** A `data-root` pointing at a directory that does not exist, is not writable, or is on a filesystem the chosen storage driver rejects. ## Step 3 — restore service before you finish the investigation The billing cron is not running while you read logs. Recovery is deliberately blunt: ``` sudo mv /etc/docker/daemon.json /etc/docker/daemon.json.bad sudo systemctl start docker ``` With no configuration file, the daemon starts on its defaults. Containers whose restart policy is `always` or `unless-stopped` come back on their own; anything created with the default policy needs `docker start`. Keep the broken file — it is the evidence, and you will edit it rather than retype it. This is also the moment to notice what the outage cost: a cron-style container that missed its window does not catch up by itself, so check whether the run needs to be triggered manually once the engine is healthy. ## Step 4 — fix it with a check that does not risk the engine `dockerd --validate` parses the configuration file and exits without starting an engine, so it is safe to run on a host that is currently serving. It catches malformed JSON and unrecognised keys — most of what breaks a start-up — before you apply anything. When the journal is unhelpfully terse, run the daemon in the foreground with `dockerd --debug` and read its output directly; stop it with Ctrl-C and let systemd own the process again. A sanity check on the file itself costs nothing: `python3 -m json.tool /etc/docker/daemon.json` will reject a trailing comma faster than a restart will. ## Step 5 — separate the file from the unit Not every daemon start-up failure is the JSON. If someone has added a systemd drop-in under `/etc/systemd/system/docker.service.d/`, its flags are merged into `ExecStart`, and an edit there does nothing until `systemctl daemon-reload` is run. `systemctl cat docker` prints the effective unit including drop-ins — read it before concluding the file is at fault. The rule to keep is: set an option in `daemon.json` *or* on the command line, never both. ## Why this is a senior question It tests three things at once: that you know the engine's failure is visible in the host's service logs rather than through the Docker CLI; that you restore service before completing the diagnosis, because a daemon that will not start is a full-host outage and not a puzzle; and that you close the loop with a validation step so the same edit cannot take the host down twice. The candidate who reaches for `docker logs` here — a command that requires the daemon that is dead — has not operated a real host.

  • What does the error 'specified both as a flag and in the configuration file' mean, and how do you fix it?
    It means the same option is set twice: once on the `dockerd` command line in the systemd unit, and once as a key in `daemon.json`. The daemon refuses to pick a winner and exits. `systemctl cat docker` shows the unit's `ExecStart` flags and any drop-ins; remove the setting from one side — usually leave the unit alone and drop the key from the file, or edit the unit through a drop-in and re-run `systemctl daemon-reload`.
  • The daemon starts again but docker images is now empty. What happened?
    Almost always a changed `storage-driver` or `data-root`. Neither migrates anything: the daemon initialises a fresh store and simply cannot see images written under the previous driver or in the previous directory. The old data is still on disk. Reverting the key and restarting brings the images back; if the move was intended, the content has to be copied deliberately rather than assumed.
  • How do you get diagnostics when the journal shows nothing useful?
    Stop the unit and run the daemon in the foreground: `sudo dockerd --debug`. It prints configuration and start-up errors straight to the terminal, including the ones a supervisor can swallow. `dockerd --validate` is the non-destructive version when you only want the configuration checked. Once you have the message, stop the foreground process and let systemd manage the daemon again.

saying these in an interview costs you the question

  • Runs docker logs to debug a daemon that will not start
  • Never opens journalctl or systemctl status for the service
  • Deletes /var/lib/docker to make the daemon start again
  • Reinstalls the engine instead of reading the error message
  • Keeps restarting the service without changing anything
  • Edits the file again and restarts straight into production

context