What goes wrong when a systemd unit runs `docker run -d` and the container also sets --restart=always?
answer
- The command you ran is only a client
- Detached means the client exits at once
- The unit tracks the wrong process
- Two watchdogs, one container, no agreement
- Foreground run plus an explicit ExecStop
basics
~20 sTwo supervisors watch one container and systemd watches the wrong process. docker run -d exits at once, so the unit looks dead while the container runs, and the engine's restart policy revives a container systemd stopped.
solid answer
~50 sThe `docker` command is a client: it posts to the daemon over `/var/run/docker.sock`, and the container's real host parent is a containerd shim the daemon started, not the CLI and not systemd. With `-d` the client exits as soon as the container is created, so a `Type=simple` unit sees its main process die and marks the service dead — and a `Restart=` line then relaunches it, hitting `Conflict. The container name … is already in use` or spawning duplicates. Meanwhile `--restart=always` means the engine also considers itself responsible: stop the unit for maintenance and the container can reappear when dockerd restarts, with systemd convinced it is stopped. Pick one owner. Either the engine supervises it and there is no unit, or systemd supervises it: run in the foreground, no `--restart`, an explicit `ExecStop=/usr/bin/docker stop`, and a pre-start `docker rm -f` to clear stale names.
code
bash · 8 lines# the client exits immediately; the container is elsewhere
docker run -d --name invoice-worker invoice-render:2026.08.19-7c41ab2
echo "client exited with $?" # prints at once, container still up
# the host PID of the container's PID 1, and who its parent really is
pid=$(docker inspect -f '{{.State.Pid}}' invoice-worker)
ps -o pid,ppid,comm -p "$pid"
ps -o pid,comm -p "$(ps -o ppid= -p "$pid" | tr -d ' ')" # a containerd shimgo deeper
Know that docker run -d returns immediately and that the container keeps running afterwards, so a script or unit that expects the command to stay in the foreground will misread the service as finished.
Explain the layering — client, daemon, containerd shim — and why a Type=simple unit around a detached run reports dead while the container serves. State the rule that only one supervisor may own a container.
Demonstrate the failure modes that only appear at boot or during maintenance: name conflicts on restart, a stopped unit whose container returns when the daemon restarts, duplicate workers double-consuming a queue. Show the unit you would actually deploy.
Decide the fleet-wide convention and defend it: whether hosts run containers under the engine's own policy or under units, what that buys in ordering and observability, and how you keep every host from inventing its own supervision story.
### The process systemd can see is not the process you care about `docker` is a thin HTTP client. `docker run` opens `/var/run/docker.sock`, asks the daemon to create and start a container, and — in detached mode — prints the container ID and exits. The daemon hands the actual work to containerd, which starts a `containerd-shim-runc-v2` process; that shim, not the client, is the host parent of the container's PID 1, and it deliberately survives so containers outlive a daemon restart of the *client*. Check it yourself: `docker inspect -f '{{.State.Pid}}' invoice-worker` gives a host PID whose parent is a shim, in the daemon's own cgroup tree, not in the unit's. That single fact explains every symptom of the pattern in the question. ### Symptom 1: the unit is "dead" while the container runs With `ExecStart=/usr/bin/docker run -d …` under the default `Type=simple`, systemd treats the client as the service's main process. The client exits within milliseconds, exit code 0, and the unit transitions to inactive while the invoice-rendering worker happily keeps serving on port 8127. `systemctl status` now lies to everyone who reads it, and to every alert built on it. Add `Restart=always` and it gets worse rather than better: systemd sees the "service" exit and starts it again, which runs `docker run -d --name invoice-worker` a second time. That fails with `Conflict. The container name "/invoice-worker" is already in use by container …`, or, if the unit does not pin a name, quietly starts a *second* worker that fights the first over the published port or double-consumes the queue. Switching to `Type=forking` does not rescue it either: the container is not a fork of the client, so there is no child for systemd to adopt. ### Symptom 2: two supervisors, one container `--restart=always` tells the daemon it is responsible for keeping this container up, including after the daemon itself restarts. A systemd unit with a `Restart=` line says the same thing about the same container. Now: * You `systemctl stop invoice-worker` to take the service down for maintenance. If the unit has no `ExecStop`, systemd only kills the client — which, with `-d`, exited long ago — and the container is untouched. The service you "stopped" is still serving. * If the unit does stop the container, the engine still holds an `always` policy for it, so a daemon restart on that host brings the container back while systemd believes the unit is inactive. Nothing on the box agrees about the desired state. * After a reboot, both supervisors try to start it. Whichever loses reports an error and, under `Restart=always`, retries in a loop. This is the actual interview point: **restart supervision is not additive.** Two watchdogs over one object do not make it more available; they make its state undefined and produce failures that only appear at boot or during maintenance, which is exactly when you cannot afford them. ### Choosing one owner **Option A — the engine owns it.** Run with `--restart=always` (or `unless-stopped`) and write no unit at all. The daemon restarts the container on failure and on its own startup, so as long as the engine's own service is enabled the container comes back after a reboot. This is the least machinery and it is a perfectly respectable answer for one host. What you give up is ordering and dependency expression at the host level, and the container's lifecycle is invisible to the host's service tooling. **Option B — systemd owns it.** Then the unit must actually track the container: ```ini [Service] ExecStartPre=-/usr/bin/docker pull invoice-render:2026.08.19-7c41ab2 ExecStartPre=-/usr/bin/docker rm -f invoice-worker ExecStart=/usr/bin/docker run --rm --name invoice-worker \ -p 127.0.0.1:8127:8000 invoice-render:2026.08.19-7c41ab2 ExecStop=/usr/bin/docker stop -t 30 invoice-worker ``` Three things make this work. The run is **in the foreground**, so the client lives as long as the container and the unit's state tracks something real. There is **no `--restart` flag**, so the engine does not supervise — note that Docker refuses `--rm` together with `--restart` anyway, which is the CLI telling you the two models do not mix. And there is an explicit **`ExecStop` that asks the engine to stop the container**, rather than relying on a signal to the client reaching the container's PID 1. The pre-start `docker rm -f` clears a container left behind by an unclean shutdown, and the leading `-` keeps a first-ever start from failing when there is nothing to remove. The unit also has to be ordered after the Docker engine's own service, or it will race the daemon at boot. One side effect worth knowing: with the foreground form, the container's stdout is relayed by the client and captured by the host's journal *as well as* by the engine's log driver, so the same lines are stored twice. ### How to answer this in an interview Name the layering — client, daemon, shim, container — say that systemd supervises only the client, then state the rule: one supervisor per container, chosen deliberately, with the unit either absent or written to genuinely own the container's lifecycle.
- If systemd is going to own the container, what must the unit do that a bare `docker run` line does not?It has to run the container in the foreground so the unit's main process lives as long as the container, stop the container explicitly on shutdown rather than assuming a signal to the client suffices, clear a stale container of the same name before starting, and be ordered after the Docker engine's own service so it does not race the daemon at boot.
- Which host process is the container's parent, and why does that matter here?A `containerd-shim-runc-v2` process started by the daemon on containerd's behalf; the shim is deliberately detached from whatever client asked for the container so containers survive client and daemon churn. It matters because it is why systemd's supervision of the CLI tells you nothing about the container, and why the unit's cgroup does not contain the workload.
- When is `--restart=always` with no systemd unit the better answer on a single host?When the only requirement is "come back after a crash or a reboot". The engine already does that, its own service is enabled at boot, and a unit adds a second source of truth for no gain. Reach for a unit when you need host-level ordering against other services, a pre-start pull, or the container's state to be visible to the host's service tooling.
Two nurses each independently deciding when to resuscitate the same patient, while the chart on the door is filled in by someone who only watched the messenger walk out of the room.
saying these in an interview costs you the question
- Thinks stopping the unit always stops the container
- Treats `Restart=` and `--restart` as the same knob
- Says setting both is safer, belt and braces
- Believes the docker CLI is the container's parent process
- Suggests `Type=forking` to track a detached container
- Writes a unit with no `ExecStop` for the container