You find a container reported in the `dead` state, and `docker rm` refuses to remove it. What does that state actually mean, and how do you investigate and clean it up?
answer
- dead = teardown failed partway, not a normal exit
- Usual cause: busy or hung mount, storage-driver error
- inspect .State.Error, then journalctl -u docker
- grep container id in /proc/mounts; lsof/fuser the merged dir
- rm -f → unmount → restart daemon; recurrence = host problem
basics
~20 sdead means the daemon started tearing the container down and could not finish — usually a stuck mount or storage-driver error. The container cannot be started. Try docker rm -f; if that fails, find and unmount the leftover mounts under /var/lib/docker, then remove again, restarting the daemon if necessary.
solid answer
~60 s`dead` is not part of the normal lifecycle. It means the daemon attempted removal or teardown and the operation failed partway — most often because the container's filesystem could not be unmounted (a busy overlay/aufs mount, a bind mount whose source vanished, an unresponsive network filesystem), or because the daemon was killed mid-removal, or a storage-driver error occurred. The container is unusable: it cannot be started, and a plain `docker rm` often fails too. Investigation order: 1. `docker inspect -f '{{.State.Status}} {{.State.Error}}' <c>` for the recorded error. 2. Daemon logs — `journalctl -u docker` — around the failure timestamp; that is where the mount or driver error appears. 3. `mount | grep <container-id>` or `grep <id> /proc/mounts` to find leftover mounts, plus `lsof`/`fuser` on the merged directory to find what is holding it. Cleanup: `docker rm -f <id>`; if it still fails, unmount the stale mounts by hand and retry; as a last resort restart the daemon, which re-runs cleanup on startup. Recurrences usually point at a real host problem — a dead NFS mount, a full or failing disk, or a storage-driver bug — not at the container.
code
bash · 5 linesdocker ps -a --filter status=dead
docker inspect -f '{{.State.Status}} | {{.State.Error}}' <id>
journalctl -u docker --since '30 min ago' | grep -i <short-id>
grep <id> /proc/mounts
df -h /var/lib/docker && df -i /var/lib/dockergo deeper
It is enough to recognize that dead is abnormal — the container cannot be started — and that docker rm -f is the first thing to try.
Explain that it means a failed teardown, usually a busy mount or storage error, and know to check .State.Error and the daemon logs.
Work the full path: inspect, daemon logs, mount table, lsof/fuser, escalating cleanup — and identify the host-level root cause rather than just clearing the container.
Treat it as a host-health signal: alert on dead containers, audit host agents that touch /var/lib/docker, and decide whether the node should be drained and replaced rather than repaired.
## What `dead` is Docker's normal states form a path: `created → running → exited`, with `paused`, `restarting` and `removing` along the way. `dead` sits outside that path. It is the daemon's way of recording *"I tried to tear this container down and I could not finish, so I am not going to pretend it is usable."* You reach it when a removal or cleanup operation fails partway through. Common causes: - **The container's filesystem cannot be unmounted.** The overlay2 (or older aufs/devicemapper) mount is busy: a process on the host still holds a file descriptor into the merged directory, a nested mount was created inside it, or a bind-mount source is on a hung network filesystem (a dead NFS or FUSE mount will block unmount indefinitely). - **The daemon died mid-removal** — OOM-killed, crashed, or the host was hard-rebooted while `docker rm` was in flight — leaving the container object half-deleted. - **A storage-driver or filesystem error**, including a full disk (`no space left on device` during teardown) or an underlying device I/O error. - **Leftover state after an unclean shutdown**, where the daemon's on-disk view and the kernel's mount table disagree. A dead container cannot be started. `docker start` returns an error; `docker exec` cannot attach; and a plain `docker rm` frequently fails with a device-or-resource-busy error. ## Investigating Start with what the daemon recorded: ```bash docker inspect -f '{{.State.Status}} | {{.State.Error}} | {{.State.ExitCode}}' <id> docker inspect -f '{{json .Mounts}}' <id> docker ps -a --filter status=dead ``` `.State.Error` often names the failure directly. Then read the daemon's own log around that time — this is where the real error lives: ```bash journalctl -u docker --since '30 min ago' | grep -i <short-id> ``` Then look at the host's mount table for anything still referencing the container or its layers: ```bash grep <container-id> /proc/mounts mount | grep -E 'overlay|<container-id>' lsof +D /var/lib/docker/overlay2/<layer-id>/merged 2>/dev/null fuser -vm /var/lib/docker/overlay2/<layer-id>/merged ``` Also check the obvious host-level causes, because they are frequently the real story: `df -h /var/lib/docker` and `df -i` for space and inodes, `dmesg -T | tail` for I/O errors or OOM kills, and `stat` on any bind-mount source that lives on a network filesystem. ## Cleaning up Escalate in order, least invasive first: 1. **`docker rm -f <id>`** — force removal. This clears the majority of cases. 2. **Unmount the leftovers, then retry.** If `rm -f` reports the device is busy, kill or restart whatever process `lsof`/`fuser` identified, then `umount` the stale mount points found in `/proc/mounts` (a lazy `umount -l` may be needed for a hung network mount), and run `docker rm -f` again. 3. **Restart the Docker daemon.** `systemctl restart docker` re-runs container cleanup on startup and resolves cases where the daemon's in-memory state is inconsistent. Note this affects every container on the host, so treat it as a maintenance action, not a reflex. 4. **Last resort: manual state removal.** Deleting `/var/lib/docker/containers/<id>` by hand with the daemon stopped is possible but risks leaving orphaned layers and metadata inconsistency; only do it with the daemon stopped and after the mounts are gone, and prefer a documented recovery over improvisation. ## Do not stop at the symptom A single dead container after an unclean reboot is noise. A host producing them repeatedly is telling you something: - a **hung network mount** used as a bind source, which will keep blocking unmounts, - a **full or failing disk** under `/var/lib/docker` — check space *and* inodes, - **something on the host reaching into container layer directories** (a backup agent, an antivirus scanner, or a security tool walking `/var/lib/docker`), which is a classic cause of busy mounts, - a **storage-driver or kernel bug**, worth checking against the driver in use (`docker info | grep -i 'storage driver'`) and the kernel version. Operationally, treat `dead` as an alertable condition: `docker ps -a --filter status=dead -q | wc -l` is a cheap check, and any non-zero result on a fleet host deserves a look, because the same underlying problem that stranded one container will strand the next one.
- What distinguishes `dead` from `exited` with a non-zero code?`exited` is a normal outcome: the main process ran and terminated, the exit code is recorded, and the container object is fully intact — you can inspect it, read its logs, copy files out and start it again. `dead` means the daemon could not complete teardown, so the container is in an inconsistent state and cannot be started at all. One is about the application; the other is about the daemon and the host's storage.
- Why would `umount` fail on a container's overlay mount, and what do you do about it?Because something still holds it: a host process with an open file descriptor under the merged directory (backup agents and security scanners are frequent culprits), a nested mount created inside it, or a bind-mount source on a hung network filesystem that blocks the unmount indefinitely. Identify the holder with `lsof`/`fuser`, stop it, and retry; for a hung network mount a lazy unmount (`umount -l`) detaches it so cleanup can proceed.
- A host produces dead containers repeatedly. What do you check beyond the containers themselves?Host-level storage health: free space and inodes under `/var/lib/docker`, `dmesg` for I/O errors, and the storage driver and kernel version for known bugs. Also look for host agents walking `/var/lib/docker` and holding mounts open, and for bind mounts backed by network filesystems. Repeated dead containers are a host symptom, not a container symptom.
saying these in an interview costs you the question
- Treating `dead` as a synonym for stopped or crashed
- Jumping straight to deleting files under /var/lib/docker before trying rm -f and unmounting
- Restarting the Docker daemon as the first step on a busy host
- Ignoring a recurring pattern instead of investigating disk, inodes, or host agents holding mounts
- Assuming a dead container can be started again once the host quiets down