skip to content

Why does each running container get its own long-lived shim process sitting between the container manager and the container's init process, and what does that buy you when the container manager or Docker daemon is restarted?

level: middleimportance: must knowfreq 45%

answer

  1. runc exits → someone must stay
  2. shim = parent, keeps exit code + stdio
  3. containerd/dockerd not in the process tree
  4. Docker needs live-restore: true
  5. orphaned shims after an unclean manager crash

basics

~20 s

The shim is the container's real parent. It keeps the stdio/TTY open, waits on the process so the exit code is never lost, and holds the container alive independently of the daemon. Because the daemon is not in the process tree, containerd or dockerd can restart or upgrade while containers keep running, then re-attach through the shims.

solid answer

~50 s

`runc` exits right after exec'ing the entrypoint, so something must remain. That is the **shim** (`containerd-shim-runc-v2`), one per container, started detached from containerd. It exists to: - **Be the parent.** It `wait()`s on the container's init process, so the exact exit status is captured rather than lost to re-parenting to host PID 1. - **Own the I/O.** It holds the pipes or the TTY master, so logs are not truncated and `docker attach`/`exec` survive a manager restart. - **Decouple lifetime.** containerd and dockerd are not in the container's ancestry, so they can be stopped, upgraded and restarted without signalling containers. On restart, containerd reconnects to the existing shims via their sockets and rebuilds its task view; Docker needs `live-restore` enabled for the daemon side to behave the same. Practical payoff: zero-downtime engine upgrades, and the fact that a dead daemon does not mean dead workloads. The costs are a per-container process (a few MB of RSS) and shims outliving a crashed manager as orphans that must be reconciled.

code

bash · 4 lines
bash
docker run -d --name web -p 8080:80 nginx:alpine
ps -ef --forest | grep -B1 -A3 containerd-shim
systemctl restart containerd
curl -sf localhost:8080 >/dev/null && echo "still serving"

go deeper

for a junior

Know that a per-container helper process stays alive and that this is why containers survive a daemon restart.

for a middle

Explain the three concrete jobs — parent/exit status, stdio ownership, lifetime decoupling — and that Docker needs live-restore enabled.

for a senior

Bring in the operational picture: zero-downtime engine upgrades, reconciliation after an unclean crash, orphaned shims, shim version skew and log buffering.

for a principal

Treat it as blast-radius design — the control plane must be restartable independently of the data plane — and connect it to node upgrade strategy across a fleet.

## The problem the shim solves When `runc` starts a container it creates the namespaces, applies cgroups and security settings, then `execve`s the entrypoint — and exits. Something has to remain behind, because a running process needs a parent that will: - collect its **exit status** (only the parent gets it from `wait()`; if a process is re-parented to host PID 1 the status is reaped by init and lost to the container system), - hold the **stdout/stderr pipes or the TTY master**, otherwise the writing end closes and the process gets `SIGPIPE` or its logs vanish, - accept **exec** requests into the container later, - keep the container's **cgroup and namespaces** referenced and clean them up correctly at the end. The obvious candidate is the daemon itself. That is exactly what early Docker did, and it is why restarting Docker used to kill every container on the host. ## The shim as the answer containerd instead starts a small **shim** process per container (`containerd-shim-runc-v2`), double-forked so it is *not* a child of containerd. The shim: 1. invokes the OCI runtime binary (`runc create`, then `runc start`) — or, for other runtime handlers, whatever binary implements that handler; 2. becomes the parent of the container's init process; 3. owns the stdio: pipes into the logging path, or the TTY master for interactive containers; 4. exposes a small ttrpc API over a socket so containerd can start execs, resize TTYs, send signals and receive exit events; 5. reaps the process, reports the exit code and container-exit event, and tears down mounts and the cgroup afterwards. The v2 shim API allows a shim to serve **multiple containers in one sandbox**, which is what pod-level and VM-based runtimes (Kata) use; for plain Docker it is effectively one per container. ## What the decoupling buys **Manager restarts are non-events for workloads.** Since containerd is not in the container's process tree, `systemctl restart containerd` leaves containers running. On start-up containerd rediscovers the shims by their socket addresses recorded in its state directory, reconnects, and re-subscribes to exit events. Anything that exited while containerd was down is reconciled from the shim's recorded exit status. **Docker Engine gets the same property, but it is a setting.** Restarting `dockerd` only leaves containers running if `"live-restore": true` is set in `/etc/docker/daemon.json`. Without it, the daemon stops containers on shutdown. Live-restore carries caveats worth naming: it is not supported together with swarm mode, and a daemon restart that changes bridge/network configuration or a major daemon version upgrade may still require containers to be recreated. Logs continue to be buffered by the shim while the daemon is away. **Zero-downtime engine upgrades** follow directly. You can patch containerd or Docker on a node while application containers keep serving traffic, which matters enormously on Kubernetes nodes where kubelet and the runtime get updated independently of workloads. **Crash isolation.** A daemon panic does not take down customers' processes. The blast radius of a bug in the control plane is the control plane. ## Costs and failure modes - **Per-container process cost.** A few megabytes of RSS and a PID per container; at hundreds of containers per node this is real but small. - **Orphaned shims.** If containerd dies uncleanly or state is wiped, shims can survive with no manager to talk to. They keep containers alive but invisible to `docker ps`/`ctr tasks ls` until reconciliation, and in bad cases must be cleaned up manually after stopping the container. - **Version skew.** The shim binary in use is the one that was running when the container started; upgrading containerd does not upgrade in-flight shims, so old shims linger until their containers are recreated. That is normally fine but is a real consideration when a shim-level bug is being patched. - **Log gap.** While the daemon is restarting, the shim buffers; a very long daemon outage with a very chatty container can hit buffer limits. ## How to see it `ps -ef --forest` on a Docker host shows a flat set of `containerd-shim-runc-v2` processes, each with the container's process tree beneath it, and `containerd`/`dockerd` sitting entirely to one side. Killing `dockerd` and watching `curl` against a published port keep succeeding is the clearest possible demonstration — and a good thing to have actually tried before the interview.

  • Why not just let the container process be re-parented to host PID 1 when runc exits?
    Because the container system would lose the exit status — host init reaps the process and nobody records whether it exited 0 or 137. It would also lose ownership of the stdio pipes and TTY, breaking log collection and attach, and leave nobody responsible for tearing down mounts and the cgroup. The shim exists to hold exactly those responsibilities.
  • You restarted the Docker daemon and all your containers stopped. What was misconfigured?
    `live-restore` was not enabled in /etc/docker/daemon.json, so the daemon stops containers on shutdown rather than leaving the shims running. Setting `"live-restore": true` and restarting the daemon once makes subsequent restarts non-disruptive. Note that it is incompatible with swarm mode and that some network or major-version changes still require recreating containers.

The shim is the babysitter who stays in the house; the parents (daemon) can leave and come back without waking the child, and they still learn exactly how the evening ended.

saying these in an interview costs you the question

  • Thinking the daemon supervises containers directly, so a daemon restart must kill them.
  • Assuming Docker keeps containers running across a daemon restart by default — live-restore must be enabled.
  • Claiming runc stays resident as the supervisor.
  • Saying the shim only exists for interactive TTY containers.
  • Believing a restarted containerd re-creates containers rather than re-attaching to existing shims.

context