On a Linux host running Docker Engine, walk through the chain of components involved in starting a container — from the daemon down to the process that becomes PID 1 inside the container — and say what each layer is responsible for.
answer
- dockerd → containerd → shim → runc → PID 1
- containerd = content store + snapshotter + tasks
- one shim per container, it is the parent
- runc exits after exec; it is not a supervisor
- "OCI runtime create failed" = bottom of the stack
basics
~20 sdockerd takes the API request and handles images, networks and volumes, then asks containerd to run the container. containerd manages the snapshot and metadata and starts a containerd-shim per container. The shim invokes runc, which creates namespaces and cgroups and execs the entrypoint. runc then exits; the shim stays as the container's parent.
solid answer
~50 sFour layers, each with a distinct job: 1. **dockerd** — the API front end and Docker-level concepts: image pull orchestration, networks (bridge, iptables), volumes, restart policies, log drivers. It builds the container spec and calls containerd over gRPC. 2. **containerd** — the core container manager: content store and snapshotter (prepares the rootfs from the image layers), container/task metadata, image pull, and generating the OCI `config.json`. It starts one shim per container. 3. **containerd-shim-runc-v2** — one process per container. It invokes the runtime, then remains as the container's parent: it owns the stdio/TTY, reaps the process and reports the exit status back to containerd. 4. **runc** — the OCI runtime. Given the bundle (rootfs + `config.json`), it creates namespaces, applies cgroup limits, seccomp, capabilities and LSM labels, pivots root, and execs the entrypoint. **Then runc exits.** So the container's process tree is `shim → container PID 1`, not `dockerd → container`. That is what lets the daemon restart without killing anything.
code
bash · 4 linesps -ef --forest | grep -A3 containerd-shim
ctr -n moby containers ls
ctr -n moby tasks ls
systemctl status docker containerdgo deeper
Be able to list the four components in order and say roughly what each does.
Explain each layer's specific responsibility, that runc exits after exec, and that the shim is the container's parent.
Use the chain diagnostically — read error messages to the right layer, inspect with ctr, and explain the design reasons for the split.
Discuss it as a modularity boundary: what reusing containerd bought the ecosystem, how runtime pluggability and CRI follow from it, and what you standardize on across a fleet.
## The chain ``` docker CLI → dockerd → containerd → containerd-shim-runc-v2 → runc → your process (PID 1 in the container) ``` Each arrow is a real process boundary, and each layer exists because it owns something the others should not. ## dockerd — Docker's own concepts `dockerd` serves the Engine API and owns everything that is *Docker's* idea rather than a generic container idea: named networks and the iptables/bridge plumbing behind them, volumes and volume drivers, restart policies, logging drivers, the build front end, `docker compose`-visible metadata, and the CLI-facing model of a "container" with a name and ports. It does not create namespaces. Having assembled what to run, it makes a gRPC call to containerd. ## containerd — the container manager `containerd` is the layer that does the generic work: - **Content store** — the blobs pulled from a registry, addressed by digest. - **Snapshotter** — turns image layers into a usable rootfs, typically an overlayfs mount with a writable upper layer. This is what makes the "rootfs" in the bundle exist. - **Metadata** — containers, their specs, tasks (a running instance), and namespaces (containerd's own multi-tenancy concept, e.g. `moby` for Docker, `k8s.io` for Kubernetes). - **OCI spec generation** — merging image config defaults with the caller's options into `config.json`. - **Task management** — start, exec, kill, and collecting exit status via shims. containerd is deliberately usable on its own; Docker is one client of it. ## The shim — one supervisor per container For each container, containerd starts a **shim** process (`containerd-shim-runc-v2`) and detaches it. The shim then: 1. calls the OCI runtime binary to `create` and `start` the container; 2. becomes the parent of the container's init process, so it can `wait()` on it and know the exact exit code; 3. holds the container's stdio pipes or the TTY master, so logs and `docker attach` keep working; 4. keeps an FD/socket open back to containerd, reconnecting after a containerd restart; 5. reports exit and cleans up when the container ends. The v2 shim API also allows one shim to serve a group of containers (used for pods and for VM-based runtimes like Kata, where a single sandbox hosts several containers). ## runc — the OCI runtime `runc` receives the bundle: a directory containing the prepared `rootfs/` and `config.json`. It then, in order: clones into new namespaces (`mount`, `pid`, `uts`, `ipc`, `net`, optionally `user` and `cgroup`), places the process in the configured cgroup so limits apply from the start, sets up mounts (`/proc`, `/sys`, `/dev`, binds and volumes), masks sensitive kernel paths, applies capabilities, `no_new_privs`, the seccomp filter and AppArmor/SELinux labels, `pivot_root`s into the container rootfs, and finally `execve`s the entrypoint. The crucial detail: for a normal container **runc exits immediately after start**. It is not a supervisor. The container process is re-parented to the shim. `ps -ef --forest` on a Docker host shows shims with container processes underneath and no runc in sight. ## Why it is split this way - **Daemon restarts don't kill containers.** Because the shim, not dockerd, is the parent, upgrading or restarting Docker leaves workloads running (with live-restore configured on the Docker side); the daemon re-attaches through the shims. - **Reuse.** Kubernetes' kubelet talks to containerd directly through the CRI plugin, skipping dockerd entirely — possible only because the container-management logic lives in containerd, not the Docker daemon. - **Runtime pluggability.** Because the bottom layer is just an OCI runtime binary invoked by the shim, `runc` can be replaced by `crun`, `runsc` or a VM-based runtime without touching anything above. - **Memory.** A supervisor that stayed resident per container and carried the full runtime would be expensive; the shim is small, and runc's memory is released when it exits. ## Debugging value Knowing the chain tells you where to look: - `docker` errors mentioning the API → the CLI/dockerd hop. - "OCI runtime create failed: … exec: \"/app\": stat /app: no such file or directory" → runc, at the very bottom: the entrypoint path does not exist in the rootfs. - Containers running but `docker ps` failing → dockerd is down, shims are fine. - `ctr -n moby containers ls` and `ctr -n moby tasks ls` inspect the containerd layer beneath Docker.
- Where in this chain is the container's rootfs actually assembled?In containerd, by its snapshotter. containerd unpacks the image layers from its content store and prepares a snapshot — normally an overlayfs mount stacking the read-only layers with a writable upper directory — and that mount point becomes the `rootfs/` in the bundle. runc receives an already-assembled rootfs and only pivots into it.
- What does an error like "OCI runtime create failed: exec: /app/server: no such file or directory" tell you about where the failure happened?It came from the bottom of the stack: runc successfully set up namespaces and mounts but could not exec the configured entrypoint inside the container's rootfs. Typical causes are a wrong path or a binary that was never copied into the final stage, a script missing its execute bit or a CRLF shebang, or a dynamically linked binary in a distroless/alpine image lacking its loader. It is not a daemon or containerd problem.
dockerd is the front-of-house manager taking the order, containerd is the kitchen that preps ingredients and assigns a station, the shim is the line cook who stays with the dish, and runc is the single decisive action of putting it on the heat and walking away.
saying these in an interview costs you the question
- Saying dockerd creates the namespaces and cgroups itself.
- Believing runc stays running as the container's supervisor.
- Thinking the container process is a child of dockerd, so a daemon restart must kill it.
- Placing image layer unpacking in runc rather than containerd's snapshotter.
- Treating containerd as "part of Docker" that cannot be used independently.