skip to content

dockerd, containerd & runc

The layered runtime: dockerd is the API front end, containerd manages images and lifecycles, a per-container shim survives daemon restarts, and runc sets up namespaces and cgroups then exits. Asked because it explains why Docker itself is not required.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

5

The `docker` command-line tool does not itself create containers. Describe the client/daemon architecture behind it and what actually happens when you run a docker command on a Linux host.

level: juniorimportance: must knowfreq 52%

answer

  1. CLI = HTTP client over /var/run/docker.sock
  2. dockerd = control plane, not executor
  3. DOCKER_HOST / docker context → remote engine
  4. docker group == root on the host
  5. shell is not the container's parent

basics

~20 s

The docker CLI is a thin HTTP client. It sends a request over the Unix socket /var/run/docker.sock to the long-running daemon dockerd, which does all the real work — pulling images, creating containers, managing networks — and streams results back. Containers are children of the daemon's stack, not of your shell.

solid answer

~50 s

Docker Engine is **client/server**. - `docker` (the CLI) builds an HTTP request against the Docker Engine API and sends it, by default over the Unix socket `/var/run/docker.sock` (or a TCP/SSH endpoint via `DOCKER_HOST`). - `dockerd` (the daemon) receives it, and owns image storage, networks, volumes, build coordination and container lifecycle. It delegates the actual container execution downwards to containerd and runc. Consequences that matter: - Your shell is not the container's parent. Closing the terminal does not stop the container; `docker run` just attaches streams over the API. - The CLI can drive a remote engine — same commands, different `DOCKER_HOST`. - Membership of the `docker` group is effectively root on the host, because anyone who can talk to the socket can ask the daemon to start a privileged container that mounts `/`. - If `dockerd` is down, the CLI fails with "cannot connect to the Docker daemon" even though containers may still be running.

code

bash · 6 lines
bash
curl -s --unix-socket /var/run/docker.sock \
  http://localhost/v1.45/containers/json | jq '.[].Names'

export DOCKER_HOST=ssh://ops@build-01
docker ps
docker version

go deeper

for a junior

Say clearly that the CLI is a client that sends API requests over a socket to the dockerd daemon, which does the real work.

for a middle

Add DOCKER_HOST/contexts, API version negotiation, and that dockerd delegates execution further down to containerd and runc.

for a senior

Lead with the security and operational consequences: socket access is root, daemon downtime is separable from container uptime, remote engines in CI.

for a principal

Frame daemon access as a host-level trust boundary — rootless mode, socket proxies, who may reach the API, and what that implies for shared build infrastructure.

## Two programs, not one People say "Docker" as if it were a single binary. On Linux it is at least two: the **client** (`docker`) and the **daemon** (`dockerd`). The client holds essentially no logic about containers. It parses your flags, turns them into an HTTP request against the **Docker Engine API**, sends it, and renders the response. The transport is chosen by `DOCKER_HOST` (or the active `docker context`): - default on Linux: the Unix socket `unix:///var/run/docker.sock` - `ssh://user@host` — tunnel to a remote daemon - `tcp://host:2376` — TCP, which should always be mutual-TLS So `docker ps` is roughly `GET /containers/json`, `docker run` is a `POST /containers/create` followed by `POST /containers/{id}/start` plus a stream attach, and `docker build` uploads the build context and streams back progress. ## What the daemon owns `dockerd` is a long-running privileged process (root, unless you deliberately run rootless mode). It owns: - the local **image store** and pulls/pushes to registries - **container** records, names, labels, restart policies, logging drivers - **networks** (bridge creation, iptables rules, DNS for user-defined networks) and **volumes** - the **build** front end, delegating to BuildKit - the API itself, including its version negotiation with older clients What it does *not* do is create namespaces and cgroups itself. It hands that down to containerd, which invokes a shim and `runc`. The daemon is the control plane, not the executor. ## Consequences you are expected to draw **Your shell is not the parent.** `docker run -it alpine sh` looks like a subprocess, but the container process was created several layers below the daemon; the CLI merely attached stdin/stdout/stderr to the API stream. Detaching (`Ctrl-P Ctrl-Q`) or losing the terminal does not kill the container. Conversely `docker stop` sends `SIGTERM` to the container's PID 1 via the daemon, not via your shell's job control. **Remote control is free.** The same CLI drives a daemon on another machine by switching context. This is also why CI systems can build "locally" against a remote engine. **Socket access equals root.** Anyone who can write to `/var/run/docker.sock` can ask the daemon — which runs as root — to start a container with `--privileged` and `-v /:/host`, and thereby read or modify anything on the host. Adding a user to the `docker` group is therefore a root grant, and mounting the socket into a container hands that container the host. (Rootless mode and socket-proxy filtering are the usual mitigations.) **Failure modes are separable.** If `dockerd` crashes or is being upgraded, the CLI reports `Cannot connect to the Docker daemon at unix:///var/run/docker.sock. Is the docker daemon running?` — but running containers may well still be serving traffic, because they are supervised by per-container shim processes rather than by the daemon itself. **API versioning is real.** The client negotiates an API version with the daemon; a very new client against an old daemon downgrades, and a client asking for features the daemon lacks gets a clear version error. `docker version` prints both sides, which is the first thing to check when a flag "does not exist" on one machine. ## Where Docker Desktop differs On macOS and Windows there is no Linux kernel to host containers, so Docker Desktop runs a Linux VM; `dockerd` lives inside that VM and the socket you talk to is a proxy into it. This is why bind-mount performance and file-permission behaviour differ from Linux, and why "localhost" semantics need care. The client/daemon split is the same — the daemon is simply one virtualization boundary away. ## Quick self-check You can bypass the CLI entirely and prove the model: ``` curl --unix-socket /var/run/docker.sock http://localhost/v1.45/containers/json ``` If that returns JSON, you have just done what `docker ps` does.

  • Why is adding a user to the `docker` group equivalent to giving them root on that host?
    The daemon runs as root and will do whatever the API asks. A member of the group can reach the socket and request a container with `--privileged` and the host root bind-mounted, then read or modify any file, including /etc/shadow or systemd units. There is no privilege check inside the daemon that distinguishes group members from root, so the group is a root grant in practice.
  • You close the terminal that ran `docker run`. Is the container stopped?
    Not by itself. The container's process was created below the daemon and is supervised by its own shim, so closing the terminal only tears down the attached I/O stream. The container keeps running unless it was started with `--rm` and its main process exits, or unless the process was tied to the terminal and receives a hangup through an attached TTY.

saying these in an interview costs you the question

  • Describing `docker` and `dockerd` as one program.
  • Thinking the container is a child process of your shell.
  • Believing docker group membership is a limited, unprivileged convenience.
  • Assuming the CLI can only ever manage a local engine.
  • Concluding that a "cannot connect to the daemon" error means all containers are down.

context

open as a page

On a Linux host running Docker Engine, walk through the chain of components involved in starting a container — from the daemon down to the process that becomes PID 1 inside the container — and say what each layer is responsible for.

level: middleimportance: must knowfreq 50%

basics

~20 s

dockerd takes the API request and handles images, networks and volumes, then asks containerd to run the container. containerd manages the snapshot and metadata and starts a containerd-shim per container. The shim invokes runc, which creates namespaces and cgroups and execs the entrypoint. runc then exits; the shim stays as the container's parent.

open as a page

Why does each running container get its own long-lived shim process sitting between the container manager and the container's init process, and what does that buy you when the container manager or Docker daemon is restarted?

level: middleimportance: must knowfreq 45%

basics

~20 s

The shim is the container's real parent. It keeps the stdio/TTY open, waits on the process so the exit code is never lost, and holds the container alive independently of the daemon. Because the daemon is not in the process tree, containerd or dockerd can restart or upgrade while containers keep running, then re-attach through the shims.

open as a page

containerd can be run on its own without Docker Engine. What is containerd responsible for by itself, and what is the Container Runtime Interface (CRI) that lets an orchestrator's node agent drive it directly?

level: seniorimportance: should knowfreq 42%

basics

~20 s

On its own, containerd handles image pull and storage, snapshots/rootfs preparation, container and task lifecycle via shims, and low-level networking hooks — exposed over a gRPC API on a Unix socket. CRI is a standard gRPC API (image and runtime services) that containerd implements as a plugin, so a node agent can drive it without Docker Engine.

open as a page

The Docker daemon on a Linux host is unresponsive, yet the application containers on that host are still serving traffic. Explain why that is possible and how you would inspect and manage those containers while the daemon is down.

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Containers are parented by per-container shim processes, not by the daemon, so they keep running when it hangs. Inspect them one layer down with ctr -n moby containers ls / tasks ls, or from the OS with ps --forest, nsenter into the namespaces and the cgroup tree. Fix or restart the daemon; with live-restore it re-attaches without stopping anything.

open as a page