skip to content

Why is giving a container access to /var/run/docker.sock considered equivalent to giving it root on the host, and how would an attacker turn that access into a full host takeover?

level: seniorimportance: must knowfreq 55%

answer

  1. daemon is root, API grants root things
  2. run -v /:/host --privileged then chroot /host
  3. no kernel exploit needed, API by design
  4. in-container hardening is bypassed
  5. persist: ssh key, cron, systemd, SUID

basics

~20 s

The daemon runs as root and its API can create a container that bind-mounts the host's / and runs privileged. From inside that container an attacker chroots into the host filesystem or writes to it, gaining full root. The socket is a root-equivalent control plane, not a limited API.

solid answer

~50 s

`dockerd` runs as root, and the Engine API lets a caller create arbitrary containers with arbitrary mounts and privileges. So a process holding the socket does not need a kernel exploit; it just asks the daemon, politely, to do root things. The canonical escape: call `POST /containers/create` for an image that mounts the **host root** (`-v /:/host`) with `--privileged`, start it, then `exec` a shell that `chroot /host`. Now you are root in the host filesystem: write a cron job, add an SSH key, drop a SUID binary, read `/etc/shadow`, or install a systemd unit. Everything the host root can do, you can do. This is why the socket is a **root-equivalent control plane**, and why `:ro` on the mount, a non-root user *inside* the socket container, or dropped capabilities on that container are irrelevant: the privileges are exercised by the daemon on the host, not by your container.

code

bash · 4 lines
bash
# With the CLI reachable inside the compromised container:
docker run -v /:/host --privileged -it alpine chroot /host sh
# Now root on the host filesystem:
echo 'attacker-key' >> /host/root/.ssh/authorized_keys

go deeper

for a junior

Know the headline: socket access equals host root because the daemon is root and will run anything you ask.

for a middle

Walk the -v /:/host --privileged + chroot escape and explain why it needs no exploit.

for a senior

Diagnose it in real systems (CI runners, node sockets), explain why in-container hardening fails, and prescribe proxies/rootless.

for a principal

Treat socket exposure as an architectural boundary violation on multi-tenant hosts; design so no untrusted workload ever holds root-equivalent control.

## The threat model The Docker daemon is a root-privileged process that will, on request via its API, create containers with any configuration: any image, any bind mount, any capability set, `--privileged`, host namespaces, arbitrary devices. There is no per-request authorization beyond 'can you reach the socket'. Therefore **reaching the socket is authorization to do anything root can do**, mediated by a helpful daemon. No kernel CVE, no container-runtime bug is required; this is the intended API being used as designed. ## The canonical escape, step by step 1. **Confirm access.** Hit the API: `curl --unix-socket /var/run/docker.sock http://localhost/v1.44/info`. 2. **Create a container that mounts the host root.** Ask the daemon to run any image with `Binds: ["/:/host"]` and `Privileged: true`. The `/` here is the **host** root, because the daemon resolves paths on the host. 3. **Enter it and chroot.** `chroot /host sh`. You are now operating on the host's real filesystem as root. 4. **Persist.** Add an SSH key to `/host/root/.ssh/authorized_keys`, write a cron entry, install a systemd service, or drop a SUID-root binary. Any of these is a durable host foothold. With the Docker CLI available inside the compromised container it is a one-liner: `docker run -v /:/host --privileged -it alpine chroot /host sh`. Without the CLI, the same result comes from raw API calls, so removing the binary is not a mitigation. ## Why the usual container hardening does not help - **Non-root user in the socket container?** Irrelevant. The *daemon* does the privileged work; your UID inside the container never enters the equation. - **Dropped capabilities / seccomp on the socket container?** They constrain that container's own syscalls, not the sibling the daemon spawns with `--privileged`. - **`:ro` on the socket mount?** Protects only the file, not the API verbs. - **User namespaces on the socket container?** The spawned sibling is created by the daemon under its own configuration. The throughline: every defence that operates *inside* the mounting container is bypassed because the attacker is not escaping the container, they are commanding the host's root daemon to build a new, fully-privileged one. ## Blast radius On a single-tenant build box this 'only' means a compromised job owns the box. On a shared host (multi-tenant CI, a Kubernetes node where a pod mounts the node's socket, a shared runner), it means one workload can read other tenants' source, secrets, and images, and pivot to the orchestrator's credentials sitting on that node. That is the real reason socket exposure is treated as a top-tier finding. ## What actually reduces it Don't mount the raw socket into untrusted or internet-adjacent workloads. Where a component genuinely needs a slice of the API, put a **socket proxy** in front that allow-lists only the needed endpoints (e.g. read-only `GET` for a metrics agent). Prefer **rootless Docker** or a rootless builder (BuildKit/buildx, Buildah, Kaniko) so even a compromised daemon is unprivileged. These are expanded in the CI-pattern and socket-proxy questions.

  • If the container mounting the socket runs as a non-root user and drops all capabilities, is it safe?
    No. Those controls limit that container's own process. The privileged sibling is created by the host daemon under whatever config the API call requests, so the attacker simply asks for `--privileged` and a host-root mount. In-container hardening never touches the daemon's behaviour.
  • How is this different from a container escape via a kernel or runtime vulnerability?
    A kernel/runtime escape abuses a bug to break isolation. Socket access needs no bug at all: it is the documented API being used as intended. That makes it more reliable and version-independent, and it is why it is treated as configuration risk rather than a CVE.
  • In Kubernetes, where does this same risk show up?
    When a pod bind-mounts the node's container-runtime socket (or `hostPath` to `/var/run/docker.sock` / the CRI socket), or runs privileged with host mounts. Same outcome scoped to that node, plus access to the kubelet's credentials, so it can pivot cluster-wide.

saying these in an interview costs you the question

  • Claiming a kernel exploit is required (it is not; the API is enough)
  • Thinking a non-root user or dropped caps inside the socket container prevents it
  • Believing removing the docker CLI blocks the attack (raw curl works)
  • Assuming `:ro` mount neutralises the risk
  • Underestimating blast radius on shared/multi-tenant hosts and CI nodes

context