skip to content

Privileged Containers & Escape Paths

What --privileged really hands over - every capability, host devices, an unmasked /proc, no seccomp or AppArmor - and the other ways out: host PID or network namespaces, bind-mounting the host root, and runtime CVEs. Breakout reasoning is the screen.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

5

Why does bind-mounting the host root (-v /:/host) into a Docker container hand over the whole host?

level: middleimportance: must knowfreq 55%

answer

  1. A mount hands over the resource
  2. container UID 0 = host UID 0
  3. authorized_keys, cron.d, sudoers
  4. read-only still leaks /etc/shadow
  5. --device on raw disk bypasses permissions

basics

~20 s

Bind-mounting the host root gives the container's root user read-write access to every host file. It can append an SSH key, edit /etc/passwd or sudoers, or drop a cron/systemd unit -- arbitrary code execution as root on the host, with no kernel exploit and no --privileged needed.

solid answer

~40 s

A bind mount is not an exploit; it hands the container the host's files as a shared resource. Without user-namespace remapping, **container UID 0 is host UID 0**, so `-v /:/host` lets the container write anywhere root can: append to `/host/root/.ssh/authorized_keys`, drop a file in `/host/etc/cron.d`, add a systemd unit, or edit `/host/etc/sudoers` -- any of which is root code execution on the host. Even `-v /:/host:ro` is not safe: read access alone leaks `/etc/shadow`, TLS private keys, cloud credential files and kubeconfigs. Passing a raw block device with `--device /dev/sda` is worse still -- the container reads and writes the host filesystem's bytes directly, bypassing file permissions and even the mount's read-only flag. The rule: a container that can see the host filesystem or its disks is a container that owns the host.

code

bash · 4 lines
bash
# host filesystem is now at /host; container root == host root
docker run --rm -v /:/host alpine sh -c \
  'mkdir -p /host/root/.ssh && \
   echo "ssh-ed25519 AAAA... attacker" >> /host/root/.ssh/authorized_keys'

go deeper

for a junior

Recall that -v /:/host mounts the whole host filesystem into the container and that the container's root can then read and change host files -- it is a full-compromise footgun, not a convenience.

for a middle

Explain why container UID 0 equals host UID 0 without userns-remap, and name concrete write vectors (authorized_keys, cron.d, sudoers) and why read-only still leaks /etc/shadow and keys.

for a senior

Diagnose a run spec or Compose file that mounts too much, distinguish a raw --device disk grant from a bind mount, and prescribe mounting only the needed subdirectory or a named volume.

for a principal

Own the guardrail: how you prevent host-root and raw-device mounts fleet-wide, and how you reason about the residual trust when a workload legitimately needs a specific host path.

### It is a handover, not a bug It is tempting to think of a container breakout as always exploiting some flaw. Bind-mounting the host filesystem is not that -- it is simply *giving the container the resource*. `-v /:/host` tells Docker to mount the host's entire filesystem tree at `/host` inside the container. There is nothing left to exploit; the host's files are now the container's files. ### Why container root is host root here By default Docker does not enable user-namespace remapping, so the UID a process runs as inside the container is the **same UID number** the host kernel sees. A process running as UID 0 in the container performs file operations as UID 0 on the host. When that process can reach host files through a bind mount, ownership and permission checks pass exactly as they would for real root on the host. (User-namespace remapping changes this calculus, but that is a separate control.) ### Write vectors -- instant root code execution Given write access to the host tree, an attacker has many one-liners to root on the host: - Append an attacker key to `/host/root/.ssh/authorized_keys` and SSH in. - Drop a job in `/host/etc/cron.d/` or a `systemd` service/timer that runs as root. - Add an entry to `/host/etc/passwd` or a NOPASSWD rule to `/host/etc/sudoers`. - Replace a host binary that root runs, or edit `/host/etc/ld.so.preload`. Consider a **chat-message fan-out service** container that a team started with `-v /:/host` 'so it could read host logs'. A single code-execution bug in that service is now a full host takeover: the attacker writes `/host/etc/cron.d/x` and waits one minute. ### Read-only is still a breach People reach for `-v /:/host:ro` thinking read-only makes it safe. It does not. Read access to the host root exposes every secret on the machine: `/etc/shadow` (crackable password hashes), server TLS private keys, `~/.aws/credentials` and other cloud tokens, SSH host and user keys, kubeconfigs, and application `.env` files. Any one of those is typically enough to pivot to root elsewhere. The `:ro` flag also does not constrain a raw device. ### --device and raw disks `--device /dev/sda` (or any raw block/partition device) hands the container the **block device itself**. Now the container can read or write the host filesystem's raw bytes with tools that ignore the mounted filesystem's permissions -- scraping data a running FS would deny, or writing directly to modify host files. Because it operates below the filesystem layer, a read-only *bind* flag is irrelevant; the device is the disk. Even a seemingly narrow device grant (a GPU, a serial port) deserves scrutiny, but a raw disk or partition is categorically a host compromise. ### The takeaway Neither `-v /:/host` nor `--device` on a host disk requires `--privileged`; they are their own escape surface. Mount only the specific directory a workload needs, prefer a named volume, never the host root, and never a raw disk device into an untrusted container. When you see either in a run spec or Compose file, treat the container as root-equivalent to the host.

  • Is mounting the host root read-only (-v /:/host:ro) safe, then?
    No. Read access alone exposes /etc/shadow, TLS and SSH private keys, cloud-credential files and kubeconfigs -- enough to pivot to root somewhere even if you cannot write. The :ro flag also does not constrain a raw device passed with --device. Read-only reduces the write vectors but still counts as exposing every secret on the host.
  • How does --device /dev/sda differ from a bind mount as an escape?
    It hands the container the raw block device rather than a mounted path. The container reads and writes the host filesystem's bytes directly, bypassing file permissions and even a read-only mount flag -- it can scrape or corrupt data the running filesystem would refuse. Any raw disk or partition device into an untrusted container is a host compromise.

Bind-mounting the host root is not picking the lock -- it is handing over the master key to the whole building; asking for it read-only just means you may copy every document inside rather than also rewrite them.

saying these in an interview costs you the question

  • Thinks a read-only host bind mount is safe
  • Believes the container's root is not the host's root by default
  • Assumes you need --privileged to abuse a host bind mount
  • Thinks mounting only /var/log is low-risk without checking symlinks
  • Confuses a named volume with a bind mount of the host root

context

open as a page

What do Docker's --pid=host, --net=host and --ipc=host flags each expose to a container, and why are they escape risks?

level: middleimportance: should knowfreq 50%

basics

~20 s

Each drops one isolation boundary: --pid=host reveals and lets the container signal every host process, --net=host puts it directly on the host's network stack, and --ipc=host shares host shared memory. All three erode container-host isolation and need no --privileged to be dangerous.

open as a page

How can a privileged Docker container with CAP_SYS_ADMIN break out to the host using the cgroup v1 release_agent?

level: seniorimportance: should knowfreq 40%

basics

~20 s

With CAP_SYS_ADMIN (which --privileged grants) on a cgroup v1 host, the container mounts a cgroup hierarchy, sets its release_agent to a script on a host-visible path, and empties a child cgroup. The kernel then runs that script as root in the host namespace -- a full breakout, needing no CVE.

open as a page

A GPU inference service genuinely needs host device access; how do you contain the blast radius rather than reaching for --privileged?

level: principalimportance: should knowfreq 38%

basics

~20 s

Grant the minimum, not the superset: expose only the specific device and the few capabilities the workload needs, never blanket --privileged. Then assume breakout is possible and shrink what a breakout reaches -- a dedicated low-trust host or node pool, no other tenants or secrets, a minimal cloud identity, network segmentation, and a patched runtime.

open as a page

What kind of vulnerability were the runc escapes CVE-2019-5736 and CVE-2024-21626, and what did they let an attacker do?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Both are escapes in runc, the low-level binary Docker uses to start containers. CVE-2019-5736 let a malicious container overwrite the host runc binary via /proc/self/exe and gain root on the next exec. CVE-2024-21626 (Leaky Vessels) used a leaked host file descriptor plus a crafted WORKDIR to break out to the host filesystem, at build or run time.

open as a page