skip to content

What does the `--privileged` flag actually grant a Docker container, and why is running with it considered equivalent to giving away root on the host?

level: juniorimportance: must knowfreq 60%

answer

  1. all caps + seccomp unconfined + AppArmor off
  2. cgroup device controller wide open
  3. /sys read-write
  4. mount /dev/sda1 → host root
  5. fix: --cap-add / --device / targeted profile

basics

~20 s

--privileged grants all Linux capabilities, disables the default seccomp and AppArmor/SELinux confinement, lifts the cgroup device restrictions so all host devices are usable, and mounts /sys writable. With that, a process can mount the host disk or load kernel modules — effectively host root.

solid answer

~50 s

`--privileged` is not "a bit more permission"; it removes essentially every isolation layer except namespaces: - **All capabilities** granted, including `CAP_SYS_ADMIN` and `CAP_SYS_MODULE`. - **Seccomp set to unconfined**, so blocked syscalls like `mount`, `keyctl`, and `bpf` become available. - **AppArmor/SELinux confinement dropped** to unconfined. - **The cgroup device controller allows all devices**, and `/dev` is populated from the host — so the raw host disk is readable and writable. - **`/sys` is mounted read-write.** Escape is then trivial and well documented: `mount /dev/sda1 /mnt` and edit anything on the host filesystem, or `insmod` a kernel module. Note filesystem and PID namespaces are still applied, so it is not literally "no isolation" — but every mechanism that would *stop* you crossing them is gone. The fix is to grant only the specific thing needed: `--cap-add` for one capability, `--device` for one device, a `--mount` for one path, or a targeted seccomp profile.

code

bash · 13 lines
bash
# DANGEROUS: this is host root
docker run --privileged -it debian bash
#   inside: fdisk -l ; mount /dev/sda1 /mnt ; ls /mnt/root

# scoped alternatives
docker run --cap-drop ALL --cap-add NET_BIND_SERVICE nginx        # bind :80
docker run --device /dev/ttyUSB0 my-serial-app                    # one device
docker run --cap-drop ALL --cap-add NET_RAW --cap-add NET_ADMIN \
           nicolaka/netshoot tcpdump -i eth0                      # packet capture

# audit what is running privileged
docker ps -q | xargs -r docker inspect \
  --format '{{.Name}} privileged={{.HostConfig.Privileged}}'

go deeper

for a junior

Know that --privileged removes the container's protections and is effectively host root; avoid it and ask what specific permission is needed.

for a middle

Enumerate the four things it turns off — capabilities, seccomp, MAC, device cgroup — and give the scoped --cap-add/--device alternatives.

for a senior

Demonstrate a concrete escape path, show how to diagnose the actual denial with dmesg/strace, and describe auditing and admission policy to keep it out of production.

for a principal

Treat privileged workloads as a named risk acceptance with an owner and an exception list, and design the platform so legitimate node agents are the only holders.

## What container isolation is made of A container is not one mechanism. It is a stack: 1. **Namespaces** — separate views of PIDs, mounts, network, users, IPC, UTS. 2. **Capabilities** — root's powers split into ~40 discrete privileges; Docker grants a small default subset. 3. **cgroups** — resource limits, and importantly the *device* controller, which decides which device nodes may be opened. 4. **seccomp** — a syscall filter; Docker applies a default profile blocking a few dozen dangerous syscalls. 5. **MAC (AppArmor or SELinux)** — a mandatory access control profile restricting file, mount and ptrace operations regardless of Unix permissions. 6. **Filesystem shape** — masked and read-only paths under `/proc` and `/sys`. `--privileged` turns off items 2 through 6. Only namespaces remain, and namespaces alone are a *view*, not a wall — with the right capabilities you can reach around them. ## Concretely, what changes **All capabilities.** The default set is fourteen; privileged gives the full bounding set. `CAP_SYS_ADMIN` alone is often called "the new root": it permits `mount`, namespace manipulation, and much more. `CAP_SYS_MODULE` permits loading kernel modules — code running in the kernel, outside every container boundary by construction. `CAP_SYS_RAWIO` permits raw port and memory access. **Seccomp unconfined.** The default profile denies syscalls that have no business in a normal workload — `mount`, `umount2`, `reboot`, `kexec_load`, `keyctl`, `add_key`, `bpf`, `perf_event_open` under some configs. Privileged removes the filter entirely. **MAC unconfined.** The `docker-default` AppArmor profile denies writes to `/proc/sys`, mounting, and ptrace of processes outside the container. SELinux gives each container a distinct MCS category so it cannot touch another container's labelled files. Privileged disables both. **Devices.** The cgroup device controller normally allows a tiny whitelist (`/dev/null`, `/dev/zero`, `/dev/urandom`, ttys). Privileged allows everything, and the host's `/dev` entries are available. So `/dev/sda`, `/dev/nvme0n1`, `/dev/mem` are all reachable. **`/sys` writable.** Kernel tunables and subsystem controls become writable. ## Why that is host root Three well-known escapes, each a couple of commands: - **Mount the host disk.** `mount /dev/sda1 /mnt` then write an SSH key into `/mnt/root/.ssh/authorized_keys` or drop a file into the host's cron directory. The mount namespace does not stop you because you have `CAP_SYS_ADMIN` and the device is permitted. - **Load a kernel module.** With `CAP_SYS_MODULE`, `insmod evil.ko` executes attacker code in ring 0. - **Abuse a kernel handler.** Classic variants write to `/proc/sys/kernel/core_pattern` (which pipes crash handling to a program run by the host kernel) or, on cgroup v1, to a `release_agent` file. Both cause the host to execute a chosen binary as root. Because of this, "privileged container" and "root on the node" should be treated as the same statement in any threat model. If an attacker gets code execution inside a privileged container — through an application vulnerability, a poisoned dependency, a build step — they own the machine and every other container on it. ## What people actually needed instead `--privileged` usually appears because someone hit a permission error and reached for the biggest hammer. The disciplined alternatives: - Need to bind port 80 as non-root? `--cap-add NET_BIND_SERVICE`, or simply listen on a high port and map it. - Need one device (a GPU, a serial port, a USB dongle)? `--device /dev/ttyUSB0`. - Need to run `tcpdump`? `--cap-add NET_ADMIN` and/or `NET_RAW`, nothing else. - Need to `mount` a FUSE filesystem? `--device /dev/fuse` plus `--cap-add SYS_ADMIN` — still bad, but scoped, and better with a custom seccomp profile. - Need to debug with `strace`/`gdb`? `--cap-add SYS_PTRACE` and `--security-opt seccomp=unconfined` on a *throwaway* debug container, never in the deployed spec. The method is always the same: reproduce the failure, identify the exact denied operation (`dmesg` for AppArmor/SELinux denials, `strace` for `EPERM` syscalls), grant that one thing, and re-test. ## Detecting and preventing it `docker inspect --format '{{.HostConfig.Privileged}}' <ctr>` reports it per container. In an orchestrator, policy admission can reject privileged workloads outright, and that should be the default posture with a narrow, reviewed exception list — typically only node-level agents such as storage or CNI plugins. Treat every exception as a standing risk acceptance with an owner, not a checkbox.

  • Is a privileged container still isolated by namespaces?
    Yes, the namespaces are still created — the container has its own PID, mount and network views unless you also pass `--pid=host` or `--network=host`. But namespaces only change what a process *sees*; with all capabilities, no seccomp filter and full device access, the process can mount the host filesystem or load a kernel module and step outside that view. Isolation that you can dismantle from the inside is not a security boundary.
  • Someone says they need `--privileged` to run Docker inside their CI container. What do you propose?
    First check whether they need a Docker *daemon* at all — most CI image builds can use a rootless, daemonless builder such as BuildKit in rootless mode, Buildah, or Kaniko, none of which need privileged. If a real nested daemon is required, isolate it: a dedicated node pool, a short-lived VM or a sandboxed runtime, treated as a build-farm boundary. Sharing the host's Docker socket is not a safer substitute — that is a different route to the same host-root outcome.

Namespaces are the walls of the room; capabilities, seccomp and AppArmor are the locks. --privileged leaves the walls up and removes every lock, then hands you the keys to the building's plant room.

saying these in an interview costs you the question

  • Describing `--privileged` as just "adds all capabilities" and omitting seccomp, MAC and device access
  • Claiming a privileged container is still safely isolated because namespaces remain
  • Believing running as a non-root user inside the container neutralises `--privileged`
  • Using `--privileged` as a first response to any permission error instead of finding the denied operation
  • Treating sharing the Docker socket as the safe alternative to privileged

context