skip to content

Capabilities, seccomp & AppArmor

The kernel confinement layers around every container: the bounded capability set docker run grants instead of full root, the default seccomp syscall filter, and AppArmor or SELinux profiles. Used to separate running containers from securing them.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

6

What does the `--privileged` flag actually grant a Docker container, and why is running with it considered equivalent to giving away root on the host?

level: juniorimportance: must knowfreq 60%

answer

  1. all caps + seccomp unconfined + AppArmor off
  2. cgroup device controller wide open
  3. /sys read-write
  4. mount /dev/sda1 → host root
  5. fix: --cap-add / --device / targeted profile

basics

~20 s

--privileged grants all Linux capabilities, disables the default seccomp and AppArmor/SELinux confinement, lifts the cgroup device restrictions so all host devices are usable, and mounts /sys writable. With that, a process can mount the host disk or load kernel modules — effectively host root.

solid answer

~50 s

`--privileged` is not "a bit more permission"; it removes essentially every isolation layer except namespaces: - **All capabilities** granted, including `CAP_SYS_ADMIN` and `CAP_SYS_MODULE`. - **Seccomp set to unconfined**, so blocked syscalls like `mount`, `keyctl`, and `bpf` become available. - **AppArmor/SELinux confinement dropped** to unconfined. - **The cgroup device controller allows all devices**, and `/dev` is populated from the host — so the raw host disk is readable and writable. - **`/sys` is mounted read-write.** Escape is then trivial and well documented: `mount /dev/sda1 /mnt` and edit anything on the host filesystem, or `insmod` a kernel module. Note filesystem and PID namespaces are still applied, so it is not literally "no isolation" — but every mechanism that would *stop* you crossing them is gone. The fix is to grant only the specific thing needed: `--cap-add` for one capability, `--device` for one device, a `--mount` for one path, or a targeted seccomp profile.

code

bash · 13 lines
bash
# DANGEROUS: this is host root
docker run --privileged -it debian bash
#   inside: fdisk -l ; mount /dev/sda1 /mnt ; ls /mnt/root

# scoped alternatives
docker run --cap-drop ALL --cap-add NET_BIND_SERVICE nginx        # bind :80
docker run --device /dev/ttyUSB0 my-serial-app                    # one device
docker run --cap-drop ALL --cap-add NET_RAW --cap-add NET_ADMIN \
           nicolaka/netshoot tcpdump -i eth0                      # packet capture

# audit what is running privileged
docker ps -q | xargs -r docker inspect \
  --format '{{.Name}} privileged={{.HostConfig.Privileged}}'

go deeper

for a junior

Know that --privileged removes the container's protections and is effectively host root; avoid it and ask what specific permission is needed.

for a middle

Enumerate the four things it turns off — capabilities, seccomp, MAC, device cgroup — and give the scoped --cap-add/--device alternatives.

for a senior

Demonstrate a concrete escape path, show how to diagnose the actual denial with dmesg/strace, and describe auditing and admission policy to keep it out of production.

for a principal

Treat privileged workloads as a named risk acceptance with an owner and an exception list, and design the platform so legitimate node agents are the only holders.

## What container isolation is made of A container is not one mechanism. It is a stack: 1. **Namespaces** — separate views of PIDs, mounts, network, users, IPC, UTS. 2. **Capabilities** — root's powers split into ~40 discrete privileges; Docker grants a small default subset. 3. **cgroups** — resource limits, and importantly the *device* controller, which decides which device nodes may be opened. 4. **seccomp** — a syscall filter; Docker applies a default profile blocking a few dozen dangerous syscalls. 5. **MAC (AppArmor or SELinux)** — a mandatory access control profile restricting file, mount and ptrace operations regardless of Unix permissions. 6. **Filesystem shape** — masked and read-only paths under `/proc` and `/sys`. `--privileged` turns off items 2 through 6. Only namespaces remain, and namespaces alone are a *view*, not a wall — with the right capabilities you can reach around them. ## Concretely, what changes **All capabilities.** The default set is fourteen; privileged gives the full bounding set. `CAP_SYS_ADMIN` alone is often called "the new root": it permits `mount`, namespace manipulation, and much more. `CAP_SYS_MODULE` permits loading kernel modules — code running in the kernel, outside every container boundary by construction. `CAP_SYS_RAWIO` permits raw port and memory access. **Seccomp unconfined.** The default profile denies syscalls that have no business in a normal workload — `mount`, `umount2`, `reboot`, `kexec_load`, `keyctl`, `add_key`, `bpf`, `perf_event_open` under some configs. Privileged removes the filter entirely. **MAC unconfined.** The `docker-default` AppArmor profile denies writes to `/proc/sys`, mounting, and ptrace of processes outside the container. SELinux gives each container a distinct MCS category so it cannot touch another container's labelled files. Privileged disables both. **Devices.** The cgroup device controller normally allows a tiny whitelist (`/dev/null`, `/dev/zero`, `/dev/urandom`, ttys). Privileged allows everything, and the host's `/dev` entries are available. So `/dev/sda`, `/dev/nvme0n1`, `/dev/mem` are all reachable. **`/sys` writable.** Kernel tunables and subsystem controls become writable. ## Why that is host root Three well-known escapes, each a couple of commands: - **Mount the host disk.** `mount /dev/sda1 /mnt` then write an SSH key into `/mnt/root/.ssh/authorized_keys` or drop a file into the host's cron directory. The mount namespace does not stop you because you have `CAP_SYS_ADMIN` and the device is permitted. - **Load a kernel module.** With `CAP_SYS_MODULE`, `insmod evil.ko` executes attacker code in ring 0. - **Abuse a kernel handler.** Classic variants write to `/proc/sys/kernel/core_pattern` (which pipes crash handling to a program run by the host kernel) or, on cgroup v1, to a `release_agent` file. Both cause the host to execute a chosen binary as root. Because of this, "privileged container" and "root on the node" should be treated as the same statement in any threat model. If an attacker gets code execution inside a privileged container — through an application vulnerability, a poisoned dependency, a build step — they own the machine and every other container on it. ## What people actually needed instead `--privileged` usually appears because someone hit a permission error and reached for the biggest hammer. The disciplined alternatives: - Need to bind port 80 as non-root? `--cap-add NET_BIND_SERVICE`, or simply listen on a high port and map it. - Need one device (a GPU, a serial port, a USB dongle)? `--device /dev/ttyUSB0`. - Need to run `tcpdump`? `--cap-add NET_ADMIN` and/or `NET_RAW`, nothing else. - Need to `mount` a FUSE filesystem? `--device /dev/fuse` plus `--cap-add SYS_ADMIN` — still bad, but scoped, and better with a custom seccomp profile. - Need to debug with `strace`/`gdb`? `--cap-add SYS_PTRACE` and `--security-opt seccomp=unconfined` on a *throwaway* debug container, never in the deployed spec. The method is always the same: reproduce the failure, identify the exact denied operation (`dmesg` for AppArmor/SELinux denials, `strace` for `EPERM` syscalls), grant that one thing, and re-test. ## Detecting and preventing it `docker inspect --format '{{.HostConfig.Privileged}}' <ctr>` reports it per container. In an orchestrator, policy admission can reject privileged workloads outright, and that should be the default posture with a narrow, reviewed exception list — typically only node-level agents such as storage or CNI plugins. Treat every exception as a standing risk acceptance with an owner, not a checkbox.

  • Is a privileged container still isolated by namespaces?
    Yes, the namespaces are still created — the container has its own PID, mount and network views unless you also pass `--pid=host` or `--network=host`. But namespaces only change what a process *sees*; with all capabilities, no seccomp filter and full device access, the process can mount the host filesystem or load a kernel module and step outside that view. Isolation that you can dismantle from the inside is not a security boundary.
  • Someone says they need `--privileged` to run Docker inside their CI container. What do you propose?
    First check whether they need a Docker *daemon* at all — most CI image builds can use a rootless, daemonless builder such as BuildKit in rootless mode, Buildah, or Kaniko, none of which need privileged. If a real nested daemon is required, isolate it: a dedicated node pool, a short-lived VM or a sandboxed runtime, treated as a build-farm boundary. Sharing the host's Docker socket is not a safer substitute — that is a different route to the same host-root outcome.

Namespaces are the walls of the room; capabilities, seccomp and AppArmor are the locks. --privileged leaves the walls up and removes every lock, then hands you the keys to the building's plant room.

saying these in an interview costs you the question

  • Describing `--privileged` as just "adds all capabilities" and omitting seccomp, MAC and device access
  • Claiming a privileged container is still safely isolated because namespaces remain
  • Believing running as a non-root user inside the container neutralises `--privileged`
  • Using `--privileged` as a first response to any permission error instead of finding the denied operation
  • Treating sharing the Docker socket as the safe alternative to privileged

context

open as a page

Explain Linux capabilities in the context of containers: what does Docker grant by default, and how do `--cap-drop ALL` and `--cap-add` change what a containerised process can do?

level: middleimportance: must knowfreq 55%

basics

~20 s

Capabilities split root's power into ~40 separate privileges. Docker grants a default subset of 14 (CHOWN, SETUID, NET_RAW, NET_BIND_SERVICE and others) rather than all. --cap-drop ALL removes them all; --cap-add NET_BIND_SERVICE then adds back only what the workload needs.

open as a page

What does the Docker run option `--security-opt no-new-privileges` do at the kernel level, and which attack does it stop?

level: middleimportance: should knowfreq 35%

basics

~20 s

It sets the kernel's no_new_privs bit on the container process, which is inherited by all children and cannot be unset. With it, executing a setuid binary or a file with file capabilities grants no extra privilege, so a compromised unprivileged process cannot escalate that way.

open as a page

How does Docker's default seccomp profile work, what kinds of syscalls does it block, and how would you handle an application that fails because of it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

seccomp is a kernel syscall filter. Docker applies a default JSON profile that allows most syscalls and returns EPERM for a few dozen dangerous ones (mount, reboot, kexec_load, keyctl, add_key, and namespace-creating clone flags). If an app breaks, identify the denied syscall and write a narrow custom profile — never use seccomp=unconfined in production.

open as a page

You are setting a container runtime-confinement baseline for an organisation. Which restrictions do you make mandatory by default, how do you roll it out without breaking existing workloads, and how do you handle the workloads that genuinely need more privilege?

level: principalimportance: should knowfreq 30%

basics

~20 s

Baseline: drop all capabilities, no-new-privileges, runtime default seccomp, MAC profile on, read-only root filesystem, no privileged and no host namespaces. Roll out in audit/permissive mode first, fix the failures found, then enforce. Exceptions get a named owner, a scope, and an expiry.

open as a page

What role do AppArmor and SELinux play in confining containers, how do they differ from Linux capabilities and syscall filtering, and how do you tell when one of them is causing a failure?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

They are mandatory access control systems that restrict which objects a process may touch, regardless of Unix permissions. AppArmor is path-based (Docker applies a docker-default profile); SELinux is label-based, giving each container a distinct MCS category. Denials appear in the host's kernel log, not as application errors.

open as a page