skip to content

Explain Linux capabilities in the context of containers: what does Docker grant by default, and how do `--cap-drop ALL` and `--cap-add` change what a containerised process can do?

level: middleimportance: must knowfreq 55%

answer

  1. root split into ~40 privileges
  2. Docker default = 14, not all
  3. --cap-drop ALL then --cap-add one
  4. NET_RAW / DAC_OVERRIDE are real risks in the default set
  5. non-root needs file/ambient caps — or just use a high port

basics

~20 s

Capabilities split root's power into ~40 separate privileges. Docker grants a default subset of 14 (CHOWN, SETUID, NET_RAW, NET_BIND_SERVICE and others) rather than all. --cap-drop ALL removes them all; --cap-add NET_BIND_SERVICE then adds back only what the workload needs.

solid answer

~50 s

Traditional Unix has a binary split: uid 0 can do everything, everyone else almost nothing. **Capabilities** break root's authority into discrete units — `CAP_NET_BIND_SERVICE` (bind ports < 1024), `CAP_CHOWN`, `CAP_SETUID`/`CAP_SETGID`, `CAP_NET_RAW` (raw sockets, used by ping and tcpdump), `CAP_SYS_ADMIN` (mount and much more), around forty in total. A container's root is not host root partly *because* Docker grants only a default set of **14** and drops the rest from the bounding set. That still leaves real power: `CAP_NET_RAW` enables ARP/DNS spoofing on a shared bridge network, `CAP_SETUID` helps escalation chains, `CAP_DAC_OVERRIDE` bypasses file permission checks. The hardening pattern is allowlist, not blocklist: `docker run --cap-drop ALL --cap-add NET_BIND_SERVICE nginx` Drop everything, add back only what fails without it. Two nuances: capabilities apply to uid 0 in the container, so a non-root process needs file or ambient capabilities to use one; and many services need *no* capabilities at all if they listen above port 1024.

code

bash · 8 lines
bash
docker run --rm --cap-drop ALL --cap-add NET_BIND_SERVICE nginx

# inspect the effective/bounding sets of PID 1 inside a running container
docker exec ctr sh -c 'grep Cap /proc/1/status'
docker exec ctr capsh --print          # if libcap is installed

# decode a raw mask from /proc/1/status
capsh --decode=00000000a80425fb

go deeper

for a junior

Know that capabilities are pieces of root's power, that Docker grants a limited default set, and that --cap-drop ALL is the safe starting point.

for a middle

Name the important capabilities and the default-set risks, and describe the drop-all-then-add-back workflow and how to verify with /proc/1/status.

for a senior

Explain how to derive the minimum set empirically from EPERM traces, the non-root/ambient/file-capability interaction, and how capabilities compose with seccomp and MAC.

for a principal

Set the platform baseline — drop ALL by default, an explicit exception process for additions, admission policy enforcing it, and a plan for the workloads that genuinely need more.

## Why capabilities exist Classic Unix privilege is binary. A process running as uid 0 passes every kernel permission check; everything else fails most of them. That is why `ping` historically had to be setuid-root: it needs a raw socket, one narrow kernel power, and the only way to get it was to become all-powerful root. Linux capabilities (from 2.2 onwards) decompose root's authority into roughly forty independent privileges. Each kernel permission check that used to ask "is this uid 0?" now asks "does this process hold `CAP_X`?". The ones that matter most in containers: - `CAP_NET_BIND_SERVICE` — bind to TCP/UDP ports below 1024. - `CAP_NET_RAW` — raw and packet sockets; needed by `ping` and `tcpdump`, also enables ARP and DNS spoofing against neighbours on a shared bridge. - `CAP_NET_ADMIN` — configure interfaces, routes, firewall rules. - `CAP_CHOWN`, `CAP_FOWNER`, `CAP_DAC_OVERRIDE` — change ownership, bypass ownership checks, bypass file read/write permission checks entirely. - `CAP_SETUID`, `CAP_SETGID` — change process credentials; used legitimately by servers that drop privileges, and by attackers building escalation chains. - `CAP_SYS_ADMIN` — a grab-bag including `mount`, namespace operations, and many others. Widely described as "the new root". - `CAP_SYS_MODULE` — load kernel modules; game over if present. - `CAP_SYS_PTRACE` — attach a debugger to other processes. - `CAP_MKNOD` — create device nodes. ## Docker's default set When you `docker run` with no security flags, the container's init process runs as uid 0 *inside the container* but with a bounding set of only **14** capabilities: `CHOWN`, `DAC_OVERRIDE`, `FSETID`, `FOWNER`, `MKNOD`, `NET_RAW`, `SETGID`, `SETUID`, `SETFCAP`, `SETPCAP`, `NET_BIND_SERVICE`, `SYS_CHROOT`, `KILL`, `AUDIT_WRITE`. Everything else — `SYS_ADMIN`, `SYS_MODULE`, `NET_ADMIN`, `SYS_PTRACE`, `SYS_TIME` — is absent. That default is the main reason "root in a container" is weaker than root on the host. It is also *not a hardened configuration*: `NET_RAW` lets a compromised container spoof traffic to its neighbours; `DAC_OVERRIDE` lets it read any file in its own filesystem regardless of permissions, which matters when a mounted secret is protected only by file mode; `SETUID` supports privilege-manipulation tricks. ## The allowlist pattern Never enumerate what to remove — enumerate what to keep: `docker run --cap-drop ALL --cap-add NET_BIND_SERVICE nginx` This empties the bounding, permitted, and effective sets, then reinstates one. New capabilities added by future kernels are excluded automatically, and a reviewer can see the container's entire privilege surface on one line. How to find the minimum: start from `--cap-drop ALL`, run the real workload with production-like config, and observe failures. Denials show up as `EPERM` from a syscall — `strace -f -e trace=all` on a debug run, or the application's own error (`bind: permission denied`). Map the failing operation to its capability via `capabilities(7)`, add exactly that one, and repeat. Verify the result inside the container with `capsh --print` or by reading `/proc/1/status` (`CapEff`, `CapBnd`) and decoding with `capsh --decode=`. For a large fraction of services the answer is **zero capabilities**: an HTTP server listening on 8080, reading config from a mounted file, writing to stdout, needs none. ## Interaction with non-root users Capabilities are per-process credential sets, and the usual rules mean the *permitted* set is only populated automatically for uid 0. If the image switches to an unprivileged user, `--cap-add NET_BIND_SERVICE` does not by itself let that user bind port 80 — the capability is in the bounding set but not effective. Options are a file capability on the binary (`setcap cap_net_bind_service=+ep /usr/bin/myserver` at build time), an ambient capability, or the simpler engineering answer: **listen on a high port and publish it** (`-p 80:8080`), which removes the need for any capability. A related gotcha: `--cap-drop ALL` on an image whose entrypoint does `chown -R app:app /data` will fail, because `chown` across owners requires `CAP_CHOWN`. Fix the image (bake ownership at build time) rather than adding the capability back. ## Where capabilities sit among the other controls Capabilities are only one layer. seccomp filters *which syscalls* can be called at all; AppArmor/SELinux restrict *which objects* may be touched; `no-new-privileges` prevents a process gaining privileges through setuid binaries. They overlap deliberately: Docker's default seccomp profile is even capability-aware, permitting certain syscalls only when the matching capability is held. Dropping capabilities is nonetheless the highest-value single change, because it directly shrinks what a compromised process is allowed to ask the kernel for. In a declarative deployment the same control is expressed in the container's security context rather than on a CLI flag, but the semantics are identical — drop `ALL`, add back the specific ones.

  • Which capabilities in Docker's default set would you argue are most worth dropping, and why?
    `CAP_NET_RAW` first: it permits raw and packet sockets, so a compromised container can forge ARP or DNS responses against other containers on the same bridge network. Then `CAP_DAC_OVERRIDE`, which bypasses file permission checks and defeats the protection of a mounted secret whose only defence is its file mode. `CAP_SETUID`/`CAP_SETGID` are also worth removing unless the process genuinely drops privileges at startup.
  • You add `--cap-add NET_BIND_SERVICE` but the non-root process still gets EACCES binding port 80. Why?
    Capabilities are populated into a process's effective set automatically only for uid 0; adding one to the container's bounding set does not make it effective for an unprivileged user. You need a file capability on the binary (`setcap cap_net_bind_service=+ep`) or an ambient capability so it survives the exec. The simpler fix is usually to listen on a port above 1024 and publish it, which needs no capability at all.

Root is a master key; capabilities are the individual keys on the ring. Docker hands over fourteen of them by default; hardening means handing back all of them and taking only the one door you need.

saying these in an interview costs you the question

  • Believing Docker grants all capabilities by default
  • Adding `--cap-add SYS_ADMIN` as a generic fix for permission errors
  • Assuming `--cap-add` alone works for a process running as a non-root user
  • Enumerating capabilities to drop instead of dropping ALL and adding back
  • Thinking dropping capabilities makes seccomp and AppArmor unnecessary

context