What does the Docker run option `--security-opt no-new-privileges` do at the kernel level, and which attack does it stop?
answer
- prctl PR_SET_NO_NEW_PRIVS, sticky + inherited
- setuid bits and file caps ignored on exec
- blocks app-user → root-in-container escalation
- check: grep NoNewPrivs /proc/1/status
- breaks setcap-based low-port binding
basics
~20 sIt sets the kernel's no_new_privs bit on the container process, which is inherited by all children and cannot be unset. With it, executing a setuid binary or a file with file capabilities grants no extra privilege, so a compromised unprivileged process cannot escalate that way.
solid answer
~50 s`--security-opt no-new-privileges` sets `PR_SET_NO_NEW_PRIVS` (via `prctl`) on the container's init process. Once set, the bit is sticky: it is inherited across `fork` and `exec` and can never be cleared. Its effect is that `execve` will not grant the new program any privilege the caller did not already have — setuid/setgid bits are ignored, and file capabilities are not applied. The attack it blocks: an attacker gets code execution as an unprivileged user inside the container — through a web-app RCE, a vulnerable dependency, a command injection — and then looks for a setuid-root binary in the image (`sudo`, `su`, `mount`, `pkexec`, `newgrp`, an old `ping`) with a known local privilege-escalation bug. With the bit set, running it confers nothing, so the escalation dead-ends at the unprivileged user. It is cheap and safe for almost every image, so it belongs in the default baseline alongside `--cap-drop ALL`. The exceptions are images that legitimately rely on setuid or on file capabilities at runtime.
code
bash · 11 linesdocker run --rm \
--cap-drop ALL \
--security-opt no-new-privileges \
--read-only \
myapp:1.4.2
# verify inside a running container
docker exec ctr grep NoNewPrivs /proc/1/status # NoNewPrivs: 1
# what escalation targets exist in the image?
docker run --rm myapp:1.4.2 find / -xdev -perm /6000 -type f 2>/dev/nullgo deeper
Know it stops a process from gaining privileges by running a setuid program, and that it is a good default flag.
Explain the prctl bit, its inheritance and irreversibility, and the concrete escalation path it closes.
Discuss rollout and compatibility breakage (setcap binaries, sudo entrypoints), pairing it with non-root plus dropped capabilities, and how to verify via /proc/1/status.
Make it part of a mandatory baseline enforced by admission policy, with a short reviewed exception list, and note its role as a precondition for unprivileged seccomp filters.
## The escalation the flag is aimed at Suppose an image runs its application as an unprivileged user — the recommended posture. An attacker exploits the application and gets a shell as that user. What next? The classic move is *local privilege escalation inside the container*: find a program on the filesystem with the setuid bit set and owned by root. Executing such a binary makes the process run with root's uid regardless of who launched it — that is the entire point of setuid, and it is how `passwd` and `sudo` work. Distribution base images ship several such binaries by default: `su`, `sudo` if installed, `mount`/`umount`, `newgrp`, `chsh`, `gpasswd`, historically `ping`. Any one of them with an unpatched local vulnerability, or a misconfigured `sudoers`, turns "code execution as app user" into "root inside the container". Root inside the container is not host root, but it is a large step: it can read every mounted secret, use whatever capabilities the container holds, and probe for a container escape. The same applies to **file capabilities**: a binary can carry `cap_net_raw+ep` in its extended attributes and gain that capability on exec regardless of the caller. ## What the kernel flag does `prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0)` sets a per-thread flag with two properties that make it useful as a container control: 1. **It is inherited.** Children created by `fork` and programs started by `execve` keep it, so setting it once on the container's PID 1 covers everything the container ever runs. 2. **It is irreversible.** There is no way to clear it — an attacker who is already inside cannot turn it off. With the flag set, `execve` promises that the resulting process will have no privileges the caller lacked. Practically: - setuid and setgid bits are ignored — the binary runs with the caller's ids. - file capabilities are not granted. - AppArmor or SELinux transitions that would *raise* privilege are refused. You can observe it inside a container: `grep NoNewPrivs /proc/1/status` reports `1`. ## What it does not do - It does not stop a process that is *already* root in the container from using the capabilities it already holds. - It does not remove setuid binaries; they remain on disk, they simply stop being useful for escalation. - It does not prevent kernel-level container escapes that rely on a capability the container already has. So it is one layer among several. It pairs with running as a non-root user (a sibling concern) and with `--cap-drop ALL`: if the process starts unprivileged and cannot gain privilege by exec, the escalation surface inside the container is close to empty. A defence-in-depth bonus: image hardening that removes setuid bits (`find / -perm /6000 -type f -exec chmod a-s {} +` in the build, or using a distroless base with no such binaries) achieves a similar result from the image side, and doing both is standard. ## Compatibility considerations Setting the bit breaks anything that genuinely relies on privilege gain at exec: - Images whose entrypoint uses `sudo` or `su` to step down (or up) — restructure to start as the right user via `USER` or the runtime's user option. - Binaries granted a file capability at build time, e.g. `setcap cap_net_bind_service=+ep /usr/local/bin/server` so a non-root process can bind port 80. With `no_new_privs`, that capability is no longer applied on exec. The fix is to use an ambient capability, or better, listen on a high port and publish it. - Tools like `ping` in debug images may stop working, which is cosmetic but confusing during troubleshooting. Because failures show up as plain permission errors, the practical rollout is: enable in staging, watch for `EPERM`/`Operation not permitted` during startup, and keep a short, reviewed exception list. ## Adjacent detail worth knowing The same kernel bit is a precondition for unprivileged seccomp: a process without `CAP_SYS_ADMIN` may install a seccomp filter only if `no_new_privs` is set. The reason is the same — the kernel must be sure a filter cannot be sidestepped by exec'ing a setuid binary. That is why container runtimes set the bit whenever they apply a seccomp profile to an unprivileged process, and it neatly illustrates the underlying principle: the guarantee "privilege cannot increase across exec" is what makes several other confinement mechanisms sound. In declarative deployment specs the same control appears as an `allowPrivilegeEscalation: false` setting in the container's security context; the mechanism underneath is identical.
- If the container already runs as root, does `no-new-privileges` still add value?Much less. The bit prevents privilege *gain* at exec; a process that is already root inside the container with the default capability set has little to gain from a setuid binary. Its value comes from combining it with an unprivileged user and a dropped capability set, so that a compromised process has no path upward. On a root container it is still harmless and worth keeping, but it is not the control doing the work.
- An image uses `setcap cap_net_bind_service=+ep` so a non-root process can bind port 80, and it stops working when the flag is enabled. How do you resolve it?File capabilities are not granted across exec once `no_new_privs` is set, so the binary loses the capability. Either set it as an ambient capability in the runtime configuration so it is already held rather than gained, or — usually simpler and better — have the process listen on a port above 1024 and publish it with a port mapping, which removes the need for any capability.
It is a one-way valve on privilege: whatever the process has when the valve closes is the most it can ever have, no matter what program it runs next.
saying these in an interview costs you the question
- Thinking it drops capabilities or replaces `--cap-drop ALL`
- Believing the flag can be turned off later by the process itself
- Assuming it prevents container escape rather than in-container escalation
- Overlooking that it breaks setcap-based low-port binding and sudo-style entrypoints
- Claiming it only matters for containers running as root