Hardening guidance often says that granting a Linux process CAP_SYS_ADMIN is close to giving it root outright. Why is that true, and which other capabilities carry the same warning?
answer
- the drawer everything got filed into
- one syscall it gates settles the question
- overwrite a file the login path trusts
- module loading is kernel code execution
- ptrace inherits whatever you attach to
basics
~20 sCAP_SYS_ADMIN is the kernel's catch-all capability, covering mount and dozens of unrelated administrative operations, and mount alone is enough to take over a system. CAP_SYS_MODULE, CAP_SYS_RAWIO, CAP_DAC_OVERRIDE, CAP_DAC_READ_SEARCH, CAP_SYS_PTRACE, CAP_SETUID and CAP_SETPCAP are similarly root-equivalent.
solid answer
~50 sCapabilities were meant to be orthogonal slices of root, but `CAP_SYS_ADMIN` became the drawer where every new privileged operation without an obvious home was filed. It gates `mount()` and `umount()`, `pivot_root()`, `swapon()`, `sethostname()`, `quotactl()`, entering namespaces, and a long tail of device ioctls. `mount()` by itself finishes the argument: a holder can bind-mount its own file over `/etc/shadow` or `/etc/sudoers`, or mount a filesystem image it controls containing a setuid-root binary. There is no meaningful gap between that and root. The same reasoning covers a handful of others: `CAP_SYS_MODULE` loads arbitrary kernel code; `CAP_SYS_RAWIO` reaches physical memory and I/O ports; `CAP_DAC_OVERRIDE` and `CAP_DAC_READ_SEARCH` bypass file permission checks entirely; `CAP_SYS_PTRACE` attaches to a root-owned process and injects code; `CAP_SETUID` simply becomes UID 0. A capability list is only least privilege if none of these are on it.
go deeper
Know that capabilities are not equally sized, and that CAP_SYS_ADMIN is the large one — seeing it on a list means the process is effectively root regardless of its UID.
Name concrete operations CAP_SYS_ADMIN gates, especially mount, and explain the short path from one of them back to full control of the machine.
Review a capability list adversarially: identify the root-equivalent entries, ask which single operation motivated the request, and propose a narrower design such as performing the privileged step once in a parent.
Set the standard for the fleet: which capabilities are never granted without exception review, how requests are justified and recorded, and what compensating controls a workload that genuinely needs a large capability must carry.
## The design intent versus what happened The premise of capabilities is that root's power is decomposable: forty-odd independent privileges, each guarding one class of operation, so a service can hold the one it needs. That works well for the narrow ones — `CAP_NET_BIND_SERVICE` really does only lift the privileged-port restriction, and `CAP_KILL` really does only bypass the signal-permission check. It broke down for `CAP_SYS_ADMIN`. Over decades, when a new privileged operation was added and no existing capability obviously covered it, the check was written against `CAP_SYS_ADMIN`. The result is a capability that guards a large and unrelated collection of operations, and any one of them may be enough on its own. ## What CAP_SYS_ADMIN gates, and why mount ends the discussion Among many others, holding it permits `mount()` and `umount()`, `pivot_root()`, `swapon()`/`swapoff()`, `sethostname()` and `setdomainname()`, `quotactl()`, creating and entering certain namespaces, `keyctl()` operations on other users' keys, and a large set of driver ioctls. Take only `mount()`. A holder can: - bind-mount a file it controls over `/etc/shadow`, `/etc/passwd` or `/etc/sudoers`, replacing the system's own authentication data; - mount a filesystem image it created, containing a root-owned setuid binary, and execute it — unless the mount is forced `nosuid`; - mount its own `/proc` or `/sys` view to confuse other software that trusts those paths. Each of these converts the capability into UID 0. That is why security reviewers treat `CAP_SYS_ADMIN` in a capability list as equivalent to "runs as root", and why guidance to drop capabilities is worthless if this one stays. ## The other root-equivalent capabilities The same test — can a holder reach full control by a short, well-known path? — flags several more: - **CAP_SYS_MODULE** — load and unload kernel modules. Loading a module is arbitrary code execution in kernel context; nothing above it can constrain the result. - **CAP_SYS_RAWIO** — access `/dev/mem`, `/dev/kmem`, I/O ports, and raw block devices. Writing kernel memory directly, or writing the raw block device under a mounted filesystem, bypasses every check above it. - **CAP_DAC_OVERRIDE** — skip file read, write and execute permission checks. Rewrite any file on the system, including `/etc/sudoers` and root's authorized keys. - **CAP_DAC_READ_SEARCH** — skip read and directory-search checks, and use `open_by_handle_at()`. That call opens a file by an opaque handle rather than a path, which historically allowed reaching files outside a restricted root directory. - **CAP_SYS_PTRACE** — attach to any process, including root-owned ones, read its memory and inject code. You inherit whatever that process holds. - **CAP_SETUID** and **CAP_SETGID** — change UID and GID at will. `setuid(0)` is one call away. - **CAP_SETPCAP** and **CAP_SETFCAP** — manipulate capability sets and write file capabilities. The holder can grant itself, or a binary on disk, whatever it lacked. - **CAP_FOWNER** — bypass ownership checks on operations like `chmod` and setting extended attributes. - **CAP_BPF** and **CAP_PERFMON** (split out of `CAP_SYS_ADMIN` in Linux 5.8) — narrower than the original, but still enough to observe and, with BPF, influence a great deal of kernel behaviour. ## Using this in practice The useful habit is to read a capability list adversarially rather than counting entries. "We dropped everything except three capabilities" is a strong claim if the three are `CAP_NET_BIND_SERVICE`, `CAP_CHOWN` and `CAP_KILL`, and a meaningless one if `CAP_SYS_ADMIN` is among them. When someone asks for `CAP_SYS_ADMIN`, the productive next question is which specific operation they need, because it is frequently a single mount or a single ioctl that can be satisfied some other way — performing the mount once at start-up from a privileged parent, or exposing the resource through a file descriptor the service is simply handed. It is also worth remembering what capabilities do *not* do even when the set is genuinely small. They gate specific kernel operations; they do not restrict which files ordinary permissions already allow the process to read, which hosts it may connect to, or which syscalls it may attempt. Those need separate mechanisms — mandatory access control, syscall filtering, and namespaces — and a small capability set is a complement to them, not a substitute.
- A team says they need CAP_SYS_ADMIN because their service mounts a filesystem at start-up. What do you propose instead?Move the mount out of the service. Have a privileged step perform it once before the service starts and hand the process a directory it can simply use, so the running service never holds the capability. If the mount must happen repeatedly, put it behind a tiny privileged helper with a fixed, non-parameterised action rather than granting the general privilege to the whole application.
- Why is CAP_DAC_READ_SEARCH considered more dangerous than plain read access to a lot of files?Because it also permits `open_by_handle_at()`, which opens a file from an opaque handle instead of a path. Path-based confinement assumes an attacker must traverse directories you control; a handle skips that traversal, which is how this capability has historically been used to reach files outside a restricted root directory. It is a boundary bypass, not just broad read access.
- If a service holds CAP_SETUID, what has the non-root user account actually bought you?Almost nothing in terms of privilege — the process can call `setuid(0)` and become root whenever it likes. What remains is bookkeeping: the process shows a non-zero UID in listings and audit records until it makes that call. Treat CAP_SETUID in a service's set as equivalent to running as root, and ask what the process genuinely needs it for.
saying these in an interview costs you the question
- Any capability is safer than root because it is only one
- CAP_SYS_ADMIN just permits generic system administration tasks
- Dropping most capabilities is enough regardless of which remain
- CAP_DAC_OVERRIDE only affects files the user already owns
- Capabilities also restrict which syscalls a process may make