On Linux, how do capabilities split up root's privilege, and which capability lets a process that is not running as root bind a socket to TCP port 80?
answer
- root's power, sliced into pieces
- one named privilege per kernel operation
- ports under 1024 are the reserved range
- the capability name mentions binding a service
- setcap ... =+ep on the binary
basics
~20 sLinux capabilities break root's all-or-nothing power into more than forty independent privileges that can be granted to a process or an executable one at a time. CAP_NET_BIND_SERVICE is the one that allows binding to ports below 1024.
solid answer
~50 sClassic Unix has one privilege bit: UID 0 skips almost every kernel permission check and every other UID skips none. Linux capabilities slice that single bit into separate privileges — over forty of them — so a process can hold exactly the one kernel operation it needs and nothing else. Binding an Internet-domain socket to a port below 1024 is gated by `CAP_NET_BIND_SERVICE`, so a web server can listen on 80 or 443 as an ordinary user account. You attach it to the executable with `setcap cap_net_bind_service=+ep /path/to/binary`, or a supervising process can pass it down through the ambient set. It grants nothing else: raw sockets still need `CAP_NET_RAW`, and changing addresses, routes or firewall rules still needs `CAP_NET_ADMIN`. On modern kernels you can also lower the privileged-port boundary itself with the `net.ipv4.ip_unprivileged_port_start` sysctl instead of granting anything.
code
bash · 4 linessudo setcap cap_net_bind_service=+ep /usr/local/bin/api
getcap /usr/local/bin/api
sudo -u nobody /usr/local/bin/api &
grep ^Cap /proc/$(pgrep -n -f /usr/local/bin/api)/statusgo deeper
Know that ports below 1024 are reserved by the kernel and that CAP_NET_BIND_SERVICE is what lets a non-root process bind them. Be able to say that capabilities are root's power split into separate pieces.
Explain the two ways a process obtains the capability — an attribute on the executable set by setcap, or inheritance from a privileged parent — and name the neighbouring network capabilities it does not include.
Show you can verify the claim on a live host by reading the capability masks of the running process, and weigh the per-binary grant against lowering net.ipv4.ip_unprivileged_port_start or fronting the service with a proxy.
Own the policy question: which capabilities you are willing to hand out at all, who is allowed to run setcap on the fleet, and whether a host-wide sysctl change is an acceptable trade for removing a per-binary privilege.
## The problem capabilities solve In traditional Unix, privilege is a single boolean. A process whose effective UID is 0 bypasses essentially every permission check the kernel makes; a process with any other UID bypasses none of them. That forces a bad bargain: a program that needs one privileged operation — listening on port 80, changing the system clock, opening a raw socket — has to run as root, and therefore also gets the ability to read `/etc/shadow`, load kernel modules, kill any process and rewrite any file. If it is compromised, the attacker inherits all of it. Linux capabilities (introduced in 2.2 and reworked repeatedly since) break that boolean into independent, individually grantable privileges. Internally, each privileged code path in the kernel calls a check for one specific capability rather than asking "is this UID 0?". Current kernels define more than forty of them, each named `CAP_*`. ## What CAP_NET_BIND_SERVICE actually grants The kernel reserves ports below 1024 — the "privileged" or "well-known" ports — so that an arbitrary local user cannot start a fake SSH or HTTP service on a port clients trust. `CAP_NET_BIND_SERVICE` is the capability that lifts exactly that restriction: a thread holding it may bind an Internet-domain socket to a port in that range. It is worth being precise about what it does *not* grant. It does not allow opening raw or packet sockets (that is `CAP_NET_RAW`, what `ping` and `tcpdump` need). It does not allow configuring interfaces, routes, or netfilter rules (that is `CAP_NET_ADMIN`). It has nothing to do with file permissions. This narrowness is the whole point of the model. ## Where a process's capability comes from There are two practical sources. The first is the executable file. `setcap` stores a capability set on the binary, and the kernel grants it at `execve()` time regardless of who runs the program: ``` sudo setcap cap_net_bind_service=+ep /usr/local/bin/api getcap /usr/local/bin/api # /usr/local/bin/api cap_net_bind_service=ep ``` The `p` means the capability lands in the process's permitted set and `e` means it is also made effective immediately, which is what a program that is not capability-aware needs. The second is inheritance from the parent. A service manager or supervisor that itself holds the capability can arrange for the child to keep it across `exec` through the ambient set, which is how service managers grant a low port without touching the binary on disk. ## Reading what a running process holds Every thread's sets are exposed as hex bitmasks in `/proc/<pid>/status`: ``` grep ^Cap /proc/self/status # CapInh: 0000000000000000 # CapPrm: 0000000000000000 # CapEff: 0000000000000000 # CapBnd: 000001ffffffffff # CapAmb: 0000000000000000 ``` `capsh --decode=0000000000000400` turns a mask into names — bit 10 is `cap_net_bind_service`. `getpcaps <pid>` prints the same information in readable form. ## The alternative: move the boundary instead The 1024 cutoff is a kernel policy, not a law. Since Linux 4.11 the sysctl `net.ipv4.ip_unprivileged_port_start` controls where the privileged range ends; it defaults to 1024, and setting it to 80 lets *any* local process bind 80 and above with no capability at all. That is convenient and sometimes the right answer for a single-purpose host, but it is a system-wide loosening: every local user gains the ability to squat on those ports, so it trades a per-binary grant for a host-wide one. ## Why "non-root plus a capability" is not automatically safe Capabilities are only least privilege if the capability you pick is actually small. `CAP_NET_BIND_SERVICE` genuinely is. Others — `CAP_SYS_ADMIN`, `CAP_SYS_MODULE`, `CAP_SETUID`, `CAP_DAC_OVERRIDE` — let a holder recover full root by a short path, so "we run as UID 1000 with a capability" tells you nothing on its own until you know which capability. And nothing about holding a capability restricts the ordinary things the process could already do: it still reads every world-readable file, opens outbound connections and executes other programs under the normal permission rules.
- Apart from granting the capability, what other ways are there to get a service listening on port 80 as an unprivileged user?Lower the kernel's boundary with the `net.ipv4.ip_unprivileged_port_start` sysctl, which is host-wide and lets any local user bind those ports. Or have something privileged own the port and hand the traffic over: a reverse proxy in front of a high-port backend, or a supervisor that opens the listening socket itself and passes the file descriptor to the unprivileged process at start-up.
- Does CAP_NET_BIND_SERVICE let the process do anything else on the network?No. It gates exactly one check — binding an Internet-domain socket to a port below 1024. Raw and packet sockets require `CAP_NET_RAW`; configuring addresses, routes, interfaces or netfilter rules requires `CAP_NET_ADMIN`. That narrowness is why it is one of the few capabilities that is genuinely safe to hand out.
- An older service binds port 80 as root and then drops to an unprivileged user itself. Is a capability still needed?No. The capability check happens once, at `bind()`. A process that starts privileged, binds, and then drops its UID keeps the open listening socket afterwards. The downside is that the code path before the drop still runs with full root, and getting the drop wrong — forgetting supplementary groups, or a `setuid` call whose return value is ignored — is a classic source of privilege-escalation bugs.
Root is a master key that opens every door in the building. Capabilities are the key ring where each door has its own key, so the night cleaner can be handed only the one for the lobby.
saying these in an interview costs you the question
- Capabilities are just sudo rules under a different name
- A non-root process is safe whatever capability it holds
- Ports below 1024 are blocked by the firewall
- CAP_NET_BIND_SERVICE also allows sniffing traffic
- Making the binary executable is enough to bind port 80