skip to content

Your platform standard requires every Linux service to run under a non-root account with a minimal capability set. What does that requirement actually buy you, where does the capability model stop being a real boundary, and how do you decide when a workload needs something stronger?

level: principalimportance: nice to knowfreq 28%

answer

  1. removes the easy escalations
  2. says nothing about files or network
  3. only as small as its largest member
  4. the syscall surface is untouched
  5. layer it, do not lean on it

basics

~20 s

Dropping to a minimal capability set removes specific kernel privileges and shrinks what a compromised process can escalate to. It is not a containment boundary: it does not restrict ordinary file access, syscall surface or network reach, and one root-equivalent capability undoes the whole set.

solid answer

~60 s

A small capability set buys real, measurable things: it closes the common escalation paths, it turns "what is this service allowed to do to the kernel" into a reviewable declaration, and it makes a request for more privilege a visible event. What it does not buy is containment. Capabilities gate specific privileged operations; they say nothing about the files ordinary permissions already let the process read, the hosts it can connect to, the processes it can see, or the syscalls it may attempt. The set is also only as small as its largest member — leaving `CAP_SYS_ADMIN`, `CAP_SETUID` or `CAP_DAC_OVERRIDE` in place makes the other drops decorative. So I treat the capability set as one layer and ask what the workload's blast radius is. Where a compromise would matter, I add the mechanisms that actually confine: namespaces for what it can see, mandatory access control for what it can touch, syscall filtering for what it can call, and a separate account and host boundary for anything handling secrets or untrusted input.

go deeper

for a junior

Understand the two halves: dropping capabilities removes specific kernel privileges, but it does nothing about the files and network the process could already reach as an ordinary user.

for a middle

Be able to state concretely what a minimal set does not cover — file access under normal permissions, egress, syscall surface, process visibility — and why the set is only as strong as its most dangerous member.

for a senior

Argue the design for a real service: which privilege is genuinely required, whether it can be lifted into a privileged parent or a start-up step, and which additional mechanism you would add where a compromise reaches beyond the process.

for a principal

Own the standard and its exception path. Define the fleet baseline, decide which capabilities are removed from the bounding set outright, make privilege a declared and reviewed property of a service, and keep the exception process cheap enough that teams use it honestly.

## What the requirement genuinely delivers Three things, and they are worth having. First, it removes the standard escalation paths. A process without `CAP_SETUID`, `CAP_SYS_MODULE`, `CAP_SYS_PTRACE`, `CAP_DAC_OVERRIDE` and `CAP_SYS_ADMIN` cannot turn a code-execution bug into root by any of the short, well-known routes. That is a real reduction in the value of a remote bug. Second, it makes privilege declarative and reviewable. Once every service states its capability set somewhere a human reads, an unusual entry becomes a question in code review rather than an invisible property of a running host. Third, it creates a ratchet. If the default is an empty set and additions require justification, privilege stops accumulating by accident — which is the normal failure mode when the default is "runs as root and nobody wrote down why". ## Where the model stops being a boundary A capability set is an allow-list of *privileged kernel operations*. Everything a process can already do with ordinary permissions is untouched: - **Filesystem reach.** A non-root service with no capabilities still reads every world-readable file, including configuration files, other services' logs, and any secret carelessly left mode 0644. Capabilities constrain none of that; ownership, mode bits and mount namespaces do. - **Network reach.** It can still open outbound connections anywhere the network allows, exfiltrate data, and reach internal services that trust the source address. Egress policy is a different mechanism entirely. - **Syscall surface.** Capabilities gate whether a privileged call *succeeds*, not whether it can be *attempted*. Kernel vulnerabilities are typically reached through calls that need no capability at all, so a minimal set does little against a local kernel exploit. Syscall filtering is what narrows that surface. - **Visibility of other processes.** Without a PID namespace, the service still sees every process on the host and anything exposed in `/proc`. - **Non-orthogonality.** The model's own premise leaks. Several capabilities let a holder regain the rest, so the set's strength is set by its most dangerous member, not by how many were dropped. A list of one is worse than a list of five if that one is `CAP_SYS_ADMIN`. There is a second, operational failure: privilege granted through a file capability on a binary is invisible in every listing that shows permissions, is silently lost on reinstall, and is silently reintroduced by anyone with root. Privilege that lives on disk is privilege nobody reviews. ## How I decide what a workload needs I grade by blast radius rather than by workload type. **Baseline for everything**: non-root account, empty capability set, capabilities removed from the bounding set so nothing in the process tree can regain them, and no privilege carried on binaries. **One narrow privilege needed**: prefer moving the privileged act out of the service — have a privileged parent perform the mount, or open the listening socket and pass the descriptor down — before granting anything. A grant that exists only during start-up is much cheaper than one held for the process's lifetime. **Handles untrusted input, or a compromise reaches data beyond its own**: capabilities are no longer the interesting control. Add the mechanisms that limit reach — a restricted filesystem view, mandatory access control confining which paths and ports the process may use, syscall filtering, and egress restrictions. Judge the design by what an attacker with full code execution inside the process could touch, not by the length of the capability list. **Genuinely needs a root-equivalent capability**: this is an exception, not a configuration. It should be a named, time-bounded decision with an owner, an explicit statement of what compensating controls exist (isolated host, restricted network position, extra audit), and a standing question of what it would take to remove it. ## How I would roll it out Start in report mode: inventory which services currently run as root and what each one actually needs, because most need nothing. Convert the easy majority to non-root with an empty set. Handle the residue individually, one design conversation each, and expect a few to reveal a start-up-only privilege that can be lifted out of the service entirely. Make the standard's exception process cheap enough to use honestly — an expensive exception process produces quiet non-compliance, not compliance. The honest summary to give an interviewer: a minimal capability set is a cheap, high-value hygiene control that removes the easy escalations and makes privilege visible. It is not isolation, and any argument that treats it as isolation will be wrong the first time it matters.

  • A team argues that because their process is non-root with no capabilities, a remote code-execution bug in it is low severity. How do you respond?
    Ask what the process can already reach with ordinary permissions: which files it reads, which credentials are in its environment or config, which internal services accept its connections, and which data it holds in memory. That set is untouched by the capability drop, and for most services it is the whole point of the compromise. Escalation to root is one outcome among several, not the only one that matters.
  • Where would you enforce the standard so it does not decay over time?
    At the point where services are defined, not on the hosts. The privilege set should be part of the service's declared configuration, reviewed like code, with a default of empty. Host-level checks that report drift are useful as a backstop, but a rule that only exists as a wiki page and a hand-run setcap will decay within a release or two.
  • Which capability would you remove from the bounding set fleet-wide first, and why the bounding set specifically?
    CAP_SYS_MODULE is a good first candidate: essentially no application legitimately loads kernel modules, and it is unrecoverable kernel code execution. The bounding set is the right place because it is inherited by descendants and cannot be raised again, so it survives an exec of a binary that carries the capability — a permitted-set drop makes no such promise.

saying these in an interview costs you the question

  • A dropped capability set is equivalent to sandboxing the process
  • Non-root plus no capabilities means a compromise is harmless
  • Fewer capabilities always means safer, regardless of which
  • Capabilities restrict which files and hosts a service can reach
  • Capabilities limit which syscalls the process can attempt

context