What does a seccomp-BPF filter on Linux actually restrict, and what happens to a process when it makes a system call the filter rejects?
answer
- a filter at the syscall gate
- classic BPF over the entry data
- registers only, no pointer following
- the action picks errno, signal or death
basics
~20 sSeccomp installs a small BPF program that the kernel runs at the system call entry point, seeing the syscall number, the architecture and the raw argument registers. Its return action decides the outcome: allow, fail with a chosen errno, raise SIGSYS, or kill the thread or process.
solid answer
~40 sSeccomp puts a filter at the syscall gate itself. A process installs a classic-BPF program — via `seccomp(2)` with `SECCOMP_SET_MODE_FILTER`, or `prctl()` with `PR_SET_SECCOMP` — and from then on the kernel runs that program on every system call the process makes, before dispatch. The filter's input is the syscall number, the architecture, the instruction pointer and the six raw argument registers. Its return value chooses the action: `SECCOMP_RET_ALLOW` proceeds; `SECCOMP_RET_ERRNO` makes the call fail with an errno the filter picks, without ever running the handler; `SECCOMP_RET_TRAP` delivers `SIGSYS`; `SECCOMP_RET_KILL_THREAD` and `SECCOMP_RET_KILL_PROCESS` terminate. Installing a filter is one-way — filters are inherited across `fork()`, survive `execve()`, and cannot be removed — and an unprivileged process must first set `PR_SET_NO_NEW_PRIVS` so a filtered process cannot use the filter to subvert a setuid binary.
go deeper
Know that seccomp limits which system calls a process is allowed to make, that it is set up by the program itself or its launcher, and that a denied call can fail or kill the process.
Explain the mechanics: a BPF program evaluated at syscall entry over the number, architecture and raw argument registers, with a return action selecting allow, errno, signal or kill.
Show judgment about deploying one: build the call set by observation before enforcing, prefer an errno return for optional calls, and anticipate breakage when a library upgrade changes which call implements a function.
Position seccomp among the other boundaries. Be ready to argue what it genuinely contains versus what needs a stronger isolation layer, and to weigh the maintenance cost of an allow-list against the attack surface it actually removes.
## Where seccomp sits Seccomp is attached to the one chokepoint every process must pass through. Because all interaction with the outside world is a system call, restricting which system calls a process may make is a genuine, complete confinement of what it can do — unlike a restriction applied at a library or language level, which a process can route around by trapping directly. It is worth being precise about what it is not. Seccomp is not a permission system: it does not know about users, files or objects, and it makes no reference to what the process owns. It is a filter on the *interface*. Discretionary permissions, capabilities and mandatory access control all still apply, and seccomp composes with them rather than replacing them. ## The two modes The original strict mode allows only `read()`, `write()`, `_exit()` and `sigreturn()`, killing the process on anything else. It is too narrow for real programs and survives mostly as history. Filter mode is the one in use. The process supplies a classic-BPF program, either through `prctl(PR_SET_SECCOMP, SECCOMP_MODE_FILTER, &prog)` or the dedicated `seccomp()` call with `SECCOMP_SET_MODE_FILTER`, which also offers flags such as synchronising the filter across all threads. ## The filter's view of the world The BPF program is evaluated against a fixed input structure the kernel builds at syscall entry, containing: - the system call number, - the audit architecture value identifying the ABI in use, - the instruction pointer of the calling instruction, - the six raw argument register values. The critical limitation follows from that list: **the filter sees register values, not what they point at**. It cannot dereference a pointer argument. It can see that `openat()` was called and what its flags argument was, but it cannot see the pathname, because the string lives in user memory the filter is not allowed to follow — and even if it could, another thread could rewrite the string between the check and the kernel's own copy, a time-of-check-to-time-of-use race. So seccomp is good at "this program may never call `mount()`" and structurally incapable of "this program may only open files under /etc". The architecture field is not decorative. A filter that matches only on syscall numbers, without first checking the architecture, can be bypassed on a kernel that supports a second ABI, because the same number means a different call there. Checking the architecture first and rejecting anything unexpected is the standard opening of every correct filter. ## What happens on rejection The filter's return value carries an action in its high bits and, for some actions, data in the low bits: - `SECCOMP_RET_ALLOW` — proceed to the real handler. - `SECCOMP_RET_ERRNO` — do not run the handler; return failure to the caller with the errno encoded in the return value. The process sees an ordinary failed system call, which is often the kindest option, since a program that gets `EPERM` from an optional call may degrade gracefully rather than die. - `SECCOMP_RET_TRAP` — deliver `SIGSYS` to the calling thread. The handler can inspect the offending call and decide what to do, which is how some sandboxes emulate calls they do not permit natively. - `SECCOMP_RET_KILL_THREAD` — terminate the calling thread as though by `SIGSYS`. - `SECCOMP_RET_KILL_PROCESS` — terminate the whole thread group, which is usually what you want, since killing one thread of a multi-threaded program leaves it in an undefined state. - `SECCOMP_RET_LOG` — allow, but record the call, the mechanism behind building filters observationally before enforcing them. - `SECCOMP_RET_USER_NOTIF` — hand the decision to a supervising process over a file descriptor, which lets a monitor implement policy that BPF cannot express. When several filters are installed, all of them run and the most restrictive result wins; there is no way for a later filter to loosen an earlier one. ## The one-way door and no_new_privs A filter cannot be removed or relaxed once installed. It is inherited by children across `fork()` and preserved across `execve()`. That permanence is what makes it trustworthy: a compromised process cannot lift its own restrictions. It also creates a hazard the kernel closes explicitly. If a filtered process could execute a setuid binary, it could make chosen system calls fail inside privileged code and steer it into unsafe behaviour. So an unprivileged caller must first set the `PR_SET_NO_NEW_PRIVS` bit, which permanently prevents that process and its descendants from gaining privileges through `execve()`. Only a process with `CAP_SYS_ADMIN` may install a filter without it. ## Why filters are fragile in practice Allow-lists break when the C library changes which system call it uses to implement a function — the classic case being older and newer variants of the same operation, where a library upgrade quietly starts issuing a call the filter never permitted. The symptom is a program that worked yesterday dying on an unrelated library update, and it is why observation-then-enforce, generous handling of near-equivalent calls, and preferring an errno return over an immediate kill are all standard practice.
- Why can a seccomp filter not restrict which files a program opens?Because it only sees register values. The path argument to an open call is a pointer into user memory, and the filter is not permitted to dereference it — partly for safety, partly because another thread could rewrite the string between the filter's check and the kernel's own copy, a classic time-of-check-to-time-of-use race. Path-based policy therefore belongs to a mechanism inside the kernel's own path resolution, such as mandatory access control, not to seccomp.
- Why must an unprivileged process set no_new_privs before installing a filter?Because filters survive `execve()`. Without that bit, a process could install a filter that makes chosen system calls fail, then execute a setuid binary and steer the now-privileged code into unsafe behaviour by breaking the calls it relies on. The no_new_privs bit permanently prevents the process and its children from gaining privileges through exec, which removes that attack. A process holding CAP_SYS_ADMIN may skip it.
- When would you choose to return an errno rather than kill the process?When the call is optional to the program's correctness. A program that receives EPERM from a capability it merely probes for will usually fall back and keep working, whereas killing it turns a benign probe into an outage. Reserve termination for calls that indicate genuine compromise or that the program has no business making. Many real filters mix both, and start in a logging-only mode to discover the true call set first.
- What happens when more than one seccomp filter is installed on a process?All of them run on every system call, and the most restrictive returned action wins. A later filter can only narrow what is permitted, never widen it, and no filter can be removed. That makes layering safe: a runtime can add its own restrictions on top of whatever a supervisor already imposed, with no way for the inner layer to escape the outer one.
saying these in an interview costs you the question
- Says seccomp can restrict which paths a process opens
- Thinks a filter can be removed once installed
- Assumes rejection always kills the process
- Ignores the architecture check in the filter
- Confuses seccomp with file permissions or capabilities