skip to content

Syscall & Kernel Confinement

A filter on which kernel calls a workload may make, and a profile limiting which files and devices it may touch. Asked because switching one off to fix a crash removes a whole wall quietly.

on this pageshow

questions

4

A container runtime applies a default syscall allowlist — what happens to a call it does not permit, and why do most workloads never notice?

level: middleimportance: must knowfreq 62%

answer

  1. the kernel refuses, not the platform
  2. an allowlist, not a blocklist
  3. refusal arrives as an ordinary error
  4. the shipped default denies the rare tail
  5. the account it runs as is irrelevant

basics

~20 s

A call outside the allowlist is refused in the kernel before it runs, so the process gets an ordinary error return from deep inside a library. Most workloads only ever make the common calls a shipped default already permits.

solid answer

~50 s

The filter is installed on the workload's processes before its own code starts, and the kernel then checks every call against it. It is an allowlist, so anything not named is refused — including calls nobody thought about when the list was written. The refusal comes back to the caller as a plain error return, the same shape a bad argument would produce, so the workload's own library reports a failure in its own terms and the platform usually records nothing at all; platforms can be configured to record refusals instead of refusing, but that is not the common shipped behaviour. A shipped default permits the few hundred calls ordinary programs make and denies a long administrative tail — loading code into the kernel, changing the machine's clock, attaching to another process, administering hardware. Almost every service, worker and batch job lives entirely inside the permitted set, which is why the wall is invisible until something reaches past it.

code

pseudocode · 8 lines
pseudocode
install_filter(process, profile.permittedCalls)   # before workload code starts

on every kernel_call(name, args) from process:
    if name in profile.permittedCalls:
        perform the call and return its result
    else:
        return permission_denied to the caller
        # nothing is emitted to the platform: the container keeps running

go deeper

for a junior

Recall that a container's process can only make the kernel calls it is permitted to make, that the list of permitted ones comes with the platform, and that anything else simply fails.

for a middle

Explain deny-by-default, that enforcement happens in the kernel on every call, and that the refusal reaches the process as a plain error return rather than as a platform event.

for a senior

Show that you have met this in production: the failure lands inside a library on an unusual code path, nothing in the platform's own signals points at it, and the fix is to permit one call.

for a principal

Frame the trade: denying a long tail of entry points costs almost every workload nothing and shrinks what a compromised process can reach, which is why the default belongs to the platform rather than to each team.

## What a syscall allowlist actually is Every containerised process reaches the kernel through a fixed set of entry points, the **kernel calls**. A **syscall allowlist** is the list of entry points one workload is permitted to use. It is installed on the process before the workload's own code begins, and from that moment the kernel checks every call against it. The word *allowlist* is the whole design. The list names what is **permitted**, and everything absent from it is refused. A call nobody considered when the list was written is refused by construction, rather than by a rule someone remembered to write against it. That is **deny by default**, and it is the property worth paying for: the calls that matter after an attacker gets code running inside your workload are precisely the ones nobody anticipates. Two facts about where this wall sits: - It is enforced **in the kernel, on every call** — not by the platform that started the container. The platform's part ended when it installed the filter; nothing in userspace is consulted afterwards, which is also why it costs almost nothing to run. - It is **independent of the account the process runs as**. A process can run under an unprivileged identity and be refused; it can hold a broad privilege set and still be refused. The filter is asked first, and all it knows is which call was attempted and sometimes the value of one argument. ## What the refusal looks like from inside The refusal is handed back to the caller the way any other failure of that call is: as an error return. The library that made the call takes a branch it already had for failure, and reports something about the operation it was attempting — initialising an accelerated decoder, opening a device, starting a helper — not about a policy. | | what happens | |---|---| | where the decision is made | in the kernel, at the moment of the call | | what the calling process gets | an ordinary error return, indistinguishable from a bad argument | | what the workload's log says | whatever its library says about its own operation | | what the platform records | usually nothing: the container did not exit and no event fired | Designs genuinely differ on the response, and it is worth saying so rather than asserting one: a filter of this kind can be configured to return an error to the caller, to kill the calling process outright, or to permit the call and record that it happened. The error return is the common shipped choice, and it is also the one that produces the confusing failure, because it is indistinguishable from the library's own bug. ## Why almost nothing notices the shipped default A default list is built from what real programs do. Reading and writing, allocating memory, creating threads, opening sockets, asking the time, starting and ending processes — that is a few hundred calls, and a general-purpose kernel offers a long tail beyond them. A default permits the first group and denies the tail. So: 1. A typical service, worker or batch job spends its entire life inside the permitted set and never observes the filter. 2. A workload that reaches into the tail finds the edge: hardware-accelerated media work, a profiler or debugger shipped inside the image, a process that inspects or manipulates other processes, a library that opportunistically tries a newer kernel interface for waiting on input and output. 3. In the last case the library often has a fallback path, so the workload does not fail — it just gets slower. That is the same event, in a quieter and more expensive form. This is the trade the default is making. Denying the tail costs almost every workload nothing, and removes from a compromised workload the entry points that turn code execution inside a container into a problem for the host. ## Why the wall exists at all A container is a process fenced off by the kernel, and the kernel is shared. Every call the workload is permitted to make is a piece of kernel code reachable from inside it, so the set of permitted calls is the size of the surface a bug in that code can be reached through. Shrinking the list shrinks that surface. It does not turn a shared kernel into a separate machine, and it is not a substitute for what the workload runs as — it is one wall among several, and the cheapest of them to keep. ## What an interviewer is listening for That you know the refusal is invisible **by construction** and not by an oversight; that you can say *allowlist* and mean deny-by-default; that you do not confuse the filter with file permissions or with the account; and that the fix for a blocked workload is to permit the one call it needs, not to take the wall away.

  • Can a workload widen its own filter once it is already running?
    No — filters of this kind are designed to be one-way. A confined process may narrow itself further, but it cannot restore a call it was refused, which is what makes the wall worth anything after a compromise. Widening means changing the workload's declared confinement in its spec and starting it again.
  • The same image runs fine on one host and has a call refused on another — what differs?
    Not the image: the filter comes from the platform's configuration and the workload's spec, neither of which ships inside the image. Either the two hosts run different platform defaults, or one spec asks for a relaxed or custom profile. A library that silently takes a fallback path on one host and the fast path on the other produces the same asymmetry.
  • If almost nothing notices the default, what is it actually buying?
    It removes the long administrative tail of kernel entry points from every workload that never needed them. That tail is where a compromised process goes next — loading code into the kernel, inspecting other processes, administering hardware — so the default costs ordinary workloads nothing and takes the cheap routes away from an attacker who already has code running.

It behaves like a switchboard with a fixed list of extensions: you may dial anything, but only listed numbers connect, and an unlisted one just returns a dead tone. The caller cannot tell a blocked number from a broken line.

saying these in an interview costs you the question

  • Thinks the platform blocks the call before the kernel sees it.
  • Describes it as a blocklist of dangerous calls rather than an allowlist.
  • Expects a distinctive event naming the filter whenever a call is refused.
  • Believes running as an unprivileged account exempts a process from the filter.
  • Says a refused call always terminates the container.
  • Assumes the filter ships inside the image rather than with the platform.
open as a page

How does a file-and-device reachability profile differ from a syscall allowlist when both can refuse the same workload?

level: middleimportance: should knowfreq 38%

basics

~20 s

They answer different questions. The allowlist decides which kernel calls may be attempted at all; the reachability profile decides which files and devices a call is allowed to touch. Widening one leaves the other exactly as it was.

open as a page

Which confinement wall explains a thumbnail worker that fails with a permission error inside its decoding library, on files that are readable?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Most often the syscall allowlist. A refused kernel call is returned to the caller as a plain permission error, so it surfaces as a library failure on an operation that has nothing to do with file permissions.

open as a page

When the platform's default confinement filter blocks a team's workload, what standard would you set for relaxing it?

level: principalimportance: should knowfreq 34%

basics

~20 s

Require the request to name the refused call, not the wall. Then permit that one call in confinement scoped to that workload, with an owner, an expiry and a review — and treat switching the filter off as an exception, not the next rung down.

open as a page