Most Linux distributions ship with unprivileged eBPF disabled via kernel.unprivileged_bpf_disabled. If the verifier already proves every program safe, why is unprivileged loading turned off, and how would you decide what privileges to grant eBPF tooling across a fleet?
answer
- the proof is about architectural execution
- CPUs speculate past checks
- the analyser is itself attack surface
- split from CAP_SYS_ADMIN in 5.8
- tracing capability equals kernel-memory read
basics
~20 sThe verifier proves architectural safety, not speculative-execution safety, and it is itself a large piece of attack surface whose bugs have been privilege escalations. Distributions therefore disable unprivileged loading and expect eBPF tooling to run with CAP_BPF plus the capability its program type needs.
solid answer
~50 sTwo separate arguments push in the same direction. First, the verifier reasons about *architectural* execution; CPUs speculate past the bounds checks it inserts, so an unprivileged program can be shaped into a Spectre-style gadget that leaks kernel memory. The verifier applies extra restrictions to unprivileged programs — index masking, barriers, refusing to let pointer values become observable — but it is a mitigation, not a proof. Second, the verifier is thousands of lines of complex kernel code, and verifier bugs have historically been directly exploitable local privilege escalations. Distributions therefore default `kernel.unprivileged_bpf_disabled` to 1, and upstream added a config option to make that the built-in default from Linux 5.16. For a fleet I would not chase unprivileged loading at all: grant loaders `CAP_BPF` plus what the program type needs — `CAP_PERFMON` for tracing, `CAP_NET_ADMIN` for XDP and tc — treat that as effectively root-equivalent for reading kernel memory, load and pin programs at startup from a small set of audited tools, and keep kernels patched.
go deeper
Know that loading eBPF programs normally requires root or a specific capability, and that ordinary users cannot load them on a typical distribution.
Explain that kernel.unprivileged_bpf_disabled gates the load path, that CAP_BPF was split from CAP_SYS_ADMIN in Linux 5.8, and that tracing and networking program types each need an additional capability.
Articulate why the safety proof is insufficient for untrusted loaders — speculative execution outruns the inserted checks, and verifier bugs have been privilege escalations — and design agents that load at startup, pin, and drop privilege.
Own the fleet-wide tradeoff: keep the hardened default, curate who may load, be candid that the tracing capabilities are root-equivalent for kernel memory, and accept the kernel-patching cadence that an eBPF-heavy platform implies.
## What the verifier's guarantee does and does not cover The verifier proves two things about the instruction stream: memory accesses are in bounds, and the program terminates. Both proofs are about the **architectural** semantics of the program — what the machine is defined to do. Modern CPUs also do things the architecture does not define: they execute past unresolved branches speculatively and roll back the results, leaving traces in the caches. That gap is the problem. A bounds check the verifier inserted is a branch, and the CPU may speculate past it, performing the out-of-bounds load transiently. The architectural result is discarded, but the cache state is not, and an attacker who can shape the program and time the cache reads kernel memory. eBPF is an unusually good vehicle for this because the attacker gets to *write the gadget* and run it in kernel context, repeatedly, with precise control. The verifier does fight back for unprivileged programs: it masks indices so out-of-bounds speculation reads in-bounds addresses, inserts speculation barriers, refuses to let raw pointer values become observable to user space, and applies stricter arithmetic rules. These are real mitigations, and they are also a moving target that has needed repeated fixing. ## The second argument: the verifier is attack surface Even leaving speculation aside, the verifier is a large, intricate piece of kernel code doing range tracking and abstract interpretation on attacker-supplied input. A bug in its arithmetic — a range it computes as narrower than reality — hands an attacker a verified program that performs arbitrary kernel reads and writes. Several such bugs have been published, each a clean local privilege escalation on kernels where unprivileged loading was enabled. That is the crux: exposing the verifier to unprivileged users means the security of the box depends on it having no bugs. Distributions decided not to take that bet. ## What the sysctl actually does `kernel.unprivileged_bpf_disabled` controls whether a process without the relevant capabilities may call `bpf()` to load programs. Debian and Ubuntu set it to 1 by default, RHEL likewise, and since Linux 5.16 the kernel has a build-time option making "off" the compiled-in default. On modern kernels the value 2 exists to mean disabled-but-re-enablable, whereas writing 1 is one-way for the lifetime of the boot — a deliberate anti-tamper property, since a hardening setting an attacker can simply flip back is not hardening. One practical consequence: this affects **loading**, not running. Programs already loaded and pinned keep working, and a user-space tool that only reads maps through pinned file descriptors does not need load privilege at all. ## The capability model Before Linux 5.8 everything needed `CAP_SYS_ADMIN`, which is close to root. 5.8 split it: - **`CAP_BPF`** — create maps and load programs. - **`CAP_PERFMON`** — additionally required for tracing programs such as kprobes and perf events. - **`CAP_NET_ADMIN`** — additionally required for networking hooks such as XDP and tc. The split is genuinely useful for reducing what a compromised agent can reach, but it is important to be honest about it in an interview: `CAP_BPF` plus `CAP_PERFMON` lets a process read arbitrary kernel memory through tracing programs. That is root-equivalent in every way that matters for confidentiality. The split reduces blast radius against *some* attacks; it does not turn an eBPF agent into an unprivileged workload. ## How I would decide for a fleet **Do not enable unprivileged eBPF.** The use case — letting ordinary users load their own programs — is rare, and the exposure is the entire verifier plus the speculation surface. Keep the sysctl at its hardened default and treat any request to lower it as needing a very specific justification. **Make loading a privileged, curated act.** A small set of audited agents (your profiler, your network observability agent, your security monitor) load at startup with `CAP_BPF` plus the one extra capability their program types need, pin what must outlive them, and drop privilege afterwards. Engineers who need ad-hoc tracing get it through a controlled path — a privileged helper, or sudo access to specific tools — rather than a general grant. **Budget for kernel currency.** Because the mitigations live in the verifier, an eBPF-heavy platform inherits a dependency on kernel updates. If your fleet cannot take kernel patches on a reasonable cadence, that is an argument for fewer eBPF agents, not for looser privileges. **Watch for the container case.** A container granted `CAP_BPF` and `CAP_PERFMON` can trace the *host* — namespaces do not scope tracing programs. Those capabilities on a container are an isolation decision, not a convenience flag, and they belong under the same review as host root.
- What exactly does the verifier do differently for unprivileged programs?It applies a stricter regime: masking indices so speculative accesses stay in bounds, inserting speculation barriers, refusing constructs where a pointer value could become observable to user space, and tightening pointer arithmetic. It also keeps the old 4,096-instruction program ceiling. These are mitigations layered on the ordinary safety proof, not part of it.
- Does granting CAP_BPF plus CAP_PERFMON to a monitoring agent meaningfully reduce risk versus running it as root?Somewhat, and less than people hope. It removes the rest of CAP_SYS_ADMIN, so the agent cannot mount filesystems or load kernel modules. But tracing programs can read arbitrary kernel memory, so the agent remains root-equivalent for confidentiality. Treat it as a privileged component, review it accordingly, and do not describe it as unprivileged.
- A team asks for CAP_BPF and CAP_PERFMON on their container so they can run bpftrace. What is your answer?That those capabilities are not scoped by namespaces — the container could trace every process on the host, including other tenants. It is a host-root-equivalent grant wearing a container flag. On a shared node I would refuse and offer a node-level agent or a privileged break-glass path; on a dedicated single-tenant node it may be acceptable with review.
- Is disabling unprivileged eBPF a problem for programs that are already loaded?No — the sysctl governs the load path only. Programs loaded and pinned by a privileged process keep running, and user-space tools that read results through pinned map file descriptors need no capability at all. That separation is what makes load-at-startup-then-drop-privilege a workable pattern.
saying these in an interview costs you the question
- Says the verifier makes eBPF safe for anyone to load
- Treats CAP_BPF as an unprivileged capability
- Ignores speculative execution as out of scope
- Thinks namespaces confine a container's tracing programs
- Assumes the sysctl can be toggled back freely at runtime