skip to content

Your bpftrace one-liner runs fine as root on a Linux host but fails to load when you run it inside a container, and when you run it from the host against a containerised workload the PIDs and file paths it prints do not match what you see inside the container. Explain both problems and how you would work around each.

level: seniorimportance: should knowfreq 28%

answer

  1. two blocked syscalls before you even start
  2. the kernel sees everything; you may not
  3. one namespace's number, another namespace's tool
  4. the path exists, just not for you
  5. filter on the cgroup, not the name

basics

~20 s

Inside a container the loading fails because the bpf and perf_event_open system calls are usually blocked by the default seccomp profile and the required capabilities are dropped. From the host, the PIDs printed are the host's, and the file paths belong to the container's mount namespace, so neither matches what the container shows.

solid answer

~60 s

Loading a tracing program needs the `bpf()` and `perf_event_open()` system calls plus privileges — `CAP_BPF` with `CAP_PERFMON`, or `CAP_SYS_ADMIN` on older kernels — and a typical container has none of these: the default seccomp profile blocks those calls outright, the capabilities are dropped, and tracefs is not mounted. So the usual answer is not to fix the container but to trace **from the host**, where you can see every process anyway, because eBPF tracing is a kernel-wide facility and namespaces do not hide events from it. That leads straight to the second problem. bpftrace's `pid` and `tid` built-ins report the identifiers in the initial PID namespace, so they are the host-side PIDs — those match `ps` on the host, not `ps` inside the container. Likewise a path read from a syscall argument is resolved in the *target's* mount namespace, so `/app/config.yaml` may not exist on the host; you reach it through `/proc/<hostpid>/root/app/config.yaml`. To scope a trace to one container you filter on its cgroup or on the set of host PIDs, not on names you saw inside it.

code

bash · 3 lines
bash
# Reach a traced process's file, and translate its PID, from the host
cat /proc/4242/root/app/config.yaml
grep NSpid /proc/4242/status

go deeper

for a junior

Know that eBPF tracing needs root-level privileges on the host and normally will not run inside an ordinary container, and that the host can see container processes anyway.

for a middle

Name the specific barriers — the blocked bpf and perf_event_open calls under the default seccomp profile, the dropped capabilities, missing tracefs and BTF — and explain that reported PIDs are host-side because the kernel identifies tasks in the initial namespace.

for a senior

Show the whole investigation: trace from the host, translate identifiers through /proc/<hostpid>/, reach container files via /proc/<hostpid>/root/, and scope with a cgroup predicate so forked children stay in scope.

for a principal

Own the access policy: who may load kernel tracing programs on a shared host, whether a privileged tracing sidecar is permitted at all, and how that capability is granted, time-boxed and audited given it is effectively host root.

## Problem one: it will not load in a container Loading and attaching a BPF tracing program is a privileged kernel operation, and a container is deliberately configured to prevent privileged kernel operations. Several independent barriers apply, and you usually hit more than one: - **Seccomp.** The default profile applied by mainstream container runtimes blocks `bpf(2)` and `perf_event_open(2)`. Even a container running as root with capabilities intact will get `EPERM` from the system call itself. - **Capabilities.** Loading tracing programs and attaching to perf events requires `CAP_BPF` together with `CAP_PERFMON` on Linux 5.8 and later, or `CAP_SYS_ADMIN` before that. Containers drop these by default. Attaching uprobes and reading process memory can additionally involve `CAP_SYS_PTRACE`. - **Filesystem visibility.** bpftrace and bcc want tracefs (`/sys/kernel/tracing`, historically under `/sys/kernel/debug/tracing`) to enumerate tracepoints, and BTF at `/sys/kernel/btf/vmlinux`. Neither is normally mounted or readable inside a container. - **sysctls.** `kernel.perf_event_paranoid` and `kernel.unprivileged_bpf_disabled` are host-wide knobs that a container cannot change, and hardened hosts set them restrictively. You *can* force it — a privileged container with the seccomp profile unconfined and the host's tracefs and BTF bind-mounted in. That is exactly how packaged eBPF tooling images are run. But recognise what you have done: a container with the ability to load kernel programs and read arbitrary kernel memory has effectively host-level power, so this is a deliberate, audited exception, not a convenience. The better default is to trace **from the host**. eBPF instrumentation is kernel-global: a probe on a syscall tracepoint sees every process on the machine, containerised or not, because namespaces virtualise identifiers and views, not the kernel code paths themselves. Nothing is hidden from you by the container; the question is only whether you can interpret what you see. ## Problem two: the identifiers do not line up **PIDs.** A container's processes live in their own PID namespace, where the entry point is PID 1. From the host the same process has a different, larger PID. The built-ins bpftrace exposes report the value in the *initial* PID namespace — the host's view — because that is what the kernel's task structures use as the global identity. So a trace that prints `pid` gives you numbers that match `ps` on the host and the runtime's own process listing, and that will not be found by `ps` inside the container. Translating goes through `/proc/<hostpid>/status`, whose `NSpid` line lists the process's id in each nested namespace. This is not a bug and it is usually what you want during an incident: you are on the host, you want host identifiers, and they let you go straight to `/proc/<hostpid>/`. **Paths.** A filename argument captured at a syscall is a string in the caller's context, resolved against the *caller's* mount namespace. If a container has its own root filesystem, `/app/config.yaml` is a perfectly valid path there and does not exist at that location on the host. To open the file the traced process opened, go through the host's view of that process's root: ``` /proc/<hostpid>/root/app/config.yaml ``` The same applies in reverse to anything you supply — a uprobe target path must be a path *the host* can resolve, so you name the binary via `/proc/<hostpid>/root/...` rather than the path the container sees. **Mounted-over and overlay paths.** With an overlay root filesystem, the file the process opened may correspond to a file in a lower layer or in the upper directory, so the host-side location under the storage driver's directories is not somewhere to reason about; `/proc/<hostpid>/root/` is the reliable route because the kernel resolves it in the target's own namespace. ## Scoping a trace to one container Since the probe is machine-wide, you need a predicate that means "this container". Options, roughly in order of robustness: - **Filter by cgroup.** Every container's processes sit in a distinct cgroup v2 subtree, and bpftrace can compare the current task's cgroup identifier against the id of a named cgroup path. This is the most durable filter, because it survives processes forking inside the container. - **Filter by the set of host PIDs.** Obtain them from the runtime's process listing or from the cgroup's `cgroup.procs` file, then filter on `pid`. Simple, but misses processes that start after you began. - **Filter by `comm`.** Convenient and often ambiguous, since another container may run the same binary. Use it to narrow, not to scope. ## The checklist When a trace "does not work" around containers, walk it in this order: can the program load at all (privileges, seccomp, tracefs, BTF); am I in the right vantage point (host, almost always); are the identifiers I am reading in the namespace I think they are (they are the host's); and is my filter actually selecting the workload (cgroup rather than name).

  • Why is tracing from the host usually the right answer rather than making the container privileged enough to trace?
    Because eBPF instrumentation is kernel-global — the host already sees every event the container generates, so nothing is gained by moving inside. Meanwhile granting the capabilities and unconfining seccomp gives that container the ability to load kernel programs and read kernel memory, which is close to host root. You would be widening a security boundary to obtain a view you already had.
  • You have a host-side PID from a trace and need the PID as the container sees it. How do you translate?
    Read `/proc/<hostpid>/status` and look at the `NSpid` line, which lists the process's identifier in each PID namespace it belongs to, outermost first. The last value is its id in its own namespace. The reverse direction, container PID to host PID, is easier obtained from the runtime's process listing or from the cgroup's `cgroup.procs`.
  • Why is filtering by cgroup better than filtering by a list of PIDs when scoping a trace to one container?
    Because the PID list is a snapshot. A container that forks workers, restarts a supervised child, or execs a new process after you started tracing will produce events from PIDs your filter never knew about. Cgroup membership is inherited by children, so a cgroup predicate keeps selecting the workload as its process set changes.

saying these in an interview costs you the question

  • Assumes namespaces hide container events from a host-side trace
  • Thinks running as root inside the container is enough to load BPF
  • Reads a traced path on the host and concludes the file is missing
  • Expects bpftrace pid values to match ps inside the container
  • Makes the tracing container privileged without treating it as an exception

context