As an ordinary non-root user on a Linux host, `perf record` fails with a permission error about performance monitoring operations. Which sysctl governs that, what do its values mean, and what additionally blocks perf inside a container?
answer
- one sysctl gates the whole subsystem
- higher number, less allowed
- user space only is the default
- a capability exists just for this
- containers add a syscall filter on top
basics
~10 sThe sysctl is kernel.perf_event_paranoid, exposed at /proc/sys/kernel/perf_event_paranoid. Higher values restrict unprivileged use: 2 permits user-space measurement only, 1 also allows kernel profiling, -1 removes restrictions. Granting CAP_PERFMON is the privileged alternative.
solid answer
~40 sPerf access is gated by `kernel.perf_event_paranoid`, readable at `/proc/sys/kernel/perf_event_paranoid`. Roughly: `3` (a Debian/Ubuntu addition) forbids unprivileged use entirely, `2` — the upstream default — allows user-space measurements only, `1` additionally allows kernel profiling, `0` allows raw and CPU-wide access, and `-1` removes the restrictions. So the usual fix is either `sysctl -w kernel.perf_event_paranoid=1` on a machine you trust, or running perf with a process that holds `CAP_PERFMON` (Linux 5.8 and later; `CAP_SYS_ADMIN` before that). Separately, `kernel.kptr_restrict` can hide kernel addresses so kernel symbols resolve to zeros even when sampling succeeds. Inside a container there is a second gate: the default seccomp profile of common runtimes blocks the `perf_event_open` syscall outright, so the capability alone is not enough.
code
bash · 9 lines# what the kernel currently allows
cat /proc/sys/kernel/perf_event_paranoid
# allow kernel profiling for this boot
sudo sysctl -w kernel.perf_event_paranoid=1
sudo sysctl -w kernel.kptr_restrict=0
# make it survive a reboot
printf 'kernel.perf_event_paranoid=1\nkernel.kptr_restrict=0\n' | sudo tee /etc/sysctl.d/99-perf.confgo deeper
Know the name kernel.perf_event_paranoid and that lowering it, or running perf with elevated privilege, is what gets a permission error out of the way.
Be able to walk the value scale from 3 down to -1, say which level enables kernel profiling versus user-space-only measurement, and show both sysctl -w and the persistent drop-in file.
Show that you treat loosening it on a shared host as a security decision, reach for CAP_PERFMON instead of blanket root where an agent needs standing access, and separate permission failures from symbolization failures when triaging.
Own the fleet policy: which host classes run permissively enough to profile in production, how that is granted and audited, and whether profiling containerised workloads happens from the host rather than by relaxing every container's seccomp profile.
## Why perf is privileged at all A CPU profiler samples the instruction pointer of whatever is running. On a shared machine, that is a side channel: sample often enough and you can infer what other users' processes are doing, and kernel-mode samples leak kernel addresses that defeat address-space layout randomisation. So the kernel gates access, and the gate is a single tunable. ## kernel.perf_event_paranoid ```bash cat /proc/sys/kernel/perf_event_paranoid sudo sysctl -w kernel.perf_event_paranoid=1 echo 'kernel.perf_event_paranoid=1' | sudo tee /etc/sysctl.d/99-perf.conf ``` The scale, from most to least restrictive: - **3** — not an upstream value; Debian and Ubuntu kernels carry a patch adding it. Unprivileged users get no perf event access at all. This is why the failure so often shows up first on Ubuntu. - **2** — the upstream default on modern kernels. An unprivileged user may measure their own processes in user space, but not profile the kernel. - **1** — additionally allows kernel profiling, which is what you need for `[k]` frames and any syscall-heavy investigation. - **0** — allows CPU-wide access, so system-wide recording with `perf record -a` becomes possible for unprivileged users. - **-1** — no restrictions. Because it is a plain sysctl, `sysctl -w` changes it for the running kernel and a drop-in under `/etc/sysctl.d/` persists it across reboots. On a developer box or a dedicated performance host, dropping to `1` is normal; on a shared multi-tenant machine it is a real security decision, not a formality. ## The capability route Rather than loosening the sysctl for everyone, you can privilege the profiler. Since Linux 5.8 the fine-grained capability `CAP_PERFMON` exists precisely for this: it grants performance monitoring access without the blanket power of `CAP_SYS_ADMIN`, which is what was required before. Running perf under `sudo` works too and is what most people do interactively, but on a host where a monitoring agent needs standing access, `CAP_PERFMON` on that one binary or service is the smaller blast radius. ## The second, quieter gate: kernel.kptr_restrict Even with sampling permitted, kernel symbol *names* can be missing. `kernel.kptr_restrict` controls whether `/proc/kallsyms` exposes real kernel addresses; when it is restrictive, unprivileged readers see zeros, and perf cannot map a kernel-mode sample to a function. The symptom is a profile that clearly has kernel time in it but shows unresolved addresses instead of function names. Setting `kernel.kptr_restrict=0` alongside a permissive paranoid level is the pair that makes kernel profiling actually readable. ## Inside a container This is where people get stuck, because the sysctl looks right and perf still fails. Three things are in play: 1. **`perf_event_paranoid` is host-wide, not namespaced.** There is no per-container value; the container inherits whatever the host kernel has. Changing it inside the container is either impossible or changes it for the whole machine. 2. **The default seccomp profile blocks `perf_event_open`.** Common container runtimes ship a seccomp allowlist that does not include the syscall perf depends on, so the call is rejected before any capability check matters. Relaxing that — running the container with an unconfined or customised seccomp profile — is required in addition to the capability. 3. **Capabilities must be granted explicitly.** A container process normally does not hold `CAP_PERFMON` (or `CAP_SYS_ADMIN`), so even with seccomp relaxed and a permissive sysctl, `perf_event_open` fails on the capability check. There is a fourth, practical point worth making in an interview: profiling a containerised process is frequently easier *from the host*. The process is an ordinary process in the host's PID namespace with a different PID, so `perf record -p <host-pid>` works with no container configuration at all. The one thing you must then solve is symbols — the binaries and their debug information live inside the container's filesystem, so perf needs to be pointed at them (perf's `--symfs` option exists for exactly this class of problem). ## Reading the error Modern perf builds print a diagnostic naming the sysctl and suggesting a value, which is a gift: read it rather than guessing. The distinction to keep straight is between *permission* failures (paranoid level, capability, seccomp) and *symbolization* failures (stripped binaries, missing debug packages, `kptr_restrict`). They look similarly like "perf isn't working" and have completely different fixes.
- Sampling now works but every kernel frame shows an unresolved address instead of a function name. What else is restricting you?`kernel.kptr_restrict`. It controls whether kernel pointers in `/proc/kallsyms` are exposed; when restricted, unprivileged readers see zeroed addresses, so perf has nothing to map kernel-mode samples onto. Setting it to 0 alongside a permissive `perf_event_paranoid` makes kernel symbols resolve. It is a symbolization control, not an access control — sampling was already permitted.
- You need to profile a process running inside a container. What is usually the simpler path?Profile it from the host. The containerised process is an ordinary host process with a different PID in the host's PID namespace, so `perf record -p <host-pid>` needs no container changes at all — no capability grant, no seccomp relaxation. The remaining problem is symbols, because the binary and its debug information live in the container's filesystem, so perf has to be pointed at that root.
saying these in an interview costs you the question
- Says the only fix is to run everything as root
- Confuses perf_event_paranoid with the ptrace_scope setting
- Thinks perf_event_paranoid can be set per container
- Blames missing symbols on a permissions problem
- Believes granting a capability alone unblocks perf in a container