skip to content

Running strace inside a Linux container fails with "ptrace: Operation not permitted" even though the shell is root. What is blocking it, and how would you trace the process anyway?

level: seniorimportance: should knowfreq 38%

answer

  1. container root is not full root
  2. capabilities, not the UID
  3. the default set drops one you need
  4. the process is ordinary on the host
  5. host PID, host privileges, host binary

basics

~20 s

Root inside a container runs with a reduced capability set that normally omits CAP_SYS_PTRACE, so the ptrace attach is refused. Either grant that capability when the container starts, or run strace from the host against the process's host PID.

solid answer

~50 s

Being UID 0 inside a container is not the same as holding every capability. Container runtimes drop most capabilities from the default set, and `CAP_SYS_PTRACE` is one of them, so attaching to any process the user does not otherwise fully own is refused. On older engines the default seccomp profile also blocked the `ptrace` syscall outright, which produced the same message; that restriction was lifted in Docker 19.03 on kernels 4.8 and newer. There are two ways forward. Grant the capability at start time — `--cap-add=SYS_PTRACE` — which requires recreating the container and so is fine in staging but often unacceptable in production. Or trace from the host: the container's process is an ordinary process in the host PID namespace with a different PID, so `strace -p <host-pid>` works with host privileges and does not need strace installed in the image at all. Paths in the resulting trace are the ones the process sees inside its own mount namespace.

code

bash · 2 lines
bash
pid=$(docker inspect -f '{{.State.Pid}}' myservice)
sudo strace -f -e trace=file -p "$pid"

go deeper

for a junior

Know that root inside a container does not carry every privilege, and that tracing needs a capability the container does not get by default. Recognising the error message as a permissions problem rather than a broken tool is the point.

for a middle

Be able to name CAP_SYS_PTRACE as the missing piece, explain that capabilities are fixed when the container is created, and say that the process is visible on the host under a different PID.

for a senior

Show the incident-safe path: find the host PID, trace from the host with host privileges and no tooling in the image, and read the trace knowing the paths are the container's. Mention seccomp and the host's ptrace_scope as the other two causes of the same message.

for a principal

Own the tradeoff between debuggability and isolation across the fleet — whether tracing capabilities are ever granted in production definitions, how host-side debugging access is controlled and audited, and what standing instrumentation removes the need to attach to a container at all.

## Root in a container is a partial root The Linux privilege model splits root's powers into capabilities. `CAP_SYS_PTRACE` is the one that lets a process attach to another process it does not already fully own. Container runtimes start containers with a deliberately reduced capability set — a default list that includes things like `CAP_CHOWN` and `CAP_NET_BIND_SERVICE` but deliberately excludes the ones that break isolation, and `CAP_SYS_PTRACE` is excluded. So a shell that reports `uid=0(root)` still cannot attach a tracer, and the error surfaces as `ptrace: Operation not permitted`. A second historical cause produces the identical message: seccomp. Docker's default seccomp profile once denied the `ptrace` system call entirely. That was relaxed in Docker 19.03 for kernels 4.8 and later, so on a current engine the capability is usually the only obstacle — but on an older platform, or a hardened profile written in-house, seccomp can still be the blocker. If adding the capability alone does not fix it, seccomp is the next thing to check. A third possibility exists on the host side. The Yama LSM exposes `/proc/sys/kernel/yama/ptrace_scope`, with 0 meaning classic permissive behaviour, 1 restricting non-privileged tracers to their own descendants, 2 requiring `CAP_SYS_PTRACE`, and 3 disabling attachment entirely until reboot. A process holding `CAP_SYS_PTRACE` is unaffected by scope 1, so this is rarely the container answer — but on a host set to 3 nothing can attach anywhere, and that is worth checking before you blame the runtime. ## Route one: grant the capability Starting the container with `--cap-add=SYS_PTRACE` restores the ability, and on an older engine you may additionally need to relax the seccomp profile. The catch is that capabilities are fixed at container creation, so this means recreating the container — which destroys the state of the incident you were investigating and, on a production workload, permanently widens the container's privileges if the flag is left in the deployment definition. It is the right answer in a development or staging environment where you can reproduce the problem on demand, and usually the wrong answer for a live incident. ## Route two: trace from the host This is the technique that matters operationally. A containerised process is an ordinary process on the host; the PID namespace only changes the number it sees for itself. From the host you have full privileges, so: 1. Find the host PID. `docker top <container>` or `crictl inspect` prints it, or you can grep the process list — the process's `/proc/<pid>/cgroup` on the host names the container's cgroup path, which is how you confirm you have the right one. 2. `strace -f -p <host-pid>` and read as usual. Three things follow from the namespace layout, and they are the details an interviewer looks for: - **strace does not need to be in the image.** You are running the host's binary, so debugging tools never have to ship in a minimal production image. This is the answer to "our image is distroless and has no shell". - **Paths in the trace are container paths.** strace prints the string the process passed, and the tracee resolves it in its own mount namespace. So `/etc/app.conf` in the output means the container's `/etc/app.conf`, which on the host lives somewhere under the overlay filesystem. With `-y`, descriptor annotations come from reading the tracee's `/proc/<pid>/fd`, so they too show the container's view. - **Threads and children are inside the container too.** `-f` follows them normally; there is nothing namespace-specific about that. ## The wider pattern This is one instance of a general rule that applies to the whole host toolkit: the isolation a container provides also blinds the tools you would normally run inside it, and the fix is almost always to observe from the host, where the kernel objects genuinely live. Tracing is the case with the sharpest error message, because the kernel refuses outright rather than quietly reporting the wrong thing. A final practical warning: everything about ptrace overhead still applies. Tracing a containerised process from the host stops it on every syscall exactly as it would anywhere else, so narrow the traced set with `-e trace=`, keep the window short, and remember that if the workload is behind an orchestrator's liveness probe, slowing it down enough to fail that probe will get the container killed and restarted underneath you.

  • Why does tracing from the host avoid needing strace inside a minimal image?
    Because you run the host's strace binary against the host PID of the process. The tracer lives entirely outside the container's filesystem; the container only needs to be a normal process on the host, which it is. That keeps debugging tools out of production images altogether, which is desirable for both image size and attack surface.
  • You add the capability and it still fails. What do you check next?
    Two things. First, seccomp: an older engine or a custom profile may deny the `ptrace` syscall regardless of capabilities. Second, the host's Yama setting — `/proc/sys/kernel/yama/ptrace_scope` at 3 disables attachment system-wide until reboot, and at 2 requires the capability you have just granted. Also confirm no other tracer already holds the process, since only one is permitted.
  • In a trace taken from the host, a path is shown as /etc/app.conf but the host has no such file with that content. Why?
    strace prints the path string the process passed, and that path is resolved in the container's mount namespace. The real file lives inside the container's filesystem, typically under an overlay upper or lower directory on the host. To inspect it, look through the container's own root — `/proc/<host-pid>/root/etc/app.conf` reaches it directly from the host.

saying these in an interview costs you the question

  • Assuming UID 0 in a container implies full kernel privileges
  • Blaming the missing strace binary rather than the capability
  • Recreating a production container mid-incident just to add a flag
  • Reading container paths in a host-side trace as host paths
  • Granting SYS_PTRACE permanently in the deployment definition

context