Why does a ptrace-based profiler fail to attach to a process inside a Docker container?
answer
- Container root is capability-limited
- The kernel guards tracing another process
- One capability, not the whole set
- Same PIDs, same files, same symbols
- A minimal image ships no profiler
basics
~10 sThe container's default privilege set excludes CAP_SYS_PTRACE, so a process that is not the target's parent is refused permission to trace it. Adding --cap-add=SYS_PTRACE to the debug run is the usual fix.
solid answer
~50 sA sampling profiler that works by tracing another process needs permission the container does not have by default. Docker drops CAP_SYS_PTRACE from the container's capability set, and on hosts where the kernel's Yama `ptrace_scope` is set to 1 only a direct parent may trace a process anyway - so a profiler started separately is denied even when it runs as the same user. Running the container with `--cap-add=SYS_PTRACE` for the debug session restores it. Two further requirements often bite: the profiler and its target must share a PID namespace and a mount namespace to see `/proc/<pid>` and the binary's symbols, and a minimal runtime image contains no profiler at all - which is what the `:debug` variants of otherwise shell-less base images exist for. Where the runtime offers an in-process alternative, such as the JVM's Flight Recorder, no ptrace privilege is needed at all.
code
bash · 4 linesdocker run --rm -it \
--cap-add=SYS_PTRACE \
--name indexer-debug \
indexer:debuggo deeper
Know that containers run with a reduced set of kernel privileges, so a tool that works on your laptop can be refused inside a container even as root, and that debugging tools sometimes need an extra flag on the run command.
Be able to name --cap-add=SYS_PTRACE as the specific grant a tracing profiler needs, explain that it is not in the default set, and say why --privileged is the wrong instrument for it.
Show the full checklist: the privilege, the shared PID and mount namespaces the tool needs to see its target and its symbols, and the fact that a minimal runtime image contains no profiler at all - so the answer is a debug variant or an in-process mechanism, used for a session and then removed.
Own the tension between a hardened default posture and on-call access: how engineers get a profiler onto a running service without any production image carrying extra privileges or tooling, and how that path stays short-lived and auditable.
### What the profiler is actually asking for Many profilers and low-level debug tools work by tracing a target process: attaching to it, reading its memory, and sampling its stacks. On Linux that is the `ptrace` mechanism, and the kernel guards it carefully because tracing a process means total control over it. Two separate gates stand in the way inside a container. **The capability gate.** A containerized process runs with a reduced capability set. `CAP_SYS_PTRACE` - the privilege that lets a process trace another it does not own the normal right to trace - is not in the default set. Attaching therefore fails with a permission error even when both processes run as the same user and even when that user is root inside the container. Root in a container is root minus a list of capabilities, and this is one of the entries on that list. **The kernel-policy gate.** Many distributions set the Yama LSM's `ptrace_scope` to 1, which restricts tracing to a direct descendant relationship: a parent may trace its child, and no one else may. A container shares the host kernel, so it shares this setting. A profiler started as a separate process is not the application's parent, so the relationship test fails - and CAP_SYS_PTRACE is exactly what lifts that restriction. There is also a historical third gate: on Docker engines before 19.03, the default seccomp profile blocked the `ptrace` syscall outright, and people worked around it by disabling seccomp for the container. Modern engines permit `ptrace` in the default profile, so today the missing capability is almost always the whole story - but you will still meet the old advice, and running an unconfined container to fix a capability problem is a bad trade. The granular fix is one flag on the debug run: ```bash docker run --rm --cap-add=SYS_PTRACE -p 127.0.0.1:8080:8080 indexer:debug ``` Reach for `--privileged` instead and you have handed the container essentially the whole capability set plus device access to solve a one-capability problem. The wider question of which capabilities a container should hold, and how to reason about the default set, is a security topic in its own right; for this workflow the relevant fact is simply that ptrace is not granted by default and can be granted alone. ### The two namespaces that must line up Permission is necessary, not sufficient. A profiler must also *see* its target. **PID namespace.** A profiler attaches by PID. PIDs are per-namespace: the application is PID 1 in its own namespace and some completely different number on the host. A tool running in a different PID namespace either cannot see the process or is looking at a different one. A profiler running inside the same container has no problem; one running elsewhere needs to be placed in the target's PID namespace deliberately. **Mount namespace.** Symbolisation needs the executable and its debug information at the paths recorded in the process. Those files exist in the container's filesystem view. A tool that sees a different mount namespace resolves stacks to hex addresses instead of function names, which looks like a broken profiler and is really a missing file. ### Minimal images have no tools A hardened runtime image - distroless-style, or built from scratch with a single static binary - deliberately contains no shell, no package manager and no profiler. Adding SYS_PTRACE to such a container grants a privilege nothing is there to use. Three responses, in rough order of preference: 1. **Use the runtime's built-in mechanism.** The JVM's Flight Recorder can be started with the process (`-XX:StartFlightRecording=...`) or enabled through the JVM's own tooling and needs no ptrace privilege, because the recording happens inside the process. Anything the runtime can do to itself avoids the whole problem. 2. **Run a debug variant of the image.** Several minimal base images publish a `:debug` tag that adds a small busybox-based shell and a few utilities to the same base, so a debug-only image is a one-word change to the `FROM` line of a debug build target rather than a different distribution. 3. **Build a dedicated debug stage.** A build target that layers the profiler and symbols onto the release stage gives you a deliberately fatter image for investigation, kept out of the release path. ### The production judgement Everything above describes an image you run *on purpose for a session*, not the posture of a service. A container that permanently holds SYS_PTRACE has a privilege an attacker who gains code execution can use to inspect and manipulate every other process in that container. The reasonable arrangement is to keep the production image minimal and unprivileged, and to make the profiling path an explicit, short-lived, auditable act - a debug variant started with the extra capability, torn down afterwards - rather than a standing grant that quietly becomes the default because it once made an incident easier.
- Why is `--cap-add=SYS_PTRACE` a better answer than `--privileged` here?It grants exactly the one privilege that was missing. `--privileged` hands the container essentially the full capability set plus device access and relaxes other confinement, so a debugging convenience becomes a container that can affect the host. If a single named capability makes the tool work, granting anything more is unjustified.
- The profiler attaches but every frame resolves to a hex address instead of a function name. What is wrong?It cannot read the executable's symbols. Either the binary in the image is stripped, or the profiler is looking at a different mount namespace than the process and so cannot open the files at the paths the process recorded. Fix it by profiling from inside the target's namespaces and using a debug variant of the image that keeps symbols.
- How would you profile a JVM service in a distroless image without adding any capability?Use a mechanism that runs inside the process rather than tracing it from outside - Flight Recorder started with the JVM through its own option, writing a recording you retrieve afterwards. No external tool attaches, so no ptrace privilege and no extra binaries in the image are required, and the runtime image stays exactly as hardened as it was.
saying these in an interview costs you the question
- Reaches for --privileged to fix one missing capability
- Thinks root inside a container has every kernel privilege
- Leaves SYS_PTRACE granted on the production container
- Ignores that PIDs differ between namespaces
- Expects a distroless image to contain a profiler
- Disables seccomp entirely instead of adding the capability