skip to content

How do you debug a shell-less Docker container from a toolbox container that shares its namespaces?

level: seniorimportance: should knowfreq 38%

answer

  1. Bring a second container, not a shell
  2. Docker can share PID, network and IPC
  3. One namespace has no sharing flag
  4. The target's init becomes PID 1
  5. Reach its files through /proc/1/root

basics

~20 s

Start a throwaway container with tools using --pid=container:<target> and --network=container:<target>, so it sees the target's processes and network while running its own shell. The mount namespace is not shared, so read the target's files at /proc/1/root.

solid answer

~40 s

Run a second container that has the tools and point it at the first one's namespaces: `docker run --rm -it --pid=container:hugo-web --network=container:hugo-web --cap-add=SYS_PTRACE busybox sh`. The shell and utilities come from the toolbox image, so the target image can stay a scratch or distroless one; `--pid=container:` means `ps` shows the target's processes with the target's entrypoint as PID 1, and `--network=container:` means you share its interfaces and loopback. What Docker gives you no flag for is the **mount namespace** — the toolbox keeps its own filesystem. You reach the target's files through the shared PID namespace instead, at `/proc/1/root/...`, along with `/proc/1/environ`, `/proc/1/cgroup` and `/proc/1/limits`. This needs only Docker API access, not root on the engine host, which is why it is usually the practical route.

code

bash · 11 lines
bash
docker run --rm -it \
  --pid=container:hugo-web \
  --network=container:hugo-web \
  --cap-add=SYS_PTRACE \
  busybox sh

# inside the toolbox:
ps -ef                                  # the target's processes, its init as PID 1
ls -l /proc/1/root/etc/nginx/           # the target's filesystem
tr '\0' '\n' < /proc/1/environ          # the target's environment
cat /proc/1/cgroup /proc/1/limits       # its cgroup path and rlimits

go deeper

for a junior

Know that the tools can come from a second container rather than the target image, and that Docker can start that container inside the first one's process and network view.

for a middle

Explain which namespaces docker run can join and which it cannot, and why /proc/1/root is the bridge to the target's filesystem once you share its PID namespace.

for a senior

Show that you can debug production without mutating the workload: no restart, no writes, a disposable toolbox, and awareness that the target must still be running for any of it to work.

for a principal

Frame it as a capability question: anyone who can start containers can join another container's namespaces, so this is a Docker-socket trust decision, and the organisation should standardise a vetted toolbox image rather than improvising during incidents.

### The idea If you cannot put tools in the target image, run them in a second container that has been placed in the *same namespaces*. Docker exposes exactly this on `docker run`: `--pid=container:<name-or-id>`, `--network=container:<name-or-id>` and `--ipc=container:<name-or-id>` join an existing container's PID, network and IPC namespaces respectively. The toolbox container is disposable, so use `--rm`. Take a static Hugo site served by an nginx binary in a minimal image, `hugo-web`, that returns 502s for a subset of paths and cannot be restarted while you are looking at it: ``` docker run --rm -it \ --pid=container:hugo-web \ --network=container:hugo-web \ --cap-add=SYS_PTRACE \ busybox sh ``` ### What you get, namespace by namespace **PID.** `ps -ef` inside the toolbox now lists the target's processes, and the target's entrypoint appears as PID 1 because you joined *its* PID namespace, where it is the init. You can see worker counts, arguments, states and parentage — for `hugo-web`, whether the 17 worker processes the config asks for actually exist. **Network.** You share the target's interfaces, its loopback and its `/proc/net`. A listener bound to `127.0.0.1` in the target is reachable from the toolbox, which is not true from any other container. **Mount — not shared.** This is the part candidates get wrong. There is no `--mount=container:` flag; a container's mount namespace is what makes its filesystem be its image, and Docker does not offer to share it. So the toolbox sees the toolbox's own `/etc`, `/bin` and `/tmp`, not the target's. ### `/proc/1/root` — the bridge to the target's filesystem Because you joined the PID namespace, the target's init is visible as PID 1, and `/proc/<pid>/root` is a magic symlink that resolves paths using that process's root directory. So from the toolbox: ``` ls -l /proc/1/root/etc/nginx/ cat /proc/1/root/usr/share/nginx/html/index.html | head tr '\0' '\n' < /proc/1/environ cat /proc/1/cgroup /proc/1/limits ls -l /proc/1/fd ``` You can read the target's config, its environment, its cgroup path, its resource limits and its open file descriptors — with the toolbox's tools. Note the ordering dependency: `/proc/1/root` is only the *target's* root because you shared its PID namespace. Without `--pid=container:`, PID 1 is the toolbox's own init and you would be reading yourself, which is a silent, confusing failure rather than an error. Access to `/proc/<pid>/root` is permission-checked like a debugger attach, so the toolbox generally needs to run as root (the default) and, in tightened setups, `--cap-add=SYS_PTRACE`. That capability is also what lets you actually attach `strace` or a debugger to the target's processes; the engine's default seccomp profile and any hardening you have layered on can still get in the way. ### What it costs and what it does not The target is untouched: no restart, no new binaries in its writable layer, no change to its image. The toolbox disappears on exit thanks to `--rm`. What it does cost is a container with elevated visibility into another workload, so on a shared engine this is a privileged action even though it needs no host root — anyone who can start containers can join any container's namespaces, which is one more reason the Docker socket is a trust boundary. ### Choosing this over `nsenter` The two routes overlap but their prerequisites differ. `nsenter` needs **root on the engine host** and gives you the host's full tool set. The toolbox route needs only **Docker API access** and gives you whatever the toolbox image ships, which you control. On a managed or remote engine, and on Docker Desktop where the host is a VM, the toolbox route is usually the only one you can actually perform. Pick a small toolbox image you trust and standardise on it, so on-call is not pulling an arbitrary image into production during an incident. ### Limits worth stating The target must be running — namespaces vanish with the last process, so a container that has already exited cannot be joined. You are reading a live process tree, so treat everything as a snapshot of a moving system. And anything you learn through `/proc/1/root` is the *current* filesystem state, including whatever the running process has already written.

  • Why can't you just share the target's mount namespace with a flag the way you share its PID namespace?
    Docker exposes no such flag. The mount namespace is what makes a container's filesystem be its image, and sharing it would blur the two containers into one root filesystem. The supported workaround is the shared PID namespace plus `/proc/1/root`, which resolves paths using the target init's root directory and therefore lets the toolbox read the target's files without joining its mounts.
  • You forgot --pid=container: and only passed --network=container:. What silently goes wrong?
    `/proc/1/root` now points at the toolbox's own init, so you read the toolbox's filesystem while believing you are reading the target's — a wrong answer with no error message. `ps` likewise shows only the toolbox's processes. Always confirm you are in the right PID namespace, for example by checking that PID 1's command line is the target's entrypoint.
  • What does --cap-add=SYS_PTRACE add here, and is it always required?
    It grants the capability behind debugger-style access, so you can attach `strace` or a debugger to the target's processes and read another process's memory. Plain reads under `/proc/1/root` often work without it when the toolbox runs as root in the same user namespace, but hardened setups, user-namespace remapping and seccomp policy can block them. Add it when you need tracing, not by reflex.
  • Does this route work against a container that has already exited?
    No. Namespaces exist only while a process holds them, so once the last process in the container is gone there is nothing to join, and `--pid=container:` fails. That is the dividing line in triage: live containers get namespace-sharing and process inspection, exited ones get post-mortem evidence such as the engine's recorded exit state, logs and filesystem contents.

saying these in an interview costs you the question

  • Thinks --pid=container also shares the filesystem
  • Expects the toolbox's tools to appear in the target
  • Reads /proc/1/root without joining the PID namespace
  • Says the target must be restarted with new flags
  • Believes this route needs root on the engine host
  • Assumes namespaces survive after the container exits

context