skip to content

What do Docker's --pid=host, --net=host and --ipc=host flags each expose to a container, and why are they escape risks?

level: middleimportance: should knowfreq 50%

answer

  1. Three flags, three walls removed
  2. Namespaces are the isolation
  3. host PID, network, IPC namespaces
  4. read /proc/<pid>/environ, reach localhost services
  5. no --privileged needed

basics

~20 s

Each drops one isolation boundary: --pid=host reveals and lets the container signal every host process, --net=host puts it directly on the host's network stack, and --ipc=host shares host shared memory. All three erode container-host isolation and need no --privileged to be dangerous.

solid answer

~40 s

A container's isolation comes from Linux namespaces; each of these flags turns one off and joins the host's. **--pid=host** shares the host PID namespace, so the container sees every host process in /proc, can `kill` them, and can read `/proc/<pid>/environ` and `/proc/<pid>/root` to steal secrets or file descriptors. **--net=host** shares the host network namespace: the container binds on host interfaces, can sniff traffic, and reaches services bound to 127.0.0.1 that were assumed unreachable (an unauthenticated database, a metrics port, a daemon on localhost). **--ipc=host** shares System V / POSIX shared memory, exposing another process's SHM segments. None of these require --privileged -- each is independently a serious weakening of the container boundary, and --pid=host in particular is the enabler for reaching host and sibling process memory.

code

bash · 3 lines
bash
# container in the host PID namespace sees and reads host processes
docker run --rm -it --pid=host alpine sh -c \
  'ps aux | head; cat /proc/1/environ | tr "\0" "\n"'

go deeper

for a junior

Recall that these three flags exist and that each removes an isolation boundary between the container and the host -- they are not performance knobs to enable by default.

for a middle

Explain that isolation is provided by Linux namespaces and that each flag joins the host's namespace instead of a fresh one. Be able to name a concrete abuse for each: reading host process env, reaching loopback-only services, touching host shared memory.

for a senior

Show how these compound into real incidents -- --pid=host plus a leaked descriptor, or --net=host plus an unauthenticated localhost daemon -- and how you would spot and refuse them in a workload spec during review.

for a principal

Own the policy question: when, if ever, a workload may take one of these, how you gate it, and how you contain a container that legitimately shares a host namespace by treating its host as a shared trust domain.

### What isolates a container A Docker container is an ordinary Linux process placed in its own set of **namespaces** -- PID, network, mount, IPC, UTS and (optionally) user. Namespaces are what make a container *look* isolated: its own process table, its own network interfaces, its own view of shared memory. The `--pid=host`, `--net=host` and `--ipc=host` flags each tell Docker *not* to create a fresh namespace of that type and to place the container in the **host's** namespace instead. Each flag removes one wall, and the walls are independent -- you can lose one without losing the others, and you do not need `--privileged` for any of them. ### --pid=host Sharing the host PID namespace means the container's `/proc` is the host's `/proc`. Concretely the container can: - **See every host process** (`ps aux` shows the whole machine), leaking command lines that often contain tokens and passwords. - **Signal host processes** -- `kill -9` a host daemon, or send SIGSTOP, causing a denial of service. - **Read `/proc/<pid>/environ`** of host and sibling-container processes, harvesting environment variables (a very common secret-delivery path). - **Reach `/proc/<pid>/root` and `/proc/<pid>/fd`**, which resolve into another process's mount view and open files -- the doorway to reading files or descriptors that belong to the host. - With `CAP_SYS_PTRACE`, **attach to and inject code into** a host process. This is why `--pid=host` is the quiet enabler behind so many breakouts: once you can address host process descriptors, other flags and bugs become weaponizable. ### --net=host The container joins the host's network namespace, so there is no NAT, no `docker0` bridge, no published-port mapping -- the process binds straight onto the host's interfaces. The risks: - **Reaching loopback-only services.** Anything an operator bound to `127.0.0.1` believing only local host processes could see it -- an unauthenticated database, an admin/metrics endpoint, an internal API -- is now directly reachable by the container. - **Sniffing and spoofing** host traffic (with the right capability), and binding privileged ports. - **Bypassing container network policy**, because there is no separate network stack to police. ### --ipc=host The container shares the host's IPC namespace: System V shared-memory segments, semaphores and message queues. It can read or corrupt shared memory that host processes use -- for example a database's shared buffers -- leaking or tampering with in-memory data. ### Why they matter as escape surfaces None of these is a bug; they are intended features for the rare workload that needs them (a profiler that must see host PIDs, a network appliance). The point for a reviewer is that they are *isolation opt-outs* that people reach for casually -- 'the app needs to see host processes', '--net=host is faster' -- and each one hands the container a lever against the host or its neighbours. They compound: `--pid=host` plus a leaked descriptor, or `--net=host` plus a localhost-bound daemon, is how a limited foothold becomes host compromise. The defence is simply not to grant them; when a workload genuinely needs one, treat the host as a shared trust domain with that container.

  • Does --net=host on its own let a container get a root shell on the host?
    No -- it breaks network isolation, not process isolation, so it is not a direct breakout. But it exposes host-local services that were assumed private (an unauthenticated Docker API on 127.0.0.1:2375, databases or admin ports bound to loopback), lets the container sniff host traffic, and disables published-port mapping. It is a pivot and information-exposure risk, and combined with a vulnerable localhost service it can lead to host compromise.
  • Why is --pid=host especially dangerous next to other containers or a secret-bearing process?
    Sharing the PID namespace lets the container read /proc/<pid>/environ and open /proc/<pid>/root and /proc/<pid>/fd for host and sibling processes, harvesting environment-variable secrets and live file descriptors. With CAP_SYS_PTRACE it can attach to those processes and inject code. It turns a single compromised container into a vantage point over every process on the machine.

Each flag is like removing one interior wall between a guest room and the rest of the house: --pid=host opens the hallway to every occupant, --net=host taps into the house's phone line, and --ipc=host shares the household's notes on the fridge.

saying these in an interview costs you the question

  • Thinks --net=host only improves performance with no security cost
  • Believes host-namespace flags still leave the container fully isolated
  • Assumes you need --privileged for these flags to be dangerous
  • Confuses --net=host with publishing a port via -p
  • Thinks --ipc=host is harmless because 'nobody uses shared memory'

context