skip to content

Which container-escape scenarios does daemon-level user-namespace remapping actually mitigate, and which Docker features stop working once it is enabled?

level: seniorimportance: should knowfreq 30%

answer

  1. mapped root loses host-owned resources only
  2. shared mapping = no container-to-container gain
  3. incompatible with --privileged, --pid=host, --network=host
  4. --userns=host silently opts out
  5. new data root, re-pull images, re-own binds

basics

~20 s

It blunts abuses of in-container root against the host: writing root-owned bind-mounted files, host device access, host-privileged kernel operations. It does not stop kernel exploits, a mounted daemon socket, or container-to-container access, and it is incompatible with host PID/network sharing and privileged mode.

solid answer

~50 s

**Mitigates:** a compromised container whose root would otherwise write host files through a bind mount, chown host paths, access devices, or perform privileged operations on non-namespaced kernel state. A process that breaks out of the mount or PID namespace but stays in the user namespace lands as an unprivileged host UID. **Does not mitigate:** kernel privilege-escalation bugs (including bugs in the namespace code itself — user namespaces have historically *widened* attack surface), anything reachable through a mounted daemon socket, or container-to-container access, since all containers share the same mapping by default. **Breaks:** privileged containers and containers sharing the host PID or network namespace must opt out with `--userns=host`; external volume and storage plugins unaware of the remapping produce unreadable files. Plus the operational cost — a new storage root, re-pulled images, and every bind mount re-owned.

code

bash · 2 lines
bash
docker run --rm --userns=host --pid=host --privileged alpine sh -c 'id -u; readlink /proc/1/ns/user'
docker inspect -f '{{.HostConfig.UsernsMode}} {{.HostConfig.Privileged}}' web

go deeper

for a junior

Know the headline: it limits what container root can do to the host, and it is not a guarantee.

for a middle

List the concrete gains (bind mounts, devices, host-privileged operations) and the concrete breakages (privileged, host PID/network, some volume plugins).

for a senior

Argue the threat model out loud: shared mapping means no inter-container gain, --userns=host is a silent opt-out to audit, and kernel bugs are unaffected.

for a principal

Decide whether it earns its cost against alternatives — unprivileged images, reduced daemon privilege, run-time policy — and set the exception process for workloads that must opt out.

## What the boundary actually is Remapping changes one thing: the credentials the kernel evaluates when a container process touches a resource *outside* its namespaces. Inside the namespace, root is still root. So the honest framing is "in-container root loses authority over host-owned things", not "the container is now safe". ## Real mitigations - **Bind-mount abuse.** The classic bad day is a container with `/etc` or `/var/run` mounted read-write, or an operator who mounts a host path for convenience. Un-remapped, container root rewrites those files freely; remapped, the writes are refused unless the host path was deliberately re-owned into the mapped range. - **Ownership games.** `chown`, setuid bits and similar tricks are confined to IDs inside the mapped block, so a container cannot manufacture a setuid-root binary that means anything on the host. - **Device and kernel operations.** Loading modules, most raw device access, changing host time, and other operations on non-namespaced state require real UID 0 with the right capability in the *initial* user namespace; a mapped root does not have it, even though `capsh` inside shows a full set. - **Escape landing zone.** A process that escapes the mount or PID namespace but remains in the user namespace is unprivileged on the host — a materially smaller second stage. ## What it does not touch - **Kernel vulnerabilities.** A bug that grants UID 0 from any UID makes the mapping irrelevant. Worse, user namespaces themselves are a well-known source of privilege-escalation CVEs, so enabling them is not a free security win — some hardened distributions restrict unprivileged namespace creation for exactly that reason. - **The daemon socket.** A container able to talk to the daemon socket can ask the daemon to start a container that opts out of the mapping. Remapping is orthogonal to that risk. - **Container-to-container.** One shared mapping means every container's root is the same host UID. Two containers on a shared volume have identical authority over each other's files. Per-tenant separation needs per-workload UIDs on top, or a different isolation technology. - **The daemon's own surface.** The daemon still runs as real root; only containers are shifted. ## Incompatibilities and cost Docker documents concrete restrictions. A container cannot share the host PID or network namespace while remapped, and privileged mode is not compatible — such containers must be started with `--userns=host`, which silently returns them to the old model, so audit for that flag. External storage or volume drivers that write files with fixed UIDs produce files the container cannot read. Everyday flags such as `--read-only` are unaffected. Operationally: the daemon moves its data root to `/var/lib/docker/<uid>.<gid>/`, so images are re-pulled and disk usage can double if you keep both trees; every existing bind mount needs re-owning; and troubleshooting gets a new failure mode engineers must learn to recognise (files showing as `nobody`). ## How to position it Remapping is most valuable when you must run legacy images that insist on being root, on hosts that are not per-workload disposable. If your images already run as an unprivileged `USER` and your containers drop capabilities and mount nothing sensitive, the marginal gain is small and the friction is real — in that situation the stronger next step is usually reducing the daemon's own privilege or tightening run-time policy, not remapping alone.

  • If a container must run with --privileged, what does enabling remapping buy you for that workload?
    Nothing — privileged mode is incompatible with the remapped namespace, so the container has to be started with --userns=host and runs exactly as it did before. The value of remapping is fleet-wide only if such opt-outs are rare and audited; otherwise you have a security control with holes you cannot see from the daemon config alone.
  • Does remapping reduce the risk of a container that mounts the Docker daemon socket?
    No. Socket access is authority over the daemon itself, which still runs as real root and can start a container with --userns=host, with privileges, or with the host root filesystem mounted. The two concerns are independent: remapping constrains a container's direct kernel and filesystem access, not what it can ask the daemon to do on its behalf.

saying these in an interview costs you the question

  • Calling remapping container-escape prevention rather than mitigation
  • Claiming it isolates containers from each other
  • Assuming privileged or host-namespace containers still get the mapping
  • Ignoring that enabling user namespaces adds its own kernel attack surface
  • Not budgeting for the storage-root change and volume re-ownership

context