skip to content

User Namespaces & Non-Root Containers

Making container root stop being host root: userns-remap maps UID 0 inside the container onto an unprivileged host UID range, while a non-root USER sidesteps the question differently. Whether container root is host root is a canonical screen.

part ofDockeroverview, primer and where to startread it →
on this pageshow

questions

5

Container hardening guides say every image should end with a USER instruction instead of leaving the process as root. What actually changes on the host when you do that, and what does it not protect you from?

level: juniorimportance: must knowfreq 72%

answer

  1. container UID 0 == host UID 0
  2. no user namespace = no translation
  3. non-root starts with empty capability set
  4. numeric USER, not a name
  5. --user 0 overrides the image default

basics

~20 s

By default the container process runs as UID 0, the same UID 0 the host kernel calls root. A USER instruction runs it as an unprivileged UID, so escapes, bind mounts and host files get no root rights. It is defence in depth, not isolation.

solid answer

~50 s

Unless told otherwise a container's main process runs as UID 0, and that is the *same* UID 0 the host kernel knows — a container is a normal process wrapped in namespaces, cgroups, a reduced capability set and seccomp, not a VM. So a runtime bug, a kernel privilege-escalation flaw, or a careless bind mount of a host path all land with root-owned power. Adding `USER 10001` (a numeric UID created in the image) makes the process unprivileged: it starts with no capabilities, cannot write root-owned files exposed through bind mounts, cannot bind ports below 1024, cannot chown or load modules. That shrinks the blast radius of an application compromise a lot. What it does not do: it does not stop kernel exploits, and it does not help if the container is privileged or has the daemon socket mounted. The cost is that writable paths and volumes must be owned by that UID.

code

dockerfile · 8 lines
dockerfile
FROM eclipse-temurin:21-jre
RUN groupadd -g 10001 app && useradd -u 10001 -g 10001 -m app
WORKDIR /app
COPY --chown=10001:10001 build/libs/app.jar /app/app.jar
RUN mkdir -p /app/cache && chown 10001:10001 /app/cache
USER 10001:10001
EXPOSE 8080
ENTRYPOINT ["java","-jar","/app/app.jar"]

go deeper

for a junior

Know that no USER means UID 0 = host root, that USER makes it unprivileged, and how to create the user and chown its directories in the Dockerfile.

for a middle

Explain what a non-root process actually loses (capabilities, privileged ports, writes to root-owned bind mounts) and how to keep the app working anyway.

for a senior

Frame it as defence in depth alongside cap-drop, read-only rootfs and tmpfs, and point out that enforcement must live where containers are launched, not just in the image.

for a principal

Talk about rolling it out across a fleet: a shared hardened base image, a numeric UID convention, a run-time policy that rejects UID 0, and how you handle images that genuinely need privilege.

## Root in a container is host root A container is not a virtual machine. It is one or more ordinary Linux processes placed in namespaces (mount, PID, network, UTS, IPC and optionally user), constrained by cgroups, and stripped down with a capability bounding set, a seccomp filter and often an AppArmor or SELinux profile. There is one kernel and one table of user IDs. When a Dockerfile does not specify a user, the process runs as UID 0. Without a user namespace, that is *literally* the host's root UID: `ps -o user` on the host shows `root`, and any file it touches through a bind mount is written as root. What keeps that from being a total takeover is only the sandbox: the default capability set is a subset of full root (no `CAP_SYS_ADMIN`, no `CAP_SYS_MODULE`), seccomp blocks a few hundred syscalls, and the mount namespace hides the host filesystem. Every one of those is a defence that has had bugs. ## What USER changes `USER 10001` (or `USER 10001:10001`) sets the UID/GID for the remaining build steps and for the container's default process. The effect at runtime: - **No capabilities.** A non-root process starts with an empty effective set, so a whole class of privilege primitives is unreachable regardless of what the daemon granted. - **Host files are protected.** Anything bind-mounted read-write is subject to normal permission checks against the *host* ownership of those files. A compromised app cannot rewrite root-owned config it happens to be able to see. - **No privileged ports.** Binding below 1024 fails unless you grant `CAP_NET_BIND_SERVICE` or change the `net.ipv4.ip_unprivileged_port_start` sysctl. - **Smaller escapes.** Several historical container escapes required in-container root to be useful. ## What it does not change It is not isolation. A kernel vulnerability that grants UID 0 from any UID still wins. Mounting the daemon socket, running `--privileged`, or sharing the host PID/network namespace hands over the machine no matter which UID you picked. And `USER` in the image is a default, not a policy: `docker run --user 0` overrides it, so enforcement has to happen where containers are launched, not only where they are built. ## Building it properly Create the user at build time and use the *numeric* ID in `USER`. Names must be resolved through `/etc/passwd`, which some runtimes and orchestrator user checks cannot do, and a numeric ID is what the kernel enforces anyway. Give the UID ownership of the directories the app writes with `COPY --chown` or a `chown` in the build, keep everything else read-only, and prefer writing to `/tmp` (mountable as tmpfs) over the image filesystem so you can add `--read-only`. A useful hardening pair is `USER 10001` plus `--read-only --tmpfs /tmp --cap-drop ALL`; the first removes the privilege, the second removes the writable surface. Verify with `docker exec <c> id` rather than trusting the Dockerfile, since a base image or an entrypoint script may have switched users on you.

  • Why does the guidance say to write USER 10001 instead of USER appuser?
    The kernel only ever enforces numeric IDs; the name is resolved through the image's /etc/passwd at container start. Tooling that wants to check whether an image is non-root, and orchestrators that reject UID 0, can read a numeric value directly but cannot resolve a name without inspecting the filesystem. A numeric UID also survives a base-image change that renames or renumbers the account.
  • If USER can be overridden at run time with --user 0, is putting it in the image pointless?
    No, but it is only a default. It is the right default because most run commands do not pass --user, and it forces you to fix ownership and writable-path problems at build time. Actual enforcement belongs to whoever starts containers — a run-time policy that refuses UID 0 — with the image default making that policy cheap to satisfy.

USER is like logging into a shared server as a normal account instead of root: the walls of the room are the same, you just carry fewer keys.

saying these in an interview costs you the question

  • Saying containers are isolated so root inside does not matter
  • Claiming container root is a different root from host root even without user namespaces
  • Thinking USER cannot be overridden at run time
  • Assuming non-root stops a container escape or a kernel exploit
  • Adding USER but leaving the app writing to root-owned directories, then fixing it with chmod 777

context

open as a page

An image now ends with `USER 10001`, and the application fails to start: it binds TCP port 80 and writes to /var/run and its own cache directory. What are your options to keep the process unprivileged?

level: middleimportance: should knowfreq 46%

basics

~20 s

Move the listener to a high port and publish 80 on the host, or grant CAP_NET_BIND_SERVICE. For writes, chown the needed directories to the UID at build time, redirect state to /tmp, and mount tmpfs or a named volume for anything else.

open as a page

A colleague enables the Docker daemon's --userns-remap option. Explain what user-namespace remapping does to the UIDs a container sees versus the UIDs the host kernel sees, and where the mapping ranges come from.

level: middleimportance: should knowfreq 44%

basics

~20 s

The daemon puts containers in their own user namespace. Inside, the process still sees UID 0; the kernel maps that to an unprivileged host UID from a subordinate range in /etc/subuid and /etc/subgid, typically belonging to a user named dockremap.

open as a page

Which container-escape scenarios does daemon-level user-namespace remapping actually mitigate, and which Docker features stop working once it is enabled?

level: seniorimportance: should knowfreq 30%

basics

~20 s

It blunts abuses of in-container root against the host: writing root-owned bind-mounted files, host device access, host-privileged kernel operations. It does not stop kernel exploits, a mounted daemon socket, or container-to-container access, and it is incompatible with host PID/network sharing and privileged mode.

open as a page

A service was changed to run as UID 10001 inside the container, and the Docker daemon has user-namespace remapping enabled. Writes to its bind-mounted data directory now fail with permission denied. How do you reason about the ownership and fix it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Work out the host-side UID: remap offset plus the in-container UID (165536 + 10001 = 175537), then chown the host directory to it. Named volumes avoid this because Docker seeds their ownership from the image.

open as a page