skip to content

On Linux, how does a user namespace let an unprivileged user become UID 0 inside it, and why does that root not give them root on the host?

level: seniorimportance: should knowfreq 40%

answer

  1. identity, not a view of a resource
  2. three numbers, written exactly once
  3. capabilities are scoped, not global
  4. the host check uses the outside id
  5. subuid ranges need a setuid helper

basics

~20 s

A user namespace maps IDs inside it to different IDs outside, written once to /proc/PID/uid_map and gid_map. You hold a full capability set inside, but only over resources the kernel owns through that namespace; host files are still checked against your real, unprivileged outside UID.

solid answer

~50 s

Creating a user namespace with `unshare(2)` or `clone(2)` and `CLONE_NEWUSER` makes the creating process fully capable *inside* it. The identity translation comes from `/proc/<pid>/uid_map` and `/proc/<pid>/gid_map`, each line being `inside-id outside-id count`, written once and then immutable. An unprivileged process may only map its own effective UID, so `unshare --user --map-root-user` produces the single line `0 1000 1` — UID 0 inside is UID 1000 outside. Wider ranges need `CAP_SETUID` in the parent namespace, which is why the setuid helpers `newuidmap` and `newgidmap` exist, backed by the ranges an administrator granted in `/etc/subuid` and `/etc/subgid`. The reason this is not a privilege escalation is that capabilities are scoped to the namespace: `CAP_SYS_ADMIN` there governs only objects owned by that user namespace and namespaces created under it. Any access to a host file is still evaluated against the outside UID, so unmapped root-owned files stay unreadable. This is exactly the machinery rootless container tooling stands on.

code

bash · 6 lines
bash
# Become UID 0 inside a fresh user namespace as an ordinary user
unshare --user --map-root-user --fork bash -c '
  id
  cat /proc/self/uid_map
  cat /etc/shadow || echo "still denied: host check uses the outside UID"
'

go deeper

for a junior

Know that a user namespace remaps user and group IDs, that being UID 0 inside is not being root on the host, and that the mapping lives in /proc/PID/uid_map.

for a middle

Explain the map format and the write-once rule, why an unprivileged process may map only its own ID, and why gid_map requires setgroups to be denied first.

for a senior

Reason about the security boundary: capabilities scoped to a namespace, host permission checks against the outside UID, and the delegated /etc/subuid ranges that rootless tooling depends on.

for a principal

Weigh the tradeoff you are accepting fleet-wide: user namespaces give real least privilege but widen unprivileged reach into kernel code, so decide where they are enabled, how ranges are allocated, and what limits you enforce.

## The idea A user namespace is the one namespace type that changes *identity* rather than a view of some resource. Inside it, a process has a UID and GID drawn from a private ID space, and it holds a capability set relative to that namespace. The kernel translates between inside IDs and outside IDs at every boundary crossing. The important consequence is that user namespaces are the root of unprivileged isolation: creating one grants the creator a full capability set within it, and that is enough to then create mount, PID, network, UTS and IPC namespaces underneath — operations that normally require `CAP_SYS_ADMIN`. That is precisely how an ordinary user can build an isolated environment without ever being root on the host. ## Writing the maps A fresh user namespace has no mapping at all; until one is written, every ID inside resolves to the overflow ID (conventionally 65534, `nobody`). The mapping is installed by writing to two files on the process: ```bash # format: inside-id outside-id count echo '0 1000 1' > /proc/$PID/uid_map echo 'deny' > /proc/$PID/setgroups echo '0 1000 1' > /proc/$PID/gid_map ``` The rules the kernel enforces are worth knowing verbatim, because they are what makes the feature safe: - Each map file may be written **once**. After that it is immutable for the life of the namespace. - A process **without** `CAP_SETUID` in the *parent* user namespace may write only a single line mapping its own effective UID. It cannot invent a range. - Writing `gid_map` as an unprivileged process requires first writing `deny` to `/proc/<pid>/setgroups`, permanently disabling `setgroups(2)` in that namespace. Without that restriction a process could use a namespace to *drop* a supplementary group and thereby gain access to a file whose group permissions were more restrictive than its other permissions. - The writer must have the right to the target IDs; a broader mapping is delegated by an administrator through `/etc/subuid` and `/etc/subgid` and applied by the setuid-root helpers `newuidmap(1)` and `newgidmap(1)`. So `unshare --user --map-root-user bash` gives you a shell where `id` reports `uid=0(root)`, and `cat /proc/self/uid_map` shows the single line proving what that 0 really is. ## Why it is not root Two separate mechanisms keep this contained. **Capabilities are namespaced.** A capability is always held *with respect to* a user namespace. Holding `CAP_SYS_ADMIN` in a child user namespace lets you act on objects that namespace owns — mounts you created in a mount namespace owned by it, interfaces in a network namespace owned by it — and nothing else. Operations that touch resources owned by the initial user namespace, such as loading a kernel module or setting the system clock, are refused because you do not hold the capability *there*. **File access uses the translated identity.** When the process touches a file on a host filesystem, the kernel maps its inside UID back to the outside UID (1000 in the example) and performs the ordinary permission check with that. A root-owned file with mode 0600 is unreadable, exactly as it was before. IDs that are not covered by the map do not translate at all — files owned by them appear as `nobody`/`nogroup` and cannot be chowned. The same reasoning covers device access: even as inside-root you cannot open a device node you were not already permitted to open, and mounting most filesystem types is still refused because those mounts would expose host-owned objects. What you *can* mount inside a user-namespace-owned mount namespace is deliberately restricted to safe types like `tmpfs`, `proc` and bind mounts of things you can already reach. ## What it buys, and what it costs The payoff is genuine least privilege. A workload can run as inside-UID 0 — satisfying software that insists on being root — while the kernel treats it as an unprivileged user for everything that leaves the sandbox. If it escapes its mount namespace, it escapes as UID 1000. The cost is kernel attack surface. Creating a user namespace hands an unprivileged user reachability into kernel code paths that previously required root — filesystem mount code, network namespace setup, and so on — and several privilege-escalation bugs have lived exactly there. That is why the kernel offers per-user limits such as the `user.max_user_namespaces` sysctl, and why some distributions ship additional restrictions on unprivileged user-namespace creation and expect administrators to opt in. ## Inspecting it `readlink /proc/<pid>/ns/user` identifies the namespace by inode; two processes showing the same value are in the same one. Reading a process's `uid_map` from outside tells you the whole translation in three numbers, which is usually the fastest way to answer "what is this thing actually running as?" when someone shows you a process that claims to be root.

  • Why must an unprivileged process write "deny" to /proc/PID/setgroups before it can write gid_map?
    To close a permission-dropping hole. Supplementary groups can *restrict* access when a file's group permissions are tighter than its other permissions. Without the restriction, a user could enter a namespace, call `setgroups(2)` to shed the group that was denying them, and gain access they did not have. Disabling `setgroups(2)` for the namespace removes that path.
  • How does rootless container tooling map more than one UID when an unprivileged process can only map its own?
    Through delegation. An administrator grants the user a range in `/etc/subuid` and `/etc/subgid`, and the setuid-root helpers `newuidmap` and `newgidmap` write the wider map on the process's behalf after checking that the requested range is within what was granted. Without those helpers you are limited to the single-line self-map.
  • A file on a host filesystem shows as owned by nobody:nogroup inside a user namespace. What does that indicate?
    Its owning UID or GID falls outside the namespace's map, so the kernel has no inside ID to present and substitutes the overflow ID. The process cannot chown it and is treated as a non-owner for permission purposes. Extending the map is the only fix, and that requires the delegated range.

saying these in an interview costs you the question

  • Says inside-root can read any host file
  • Thinks the map can be rewritten after it is set
  • Claims capabilities are global once you are UID 0
  • Believes an unprivileged user can map arbitrary UID ranges
  • Says user namespaces remove all kernel privilege risk

context