skip to content

Explain what Docker's rootless mode changes about how the daemon and its containers run, including the role of `/etc/subuid` and the `newuidmap` helper, and what ownership a host file gets when a rootless container writes it as its own root user.

level: middleimportance: must knowfreq 44%

answer

  1. dockerd runs as you, in a user namespace
  2. /etc/subuid: alice:100000:65536
  3. newuidmap/newgidmap write /proc/pid/uid_map
  4. container uid 0 = host uid 100000
  5. RootlessKit + slirp4netns + fuse-overlayfs

basics

~20 s

Rootless mode runs the daemon as an ordinary user inside a user namespace. Ranges in /etc/subuid and /etc/subgid, applied by the setuid helper newuidmap, map container uid 0 to a high unprivileged host uid — which is what host files end up owned by.

solid answer

~50 s

In rootless mode `dockerd` itself runs as an unprivileged user, started per-user (typically via `dockerd-rootless-setuptool.sh` and a systemd user unit), with its own socket under `$XDG_RUNTIME_DIR` and its data under `~/.local/share/docker`. The enabling kernel feature is the **user namespace**. Your account is allocated a subordinate range in `/etc/subuid` and `/etc/subgid`, e.g. `alice:100000:65536`. Writing a multi-range mapping into `/proc/<pid>/uid_map` requires privilege, so Docker calls the setuid helpers **`newuidmap`/`newgidmap`** from the `uidmap` package, which validate the request against those files. Inside the namespace the process sees uid 0 with a full capability set, but the kernel translates it to host uid 100000. So a file created by container root on a bind mount is owned by host uid 100000, not 0 — and the container can only touch files `alice` could already touch. The win: a container escape or a daemon compromise lands you as an unprivileged user, not root.

code

bash · 11 lines
bash
grep "^$USER" /etc/subuid /etc/subgid
# kdob:100000:65536

dockerd-rootless-setuptool.sh install
export DOCKER_HOST=unix://$XDG_RUNTIME_DIR/docker.sock

mkdir data
docker run --rm -v "$PWD/data:/data" alpine sh -c 'id -u; touch /data/f'
# 0
ls -n data
# -rw-r--r-- 1 100000 100000 0 ... f

go deeper

for a junior

Know the one-line definition: the daemon runs as a normal user and container root is a fake root mapped to a high host uid.

for a middle

Explain /etc/subuid ranges, the newuidmap helper, and predict host ownership of files written through a bind mount.

for a senior

Add the supporting stack (RootlessKit, slirp4netns, fuse-overlayfs), per-user state directories, and be precise about what rootless does not protect.

for a principal

Weigh it as a containment strategy: residual kernel and user-namespace attack surface, hardened-distro policies that disable user namespaces, and where it fits alongside VM-level isolation.

## The problem rootless mode solves Classic Docker has a root-owned daemon: a long-lived process with full privileges that builds images, mounts filesystems and configures networking on behalf of anyone who can reach its socket. A bug in it, or an escape from a container it launched, is an immediate host compromise. Rootless mode removes the root process from the picture. ## User namespaces and subordinate id ranges A **user namespace** is a kernel mechanism giving a process its own mapping from ids inside the namespace to ids outside. An unprivileged user may create one and be uid 0 within it with a full capability set — but those capabilities apply only to resources owned by the namespace, so they do not translate into power over the host. A namespace needs a *mapping*. Mapping only your own uid gives a single id, which is not enough to run a distro image that expects root plus service accounts. So Linux provides **subordinate id ranges**: `/etc/subuid` and `/etc/subgid` contain lines such as `alice:100000:65536`, meaning alice may use host uids 100000–165535 inside namespaces she creates. Writing that multi-entry map into `/proc/<pid>/uid_map` is privileged, which is where **`newuidmap`** and **`newgidmap`** (from the `uidmap` package, installed setuid-root or with file capabilities) come in: Docker's RootlessKit execs them, they validate the request against `/etc/subuid`/`/etc/subgid`, and write the map. These two small helpers are the only privileged code involved. ## What the mapping means in practice With the range above, container uid 0 is host uid 100000, container uid 1000 is host uid 100999, and so on. Consequences worth stating explicitly: - A file written to a bind mount by container root appears on the host owned by 100000 (often displayed as a bare number, since no such user exists in `/etc/passwd`). - To let a rootless container write into an existing host directory you must `chown` it into the mapped range, or use a mapping that passes your own uid straight through (Podman spells this `--userns=keep-id`). - The container can never read a host file its owning user could not already read, because the effective host identity always sits inside that subordinate range. ## The rest of the rootless stack Rootless Docker also needs privileged-looking operations for networking and mounts, and gets them from userspace components: **RootlessKit** creates the namespaces and handles port forwarding, **slirp4netns** (or pasta) provides a userspace TCP/IP stack for outbound connectivity, and storage uses native `overlay2` on modern kernels or **fuse-overlayfs** where unprivileged overlay mounts are unsupported. State lives under `~/.local/share/docker` and the socket at `$XDG_RUNTIME_DIR/docker.sock`, so each user has an independent daemon, independent images and no shared cache. ## What it protects and what it does not It collapses the blast radius of daemon bugs and container escapes to one unprivileged account. It does not protect you *within* that account — anything the user can read, a compromised container can read, including their SSH keys. It does not stop kernel exploits that break out of the user namespace itself (user namespaces have historically widened kernel attack surface, which is why some hardened distros gate them). And it is orthogonal to the application's own exposure. Rootless mode is a containment measure, not an application security control.

  • A rootless container needs to write into an existing host directory owned by your user, but gets permission denied. Why, and how do you fix it?
    The container process is not running as your uid: container root maps to the first subordinate uid, e.g. 100000, which has no write access to a directory owned by uid 1000. Either chown the directory into the mapped range, run the container process with a container uid that maps back to your host uid, or use an identity-preserving mapping such as Podman's --userns=keep-id.
  • Rootless mode still uses user namespaces — so is it the same thing as running a process as a non-root USER inside the container?
    No. A non-root USER only changes the identity inside the container; the daemon that launched it is still root on the host. Rootless mode moves the daemon itself into an unprivileged account and namespace, so the entire stack — build, storage, networking, container — is confined. They compose well and are best used together.

The subordinate id range is a block of guest badge numbers issued to you: inside your own suite the guest wearing badge 0 is 'the boss', but at the building's front desk that badge is just number 100000.

saying these in an interview costs you the question

  • Saying rootless means containers 'run as non-root' while assuming the daemon is still root
  • Not knowing files land on the host owned by a mapped high uid, then chmod 777 to work around it
  • Claiming rootless removes the need for user-namespace or capability hardening entirely
  • Assuming images and caches are shared between users — each rootless daemon has its own store

context