skip to content

After moving a CI runner to rootless Docker, a job can no longer bind port 80, `--memory` limits appear to be ignored, and throughput on published ports has dropped noticeably. Walk through the known limitations of a rootless daemon and how you would address each of these.

level: seniorimportance: should knowfreq 38%

answer

  1. ports <1024: sysctl ip_unprivileged_port_start / setcap / proxy
  2. limits need cgroup v2 + Delegate=yes
  3. slirp4netns/pasta = userspace packets = slower
  4. builtin port driver hides source IP
  5. enable-linger or the daemon dies at logout

basics

~20 s

Rootless daemons cannot bind ports below 1024, need cgroup v2 with systemd delegation for resource limits, and route traffic through a userspace network stack that costs throughput. Fix with the unprivileged-port sysctl, Delegate=yes, and a suitable port driver or host proxy.

solid answer

~50 s

Three separate limitations: **Low ports.** An unprivileged process cannot bind below 1024. Options: publish a high port and front it with a host reverse proxy; lower `net.ipv4.ip_unprivileged_port_start`; or grant `cap_net_bind_service` to the `rootlesskit` binary. On CI I prefer the proxy or high-port option. **Resource limits.** `--memory`/`--cpus` require cgroup v2 with systemd delegating a subtree to the user — a drop-in with `Delegate=yes` for `[email protected]`. On cgroup v1 or without delegation the daemon cannot apply limits. **Throughput.** Rootless networking goes through RootlessKit plus a userspace stack (slirp4netns, or pasta), which costs CPU and bandwidth versus kernel-native bridge and NAT. Mitigate by tuning MTU, choosing the port driver deliberately, or keeping hot paths inside the rootless network namespace. I would also flag the others: no image sharing between users, restricted device passthrough and AppArmor, no swarm overlay networks, and a daemon that dies at logout without `loginctl enable-linger`.

code

bash · 17 lines
bash
# cgroup v2 delegation so --memory/--cpus work
sudo mkdir -p /etc/systemd/system/[email protected]
printf '[Service]\nDelegate=cpu cpuset io memory pids\n' | \
  sudo tee /etc/systemd/system/[email protected]/delegate.conf
sudo systemctl daemon-reload

# keep the user daemon alive after logout
sudo loginctl enable-linger ci

# allow binding port 80 unprivileged (host-wide)
sudo sysctl -w net.ipv4.ip_unprivileged_port_start=80

# preserve real source IPs on published ports
systemctl --user set-environment DOCKERD_ROOTLESS_ROOTLESSKIT_PORT_DRIVER=slirp4netns
systemctl --user restart docker

docker info | grep -iA3 -e warning -e cgroup

go deeper

for a junior

It is enough to know rootless cannot bind ports under 1024 and that some features are unavailable.

for a middle

Name the three big categories — low ports, cgroup delegation for limits, userspace networking — and one workaround each.

for a senior

Triage from symptoms, distinguish kernel constraints from implementation costs, and weigh host-wide sysctl changes against a narrower proxy or capability fix.

for a principal

Decide whether the runner fleet should be rootless at all: quantify throughput and image-duplication cost against the isolation benefit, and set the standard configuration.

## Ports below 1024 Linux reserves ports under 1024 for privileged processes, so a rootless daemon fails on `-p 80:80`. Three fixes, in ascending order of blast radius: publish a high port (`-p 8080:80`) and terminate 80 at a host-level proxy or load balancer; set `sysctl net.ipv4.ip_unprivileged_port_start=80`, which lowers the reservation host-wide for every process, not just Docker; or `setcap cap_net_bind_service=+ep` on the RootlessKit binary, which is narrower but reintroduces a small privileged component. On a CI runner the first option is usually right, because nothing external needs port 80 on the runner. ## Resource limits and cgroups Applying `--memory`, `--cpus` or `--pids-limit` means writing to cgroup files the user does not own. This works only with **cgroup v2** plus **systemd delegation**: systemd hands each logged-in user a subtree of the cgroup hierarchy and the daemon writes inside it. On cgroup v1, or without delegation, limits cannot be applied and containers run unconstrained — dangerous on shared CI, where one runaway job can OOM the box. The remedy is a drop-in for `[email protected]` with `Delegate=cpu cpuset io memory pids` and a kernel booted with unified cgroups. `docker info` lists the missing capabilities in its warnings section. ## Networking cost and source addresses Rootful Docker builds a bridge and NATs with kernel machinery. Rootless Docker cannot create host network interfaces, so RootlessKit puts the daemon in its own network namespace and connects it with a userspace TCP/IP implementation — historically **slirp4netns**, increasingly **pasta**. Every packet is copied through userspace, which costs CPU and caps throughput well below native, and small default MTUs make it worse. There is a subtler symptom too: with the built-in port driver, incoming connections appear to the container as coming from 127.0.0.1, breaking IP-based logging and allow-lists; switching RootlessKit's port driver to the slirp4netns driver preserves the real source address at some performance cost. Choose the driver deliberately per workload. ## The other limitations worth naming - **Storage drivers.** `overlay2` works unprivileged on modern kernels; older ones need `fuse-overlayfs`, and the fallback `vfs` copies whole layers, exploding disk usage and build time. - **No sharing.** Each user has an independent daemon, image store and build cache, so on a multi-user box you pay N times for the same base images. - **Lifecycle.** The daemon is a systemd *user* unit; without `loginctl enable-linger <user>` it stops when the last session ends, killing long-running containers. - **Feature gaps.** Swarm overlay networking, checkpoint/restore, arbitrary device passthrough, most host-level mounts (NFS/CIFS) and AppArmor profile loading are unavailable or restricted; `ping` needs `net.ipv4.ping_group_range` widened. ## How to present it The strong answer separates *hard kernel constraints* (unprivileged ports, unprivileged mounts, cgroup ownership) from *implementation costs* (userspace networking, fuse overlays) and *operational setup* (delegation, lingering). That framing shows you understand why each limitation exists rather than reciting a compatibility list, and it makes each workaround obvious.

  • Why do published-port connections show a source address of 127.0.0.1 in a rootless setup, and what is the cost of fixing it?
    RootlessKit's default built-in port driver proxies connections from the host namespace into the container namespace, so the container sees the proxy's local address rather than the client's. Switching to the slirp4netns port driver preserves the original source address but routes traffic through the userspace stack, reducing throughput and raising CPU. Pick per workload: real client IPs where rate limiting or audit logging depends on them, raw speed elsewhere.
  • Rootless builds on an older host are extremely slow and consume huge amounts of disk. What is the likely cause?
    The storage driver has fallen back to vfs because neither unprivileged overlay2 nor fuse-overlayfs is available. vfs performs a full copy of every layer instead of stacking them, so images consume the sum of all layers and builds pay repeated copies. Install fuse-overlayfs or move to a kernel supporting unprivileged overlay mounts, then confirm the reported driver in docker info.

saying these in an interview costs you the question

  • Blaming the application for a failed bind on port 80 instead of the unprivileged-port rule
  • Assuming --memory silently works; without cgroup v2 delegation the limit is simply not applied
  • Expecting rootless networking to match kernel bridge throughput
  • Forgetting enable-linger, then wondering why containers stop when the session ends
  • Believing every rootful feature works, including swarm overlay networks and device passthrough

context