skip to content

How can a privileged Docker container with CAP_SYS_ADMIN break out to the host using the cgroup v1 release_agent?

level: seniorimportance: should knowfreq 40%

answer

  1. An empty cgroup runs a program
  2. release_agent runs on the host as root
  3. needs CAP_SYS_ADMIN to mount cgroupfs
  4. overlay upperdir bridges the paths
  5. cgroup v2 has no release_agent

basics

~20 s

With CAP_SYS_ADMIN (which --privileged grants) on a cgroup v1 host, the container mounts a cgroup hierarchy, sets its release_agent to a script on a host-visible path, and empties a child cgroup. The kernel then runs that script as root in the host namespace -- a full breakout, needing no CVE.

solid answer

~40 s

cgroup v1 has a `release_agent` file: when the last process leaves a cgroup and `notify_on_release` is set, the kernel runs the program named in `release_agent` **on the host, as root**. A `--privileged` container has `CAP_SYS_ADMIN` and is unconfined, so it can `mount -t cgroup` a fresh hierarchy inside itself. It writes a host-resolvable path -- the overlay upperdir path read from `/etc/mtab` -- into `release_agent`, drops an executable payload there, and triggers release by putting a process into a child cgroup and letting it exit. The kernel executes the payload in the host root namespace. It needs no kernel bug -- it abuses intended cgroup v1 features reachable only because the container was handed `CAP_SYS_ADMIN` and left unconfined. Dropping `CAP_SYS_ADMIN`, blocking `mount()` with seccomp/AppArmor, or running on cgroup v2 all defeat it.

code

bash · 8 lines
bash
mkdir /tmp/cgrp && mount -t cgroup -o rdma cgroup /tmp/cgrp && mkdir /tmp/cgrp/x
echo 1 > /tmp/cgrp/x/notify_on_release
host_path=$(sed -n 's/.*\upperdir=\([^,]*\).*/\1/p' /etc/mtab)
echo "$host_path/cmd" > /tmp/cgrp/release_agent
printf '#!/bin/sh\nps aux > %s/output' "$host_path" > /cmd
chmod a+x /cmd
sh -c "echo \$\$ > /tmp/cgrp/x/cgroup.procs"
cat /output

go deeper

for a junior

Recall that a --privileged container can break out to the host, and that 'privileged equals root on the host' is literally true -- you are not expected to reproduce the technique.

for a middle

Explain the ingredients: cgroup v1's release_agent runs a program on the host as root when a cgroup empties, and CAP_SYS_ADMIN lets the container mount the cgroupfs needed to set it.

for a senior

Walk the chain end to end -- mount cgroupfs, resolve the overlay upperdir, set release_agent, trigger release -- and state that it is a config/privilege abuse, not a CVE, plus the exact changes that break it.

for a principal

Frame it as evidence for policy: because unconfined CAP_SYS_ADMIN is trivially root-on-host, argue for cgroup v2, default-denied privilege, and treating any --privileged workload as sharing the host's trust boundary.

### The feature being abused In **cgroup v1**, each cgroup has two relevant files. `notify_on_release` (0 or 1) says whether the kernel should act when the cgroup becomes empty. `release_agent`, set at the root of a hierarchy, names a program the kernel runs -- **on the host, in the initial namespaces, as root** -- whenever a cgroup in that hierarchy with `notify_on_release=1` loses its last process. It was designed for cleanup automation. It is also a straight line from inside a container to root on the host. ### Why a privileged container can reach it Two things are needed. First, **CAP_SYS_ADMIN**, because mounting a cgroup filesystem is a privileged mount operation -- `--privileged` grants all capabilities including this one. Second, the container must be **unconfined**: the default seccomp profile blocks the `mount()` syscall and the default AppArmor profile restricts mounts, but `--privileged` disables both. So a privileged container can mount a brand-new cgroup v1 hierarchy inside itself and write to its `release_agent`. ### The path trick The `release_agent` program runs from the **host's** filesystem view, not the container's, so the attacker must name a path that resolves both inside and outside the container. With the overlay2 storage driver, the container's writable layer exists on the host at the overlay `upperdir`. That path is visible in the container's own mount table (`/etc/mtab` / `/proc/self/mountinfo`). So the attacker reads the `upperdir`, writes the payload script into the container (which lands under that `upperdir` on the host), and points `release_agent` at the host-side path. ### The canonical proof of concept ```bash # inside a --privileged container on a cgroup v1 host mkdir /tmp/cgrp && mount -t cgroup -o rdma cgroup /tmp/cgrp && mkdir /tmp/cgrp/x echo 1 > /tmp/cgrp/x/notify_on_release host_path=$(sed -n 's/.*\upperdir=\([^,]*\).*/\1/p' /etc/mtab) echo "$host_path/cmd" > /tmp/cgrp/release_agent printf '#!/bin/sh\nps aux > %s/output' "$host_path" > /cmd chmod a+x /cmd sh -c "echo \$\$ > /tmp/cgrp/x/cgroup.procs" # the exiting shell empties cgroup x; the kernel runs /cmd as root on the host cat /output # host 'ps aux' captured from inside the container ``` Swap the harmless `ps aux` for a reverse shell or an `authorized_keys` write and it is a full host takeover. The `-o rdma` controller is just a convenient one to mount; several controllers work. ### It is not a CVE The key point for an interview: there is **no kernel or Docker vulnerability here**. Every step uses intended behaviour. The escape exists purely because the container was given `CAP_SYS_ADMIN` and stripped of seccomp/AppArmor confinement -- i.e. because it was run `--privileged`. That is why '`--privileged` is root on the host' is not hyperbole. ### What closes it - **Don't run --privileged**; if a capability is truly needed, add only that one and never `CAP_SYS_ADMIN` casually. - **Keep CAP_SYS_ADMIN dropped** -- without it the initial `mount -t cgroup` fails and the chain never starts. - **Keep the default seccomp/AppArmor profiles** -- they block the `mount()` the technique depends on. - **Run on cgroup v2**, which has no `release_agent`; the release-notification mechanism was redesigned and this specific path is simply gone. Any one of these breaks the chain, which is why the technique is a demonstration of what `--privileged` costs rather than a bug to patch.

  • Which single change to the container's launch defeats this without moving to cgroup v2?
    Dropping CAP_SYS_ADMIN -- for example not using --privileged, or --cap-drop ALL and adding only the capabilities the app truly needs. Mounting a fresh cgroup hierarchy is a privileged mount that requires CAP_SYS_ADMIN, so without it the very first step fails. Keeping the default seccomp or AppArmor profile, which blocks mount(), stops it too.
  • Does this escape rely on a kernel or Docker vulnerability?
    No. Every step is intended behaviour: cgroup v1's release_agent is a designed feature that runs a host program when a cgroup empties. The escape is reachable only because the container was handed CAP_SYS_ADMIN and left unconfined by --privileged. It is a configuration and privilege problem, not a CVE to patch.

saying these in an interview costs you the question

  • Thinks the escape needs a kernel 0-day
  • Believes cgroup v2 hosts are still vulnerable to release_agent
  • Thinks the default seccomp profile permits mounting a cgroup hierarchy
  • Confuses release_agent (cgroup v1) with cgroup v2 release handling
  • Believes dropping CAP_SYS_ADMIN leaves the escape intact

context