skip to content

What kind of vulnerability were the runc escapes CVE-2019-5736 and CVE-2024-21626, and what did they let an attacker do?

level: seniorimportance: nice to knowfreq 28%

answer

  1. The escape is in the runtime, not config
  2. runc runs as root on the host
  3. 2019: overwrite runc via /proc/self/exe
  4. 2024: leaked fd plus WORKDIR into host
  5. fix is patching the engine

basics

~20 s

Both are escapes in runc, the low-level binary Docker uses to start containers. CVE-2019-5736 let a malicious container overwrite the host runc binary via /proc/self/exe and gain root on the next exec. CVE-2024-21626 (Leaky Vessels) used a leaked host file descriptor plus a crafted WORKDIR to break out to the host filesystem, at build or run time.

solid answer

~50 s

runc is the OCI runtime Docker (through containerd) invokes to actually spawn a container -- it is the trust boundary, and a bug in it escapes regardless of how the container is configured. **CVE-2019-5736**: runc re-executes itself via `/proc/self/exe`; a malicious container could open that path (the host runc binary) and overwrite it, so the next `docker run`/`exec` executed the attacker's code as root on the host. The fix copies runc into a sealed memfd. **CVE-2024-21626**: an internal file descriptor to a host directory (under /sys/fs/cgroup) leaked into the container; a `WORKDIR` (or `process.cwd`) of `/proc/self/fd/<n>` resolved into the host filesystem, letting the container read or write host files -- exploitable even during `docker build`. Crucially, non-root, capability-dropped, seccomp-confined containers were still vulnerable, because the hole is in the runtime. The mitigation is patching runc/containerd.

code

dockerfile · 5 lines
dockerfile
FROM alpine
# On vulnerable runc, a leaked host directory fd is reachable via /proc/self/fd/<n>;
# a WORKDIR pointing at it resolves into the HOST filesystem, not the container's.
WORKDIR /proc/self/fd/7
RUN cat ../../../../etc/shadow > /breakout || true

go deeper

for a junior

Recall that runc is the low-level program Docker uses to start containers and that keeping Docker updated matters for security -- the exact CVE numbers are not expected.

for a middle

Explain that runc runs as root on the host and is the trust boundary, so a bug in it escapes regardless of the container's user or capabilities; give the gist of the /proc/self/exe and leaked-fd mechanisms.

for a senior

Describe both mechanisms accurately, note that CVE-2024-21626 was exploitable at build time, and argue that patching runc/containerd -- not config hardening -- is the fix, with the version boundaries.

for a principal

Own the implication for platform strategy: keep the runtime current as a security control, isolate untrusted build/run so a future runtime escape has a small blast radius, and do not let config hardening create false confidence against runtime CVEs.

### Where runc sits When you `docker run`, the request flows dockerd -> containerd -> a `containerd-shim` -> **runc**, the small OCI-spec runtime that sets up namespaces, cgroups and mounts and then `exec`s your process. runc runs as **root on the host** for that instant. That makes it the container trust boundary: if the container can subvert runc, it subverts the host, and *no amount of container configuration* -- non-root user, dropped capabilities, seccomp, read-only rootfs -- can defend a flaw inside the runtime itself. That is the theme uniting these two CVEs. ### CVE-2019-5736 -- overwriting runc via /proc/self/exe To enter a container, runc re-executes itself using the magic symlink `/proc/self/exe`, which points at the running runc binary **on the host**. The flaw: a malicious image's entrypoint (or a process the attacker controlled when an operator ran `docker exec`) could open `/proc/self/exe` for writing and overwrite the host's runc binary. The next time anyone started or exec'd into a container, the tampered runc ran the attacker's code as root on the host. A single malicious image that an admin merely `exec`'d into could take the machine. The fix has runc copy itself into a read-only, sealed `memfd` before re-executing, so the on-disk binary can no longer be clobbered through `/proc/self/exe`. ### CVE-2024-21626 -- Leaky Vessels and the leaked fd In vulnerable runc versions, an internal file descriptor referring to a **host** directory (in the `/sys/fs/cgroup` area) was left open -- leaked -- into the container's process. If the container's working directory was set to `/proc/self/fd/<n>` for that leaked descriptor, the path resolved into the **host** filesystem rather than the container's, giving the process read/write access to host files outside the container. Because `WORKDIR` in a Dockerfile is honoured by runc, this was exploitable **at build time** too: merely building a malicious Dockerfile (or one `FROM` a malicious base) with a crafted `WORKDIR` could read or write host files on an unpatched engine -- not only running an untrusted image. It affected other consumers of the same runc (BuildKit, buildah, containerd). Fixed in runc 1.1.12 by closing the leaked descriptors. ### Why config hardening did not help The uncomfortable lesson: containers running as an unprivileged UID, with `--cap-drop ALL`, the default seccomp profile and a read-only rootfs were **still vulnerable** to both. All those controls constrain the *containerised process*; they do nothing about a defect in the *runtime that launches it*. A Python ML inference image built from `FROM python:slim` and pinned to a non-root user is not protected by any of that if the host's runc is unpatched. So runtime CVEs sit alongside misconfiguration as a distinct escape surface. ### The defence 1. **Keep runc/containerd/Docker patched.** These fixes ship in the engine, not your image -- CVE-2019-5736 in Docker 18.09.2, CVE-2024-21626 in Docker 25.0.2 (runc 1.1.12). This is why 'keep the engine current' is a security control, not just hygiene. 2. **Do not build or run untrusted images/Dockerfiles on shared hosts**, especially unpatched ones -- CVE-2024-21626 made even `docker build` of an untrusted Dockerfile dangerous. 3. **Isolate build/run of low-trust content** so a runtime escape, if one recurs, lands somewhere with a small blast radius. Knowing the exact CVE numbers is differentiator-level trivia; the enduring point a strong candidate carries is that the runtime is a trust boundary you patch, because container configuration cannot save you from a hole in runc.

  • Why did non-root, capability-dropped containers still fall to CVE-2019-5736?
    Because the flaw is in runc, the host binary that enters the container. Overwriting /proc/self/exe corrupts the host's own runtime regardless of the container's UID, capabilities, or seccomp profile -- those controls constrain the container process, not the runtime that launches it. The fix was runtime code (copying runc into a sealed memfd), not a container setting.
  • How was CVE-2024-21626 a build-time risk, not just a runtime one?
    The same runc executes container steps during docker build, and it honours WORKDIR. A malicious base image or Dockerfile whose WORKDIR pointed at the leaked host-directory descriptor (/proc/self/fd/<n>) could read or write host files while the image was merely being built. So on an unpatched engine, even building an untrusted Dockerfile -- not just running an untrusted image -- was dangerous.

saying these in an interview costs you the question

  • Thinks dropping capabilities or running non-root would have stopped these
  • Believes only --privileged containers were affected
  • Assumes the fix is a container config change, not patching runc
  • Thinks CVE-2024-21626 was exploitable only at runtime, not during build
  • Confuses a runc runtime escape with a kernel namespace bug

context