skip to content

If a process inside a Linux container exploits a kernel vulnerability, how does the blast radius compare with the same exploit fired inside a hardware virtual machine — and which container-level controls actually shrink it?

level: seniorimportance: must knowfreq 56%

answer

  1. Container boundary = kernel; VM boundary = hypervisor
  2. One bug vs two-bug chain
  3. Privileged, docker.sock, host mounts = the real escapes
  4. Non-root + drop caps + userns + seccomp + MAC
  5. Untrusted code → second boundary, per-tenant nodes

basics

~20 s

In a container the kernel is the isolation boundary, so a kernel privilege-escalation bug means host compromise and every co-tenant container with it. In a VM the attacker owns only that guest and must also break the hypervisor. Shrink it by cutting syscall and capability reach.

solid answer

~50 s

A container's boundary *is* the host kernel. Root inside the kernel means root on the host: all containers, their secrets, the container runtime, the node's credentials. In a VM the same exploit yields only that guest's kernel; escaping further requires a separate hypervisor bug, and the hypervisor's attack surface is far narrower than ~350 Linux syscalls plus ioctls, filesystems, netfilter and eBPF. Controls that genuinely reduce reach: - **Run as non-root** and set `no-new-privileges` — most escalation chains start from an in-container root or a setuid binary. - **Drop capabilities** to the minimum; never `--privileged`, never `CAP_SYS_ADMIN` casually. - **User namespaces**, so in-container root maps to an unprivileged host UID. - **Seccomp and AppArmor/SELinux**, to cut the reachable syscall surface. - **No host mounts** of the Docker socket, `/proc`, `/sys` or the host root. - **Patch the host kernel promptly** — it is the shared boundary. When the workload is genuinely untrusted, add a second boundary: a micro-VM or a user-space kernel.

code

bash · 11 lines
bash
docker run -d \
  --user 10001:10001 \
  --read-only \
  --cap-drop=ALL --cap-add=NET_BIND_SERVICE \
  --security-opt=no-new-privileges \
  --security-opt seccomp=/etc/docker/seccomp/app.json \
  --pids-limit=200 \
  myapp:1.4

# never in production:
# docker run --privileged -v /var/run/docker.sock:/var/run/docker.sock ...

go deeper

for a junior

Know the core fact — containers share the host kernel, so a kernel exploit can reach the host — and name two mitigations such as running as non-root and never using privileged mode.

for a middle

Explain the specific misconfigurations that cause real escapes (privileged, socket mounts, host paths, excess capabilities) and what seccomp and capability drops each remove.

for a senior

Reason about the boundary explicitly, order controls by return, own host-kernel patching and node rotation, and say when a hardened container is simply the wrong tool.

for a principal

Set the tenancy policy: which workloads get a second boundary, per-tenant node pools, blast-radius budgets, patch SLAs for the shared kernel, and the cost of the sandbox you choose.

## Why the boundary's identity decides the blast radius Security reasoning here is simple once you name the boundary. - **Container:** the boundary is the host kernel. Everything the container can do reaches the kernel directly — hundreds of syscalls, plus `ioctl` surfaces, filesystem parsers, the network stack, eBPF, and driver code. A kernel privilege-escalation bug converts container-local code execution into **host root**, and from there into every other container, their mounted secrets, the runtime socket and the node's cloud credentials. - **VM:** the boundary is the hypervisor. The same exploit gives the attacker root in *that guest's* kernel only. To reach the host they must additionally find a hypervisor bug — in virtual device emulation, typically. That surface is deliberately small and hardened, so the attack becomes a two-bug chain instead of one. That is the whole comparison: **one bug versus two, and one tenant versus all tenants.** ## What an attacker actually reaches for Realistic escapes rarely start with an exotic kernel zero-day. They start with configuration: - **`--privileged`** — effectively disables the container's restrictions; trivially escapable. - **The mounted Docker/containerd socket** — the attacker just asks the daemon for a new privileged container with the host root mounted. Not an escape, a documented API call. - **Host path mounts** — `/`, `/proc`, `/sys`, `/var/run`, or a device node. - **Dangerous capabilities** — `CAP_SYS_ADMIN`, `CAP_SYS_MODULE` (load a kernel module), `CAP_SYS_PTRACE`, `CAP_DAC_READ_SEARCH`. - **In-container root plus a kernel bug** — the classic chain; several publicised container escapes were kernel bugs reachable only with capabilities the default profile would have blocked. So the first-order hardening is not exotic sandboxes; it is not handing the boundary away. ## Controls, and what each really buys **Run as a non-root user** (`USER` in the image, `runAsNonRoot` in the platform). Removes the easiest step of most chains and blocks setuid tricks. Cheap, huge return. **`no-new-privileges`.** Stops a process gaining privileges through setuid/file capabilities after exec. **Drop capabilities.** Default runtimes already drop many; drop the rest and add back only what is needed (`NET_BIND_SERVICE` is usually all a server wants). Never grant `SYS_ADMIN` as a shortcut. **User namespaces.** Map in-container UID 0 to an unprivileged host UID, so even in-container root is unprivileged if it reaches the host boundary. Cost: complexity with volume ownership and some drivers. **Seccomp.** Filter syscalls. The default Docker profile already blocks well over 40, including whole classes of historical exploit primitives. A tailored profile is better; `--security-opt seccomp=unconfined` is the anti-pattern. **AppArmor / SELinux.** Mandatory access control on files, mounts, capabilities and ptrace — defence in depth when a syscall filter is not enough. **Read-only root filesystem, no host mounts, no daemon socket.** Removes the boring escapes entirely. **Patch the host kernel.** Because the boundary is shared, one unpatched node exposes every tenant on it. Live patching or fast node rotation is a real control, not hygiene. ## What containers do *not* fully isolate Even when configured well: many `sysctl`s are not namespaced; `/proc` and `/sys` expose host detail; the page cache, scheduler and memory bandwidth are shared, enabling noisy-neighbour effects and side channels; kernel log and audit surfaces are host-wide. VMs, by contrast, isolate essentially all of that except hardware-level microarchitectural channels. ## When the answer is a second boundary If you run genuinely hostile code — customer-submitted builds, CI for public pull requests, function-as-a-service, a code sandbox — configuration hardening is not enough on its own. Put a hypervisor or a user-space kernel underneath: micro-VMs give each workload its own kernel with a tiny device model, and user-space kernels service most syscalls outside the host kernel so far fewer host syscalls are reachable. In practice large platforms combine both: hardened containers *inside* per-tenant micro-VMs, plus per-tenant node pools so a compromise cannot cross customers. ## How to say it in an interview "Containers share the kernel, so a kernel privesc is a host and multi-tenant compromise, not a single-guest one; a VM needs a second, much rarer hypervisor bug. I reduce reach with non-root, dropped capabilities, user namespaces, seccomp and MAC, no privileged mode and no host or socket mounts, and fast host-kernel patching. For untrusted code I add a real second boundary — micro-VM or user-space kernel — and I separate tenants at the node level."

  • Why is mounting the Docker socket into a container considered equivalent to giving away host root?
    The socket is the daemon's full control API. Anything that can talk to it can create a container that is privileged and mounts the host root filesystem, then read or write anything on the node — no exploit needed. Use a rootless or API-mediated builder instead, or grant a narrowly scoped build service rather than the raw socket.
  • If seccomp and dropped capabilities are in place, is a kernel exploit still a host-level risk?
    Yes, just a less likely one. Those controls shrink the reachable surface but the workload still calls the same kernel, and a bug in a syscall you legitimately allow — or in the network stack, a filesystem parser, or an ioctl path — remains exploitable. They are defence in depth; the boundary is unchanged, which is why untrusted workloads get a hypervisor underneath.

saying these in an interview costs you the question

  • Treating a container as a security boundary equivalent to a VM
  • Using --privileged or seccomp=unconfined to make something work and leaving it
  • Mounting the container runtime socket into workloads for convenience
  • Assuming in-container root is harmless without user namespaces
  • Believing a memory or CPU limit is a security control against escapes

context