If a process inside a Linux container exploits a kernel vulnerability, how does the blast radius compare with the same exploit fired inside a hardware virtual machine — and which container-level controls actually shrink it?
answer
- Container boundary = kernel; VM boundary = hypervisor
- One bug vs two-bug chain
- Privileged, docker.sock, host mounts = the real escapes
- Non-root + drop caps + userns + seccomp + MAC
- Untrusted code → second boundary, per-tenant nodes
basics
~20 sIn a container the kernel is the isolation boundary, so a kernel privilege-escalation bug means host compromise and every co-tenant container with it. In a VM the attacker owns only that guest and must also break the hypervisor. Shrink it by cutting syscall and capability reach.
solid answer
~50 sA container's boundary *is* the host kernel. Root inside the kernel means root on the host: all containers, their secrets, the container runtime, the node's credentials. In a VM the same exploit yields only that guest's kernel; escaping further requires a separate hypervisor bug, and the hypervisor's attack surface is far narrower than ~350 Linux syscalls plus ioctls, filesystems, netfilter and eBPF. Controls that genuinely reduce reach: - **Run as non-root** and set `no-new-privileges` — most escalation chains start from an in-container root or a setuid binary. - **Drop capabilities** to the minimum; never `--privileged`, never `CAP_SYS_ADMIN` casually. - **User namespaces**, so in-container root maps to an unprivileged host UID. - **Seccomp and AppArmor/SELinux**, to cut the reachable syscall surface. - **No host mounts** of the Docker socket, `/proc`, `/sys` or the host root. - **Patch the host kernel promptly** — it is the shared boundary. When the workload is genuinely untrusted, add a second boundary: a micro-VM or a user-space kernel.
code
bash · 11 linesdocker run -d \
--user 10001:10001 \
--read-only \
--cap-drop=ALL --cap-add=NET_BIND_SERVICE \
--security-opt=no-new-privileges \
--security-opt seccomp=/etc/docker/seccomp/app.json \
--pids-limit=200 \
myapp:1.4
# never in production:
# docker run --privileged -v /var/run/docker.sock:/var/run/docker.sock ...go deeper
Know the core fact — containers share the host kernel, so a kernel exploit can reach the host — and name two mitigations such as running as non-root and never using privileged mode.
Explain the specific misconfigurations that cause real escapes (privileged, socket mounts, host paths, excess capabilities) and what seccomp and capability drops each remove.
Reason about the boundary explicitly, order controls by return, own host-kernel patching and node rotation, and say when a hardened container is simply the wrong tool.
Set the tenancy policy: which workloads get a second boundary, per-tenant node pools, blast-radius budgets, patch SLAs for the shared kernel, and the cost of the sandbox you choose.
## Why the boundary's identity decides the blast radius Security reasoning here is simple once you name the boundary. - **Container:** the boundary is the host kernel. Everything the container can do reaches the kernel directly — hundreds of syscalls, plus `ioctl` surfaces, filesystem parsers, the network stack, eBPF, and driver code. A kernel privilege-escalation bug converts container-local code execution into **host root**, and from there into every other container, their mounted secrets, the runtime socket and the node's cloud credentials. - **VM:** the boundary is the hypervisor. The same exploit gives the attacker root in *that guest's* kernel only. To reach the host they must additionally find a hypervisor bug — in virtual device emulation, typically. That surface is deliberately small and hardened, so the attack becomes a two-bug chain instead of one. That is the whole comparison: **one bug versus two, and one tenant versus all tenants.** ## What an attacker actually reaches for Realistic escapes rarely start with an exotic kernel zero-day. They start with configuration: - **`--privileged`** — effectively disables the container's restrictions; trivially escapable. - **The mounted Docker/containerd socket** — the attacker just asks the daemon for a new privileged container with the host root mounted. Not an escape, a documented API call. - **Host path mounts** — `/`, `/proc`, `/sys`, `/var/run`, or a device node. - **Dangerous capabilities** — `CAP_SYS_ADMIN`, `CAP_SYS_MODULE` (load a kernel module), `CAP_SYS_PTRACE`, `CAP_DAC_READ_SEARCH`. - **In-container root plus a kernel bug** — the classic chain; several publicised container escapes were kernel bugs reachable only with capabilities the default profile would have blocked. So the first-order hardening is not exotic sandboxes; it is not handing the boundary away. ## Controls, and what each really buys **Run as a non-root user** (`USER` in the image, `runAsNonRoot` in the platform). Removes the easiest step of most chains and blocks setuid tricks. Cheap, huge return. **`no-new-privileges`.** Stops a process gaining privileges through setuid/file capabilities after exec. **Drop capabilities.** Default runtimes already drop many; drop the rest and add back only what is needed (`NET_BIND_SERVICE` is usually all a server wants). Never grant `SYS_ADMIN` as a shortcut. **User namespaces.** Map in-container UID 0 to an unprivileged host UID, so even in-container root is unprivileged if it reaches the host boundary. Cost: complexity with volume ownership and some drivers. **Seccomp.** Filter syscalls. The default Docker profile already blocks well over 40, including whole classes of historical exploit primitives. A tailored profile is better; `--security-opt seccomp=unconfined` is the anti-pattern. **AppArmor / SELinux.** Mandatory access control on files, mounts, capabilities and ptrace — defence in depth when a syscall filter is not enough. **Read-only root filesystem, no host mounts, no daemon socket.** Removes the boring escapes entirely. **Patch the host kernel.** Because the boundary is shared, one unpatched node exposes every tenant on it. Live patching or fast node rotation is a real control, not hygiene. ## What containers do *not* fully isolate Even when configured well: many `sysctl`s are not namespaced; `/proc` and `/sys` expose host detail; the page cache, scheduler and memory bandwidth are shared, enabling noisy-neighbour effects and side channels; kernel log and audit surfaces are host-wide. VMs, by contrast, isolate essentially all of that except hardware-level microarchitectural channels. ## When the answer is a second boundary If you run genuinely hostile code — customer-submitted builds, CI for public pull requests, function-as-a-service, a code sandbox — configuration hardening is not enough on its own. Put a hypervisor or a user-space kernel underneath: micro-VMs give each workload its own kernel with a tiny device model, and user-space kernels service most syscalls outside the host kernel so far fewer host syscalls are reachable. In practice large platforms combine both: hardened containers *inside* per-tenant micro-VMs, plus per-tenant node pools so a compromise cannot cross customers. ## How to say it in an interview "Containers share the kernel, so a kernel privesc is a host and multi-tenant compromise, not a single-guest one; a VM needs a second, much rarer hypervisor bug. I reduce reach with non-root, dropped capabilities, user namespaces, seccomp and MAC, no privileged mode and no host or socket mounts, and fast host-kernel patching. For untrusted code I add a real second boundary — micro-VM or user-space kernel — and I separate tenants at the node level."
- Why is mounting the Docker socket into a container considered equivalent to giving away host root?The socket is the daemon's full control API. Anything that can talk to it can create a container that is privileged and mounts the host root filesystem, then read or write anything on the node — no exploit needed. Use a rootless or API-mediated builder instead, or grant a narrowly scoped build service rather than the raw socket.
- If seccomp and dropped capabilities are in place, is a kernel exploit still a host-level risk?Yes, just a less likely one. Those controls shrink the reachable surface but the workload still calls the same kernel, and a bug in a syscall you legitimately allow — or in the network stack, a filesystem parser, or an ioctl path — remains exploitable. They are defence in depth; the boundary is unchanged, which is why untrusted workloads get a hypervisor underneath.
saying these in an interview costs you the question
- Treating a container as a security boundary equivalent to a VM
- Using --privileged or seccomp=unconfined to make something work and leaving it
- Mounting the container runtime socket into workloads for convenience
- Assuming in-container root is harmless without user namespaces
- Believing a memory or CPU limit is a security control against escapes