Why can a workload spec with no privileged mode, no host mounts and no shared views still end in host compromise?
answer
- three routes declared, one not
- the boundary and the bug share code
- no grant required, no spec line
- patching and host replacement are controls
- clean spec is a floor, not a proof
basics
~20 sThree escape routes are grants written in the spec; the fourth is a defect in the shared kernel, reached by an ordinary call from inside the container. It needs no grant, appears in no spec, and a spec review cannot see it.
solid answer
~50 sEvery container on a host talks to the same kernel, so the kernel's call interface is permanently reachable from inside the boundary — it is not something the spec grants. A defect behind one of those calls is therefore an escape route that no review of the spec can find, and many such defects are reachable without any added privilege. What moves the exposure is how much of the call surface the workload can reach, how current the running kernel is, and how often hosts are replaced rather than left up for months. The practical consequence: a clean spec review removes every route you could have granted, which is genuinely worth doing, but it establishes a floor rather than proving isolation — and for workloads you actually do not trust, the answer is a different kind of boundary rather than a better spec.
go deeper
Remember that the kernel is shared by every container on a host, so a defect in it can be reached from inside a container that was granted nothing unusual at all.
Explain why this route never appears in a workload spec: the boundary is drawn by the same kernel the workload calls into, and calling the kernel is the normal case rather than a grant.
State the review's claim accurately — no route was granted, isolation was not proven — and name the real levers: the reachable call surface, the kernel the hosts run, and how often those hosts are replaced.
Decide where spec review stops being the right control for your estate, which workload classes get a boundary that does not rely on the shared kernel, and what host-replacement cadence you are willing to fund.
## Three routes you granted, one you did not The escape surface of a container has four routes, and they split cleanly in two. **Declared, and therefore reviewable:** 1. Running the container privileged — the boundary's restraints switched off. 2. Mounting a host path, or the runtime's control socket, into the container. 3. Sharing the host's process or network view. Each of those is a line in a workload spec. Somebody wrote it, somebody can read it, and somebody can delete it. **Undeclared:** 4. A defect in the kernel itself, reached through an ordinary call made from inside the container. No line in the spec asks for this one. It is available by construction, because the container is a process and a process talks to the kernel. ## Why the kernel is always reachable The boundary around a container is drawn *by* the kernel. The per-resource views, the trimmed privilege set, the call allowlist and the access-control profile are all features of the same kernel the workload is calling into. That means: - the code that enforces the boundary and the code that could break it are the same body of code; - every container on the host shares that one instance of it, so a defect is reachable from all of them; - a call being permitted is the normal case — workloads have to call the kernel to do anything at all. So this route does not require a misconfiguration, a careless reviewer or a generous grant. It requires a bug, and large shared codebases have bugs. ## What actually moves the exposure | Lever | Effect | |---|---| | How much of the call surface the workload can reach | Narrowing the permitted calls shrinks the set of defects reachable from inside — it reduces the route, it does not remove it. The mechanism for narrowing it is its own subject. | | How current the running kernel is | Fixed defects only stop mattering on hosts that are actually running the fix. | | How long a host stays up | A fleet whose hosts are replaced regularly runs newer code by default; one whose hosts have been up for a year is running whatever shipped a year ago. | | Who shares the host | The route's value to an attacker is whatever else is on that machine. | The second and third rows are why host patching and host replacement are isolation controls rather than housekeeping. They are the only levers that retire defects on this route. ## What a clean spec review is still worth A great deal, and the claim just has to be stated accurately: - **What it establishes:** no escape route was granted. A compromise now needs a kernel defect and the ability to reach it, rather than a checkbox somebody ticked. - **What it does not establish:** that the workload cannot reach the host. That is a different and much stronger claim, and no spec supports it. The difference is operationally large. Attacks that need only a granted route are cheap, scriptable and widely automated; attacks that need a live kernel defect are not. Removing the granted routes changes the population of attackers who can use your host as a stepping stone, which is exactly what hardening is for. ## Where the answer stops being "tighten the spec" If the workload is genuinely untrusted — somebody else's code, arbitrary submitted jobs, a hostile multi-tenant mix — then tightening the spec runs out of road, because the remaining route is the one the spec does not describe. At that point the decision is about what kind of boundary the workload gets, which is a design question about isolation rather than a review question about grants. Recognising that the spec review has hit its limit is the senior part of this answer; choosing the replacement boundary belongs to the isolation discussion. ## How to say it "Three of the four ways out are things we grant, and we granted none of them — that is a real improvement and I would insist on it. The fourth is a defect in the kernel we all share, it is reachable from any container without our help, and the only levers we have on it are how much of the kernel this workload can reach and how quickly these hosts get replaced. So this spec is clean; it is not a proof of isolation, and if this workload were untrusted I would not be arguing about the spec at all."
- If this route cannot be removed, what does the spec review still buy?It removes every route that could have been granted, so a compromise now requires a kernel defect and a way to reach it instead of a ticked box. That changes who can attack you: granted-route attacks are cheap and automated, while a working kernel escape is scarce. The review changes the price, not the possibility.
- What makes one host's exposure to this route larger than another's?How much of the kernel's call surface its workloads may reach, how current the kernel it is running is, and how long it has been up — a host replaced weekly runs recent code by default, one that has been up for a year runs whatever shipped then. What else runs on the machine sets what the escape is worth.
- Does this mean container boundaries are not worth hardening?The opposite. Most real incidents use a granted route, because those are free and this one is not. Hardening removes the cheap routes and leaves the expensive one, which is the normal shape of a security control. The mistake is only in the claim afterwards — say the routes were removed, not that isolation is proven.
saying these in an interview costs you the question
- Says a spec with no privileged mode and no host mounts proves the workload is isolated
- Assumes an escape always requires a misconfiguration somewhere
- Thinks this route needs the workload to hold the host's full privilege set first
- Treats host patching and host replacement as maintenance rather than isolation controls
- Believes narrowing the reachable call surface removes the route instead of shrinking it