skip to content

A container with strict CPU and memory ceilings triggers a fault in the host's shared kernel — what happens to its neighbours?

level: middleimportance: must knowfreq 58%

answer

  1. what a ceiling is actually accounting for
  2. count the kernels on this machine
  3. limits cap consumption, not correctness
  4. one shared path, every container
  5. the failure domain is the machine

basics

~20 s

They stop too. A ceiling bounds how much of a resource a workload consumes; it does not bound the kernel running on that workload's behalf, and one kernel serves every container on the machine, so a panic takes all of them down together.

solid answer

~50 s

A ceiling is an accounting device. It caps how much CPU time or memory a workload gets: exceed the CPU ceiling and the process is throttled and keeps running, exceed the memory ceiling and it is killed. Neither says anything about the correctness of the kernel code executing on that workload's behalf. Every container on the host goes through the same kernel, so a fault in a shared path — a device driver, a filesystem path, a network path — is a whole-machine event, and a panic ends every container regardless of how tightly each one was capped. The same shape holds for security: a flaw reachable at the call boundary is reachable from any workload on that machine. The failure domain is the host, which is why replicas of one workload are spread across hosts rather than across one.

go deeper

for a junior

Remember that the containers on one machine share one kernel, so anything that stops the kernel stops all of them at once, whatever each one's limits say.

for a middle

Explain the difference between bounding consumption and bounding blast radius: a ceiling throttles or ends one workload's processes, while a fault in shared kernel code is a machine-level event.

for a senior

Demonstrate the operational consequence — replicas spread across machines, the host treated as the unit taken out of service, and simultaneous unrelated failures on one machine read as a host-level signal rather than a workload bug.

for a principal

Argue about what the shared boundary is worth: which workloads may share a machine at all, how much of the estate one kernel fault may take with it, and what second walls you require beneath the first.

## What a ceiling actually bounds A per-container CPU or memory ceiling is enforced by the kernel's accounting and limiting mechanism on the processes inside the boundary. It answers exactly one question: *how much of this resource may these processes consume?* The two answers differ in kind, and the difference is worth stating precisely because it is so often reversed: - **A CPU ceiling squeezes.** The workload is given a time quota; once it is used up the processes wait. They run slower and stay alive, and the visible symptom is latency, not death. - **A memory ceiling kills.** There is no throttle for memory — when the workload cannot be kept under its ceiling, a process is ended. What neither one bounds is **what the kernel does while serving that workload**. Consumption and correctness are different axes, and no ceiling has any purchase on the second one. ## Why one kernel means one failure domain Every container on a host issues its calls into the same kernel instance. That kernel is not partitioned by the boundary; the boundary partitions the *views* the kernel hands out — filesystem, process table, network, ids, hostname — and accounts for consumption. The code itself is shared. So when a shared path faults — a device driver, a filesystem path, a networking path — the machine is what fails. A panic stops the machine; every container on it stops with it, in the same instant, with no regard for its limits, its priority or how well behaved it was. A hang in a shared path produces the same shape more slowly: unrelated tenants degrade together. The security shape is the same shape. A flaw at the call boundary is reachable by any workload that can make calls, and a workload that exploits it is not constrained by how small its ceilings were. That is precisely why platforms add a second wall over the call surface — a restricted set of permitted calls, or a mandatory access-control profile — rather than relying on the resource boundary for containment it was never built to give. ## Three risks that are easy to conflate | Failure | What it reaches | What actually limits it | |---|---|---| | A workload exceeding its own memory ceiling | that workload's processes | the ceiling itself, working as designed | | A fault in shared kernel code on one machine | every container on that machine | not co-locating: fewer tenants per machine, replicas on separate machines | | A defect in the kernel version the fleet runs | every machine running that version | changing the version, and narrowing what workloads may ask of the kernel | The second and third rows are both *correlated* risks, but they correlate on different things, and the mitigation for one does nothing for the other. Spreading replicas across machines is a real defence against the second and no defence at all against the third. ## Reading the pattern in production Because the failure domain is the machine, the tell is **correlation by host**: - unrelated workloads, owned by different teams, degrading at the same moment; - the same workloads healthy on other machines at the same time; - symptoms that cross the usual explanations — network, storage and scheduling all sour together. That pattern points at something the tenants share, not at code any of them wrote. It is also the practical reason to record which host served a failing request: without that field, the correlation is invisible, and the incident gets debugged as a workload problem for hours. ## What follows for design 1. **Spread replicas of one workload across separate machines.** Copies on one machine are one panic away from zero copies. 2. **Treat the host as the unit that fails and the unit you take out of service.** That is why maintenance, patching and hardware work are all expressed as emptying a machine and restarting it. 3. **Decide co-location deliberately.** The number of tenants on a machine is the size of the blast radius when its kernel fails, and for some workloads the honest answer is fewer neighbours — or a boundary that does not share this kernel at all, which is a separate decision with its own costs. 4. **Do not sell ceilings as containment.** They are capacity controls. Saying that a tightly capped workload cannot hurt its neighbours is true about consumption and false about everything else.

  • Several unrelated workloads on one machine degrade at the same moment while the same workloads elsewhere stay healthy. What does that pattern suggest?
    That the common factor is the machine, not the workload. Unrelated tenants failing together at one instant points at something they share — the kernel, the hardware, a local device — rather than at code each of them runs independently. It is the practical reason to record which host served a failing request; without it the correlation is invisible and the incident is debugged in the wrong place.
  • Replicas are spread across machines, but every machine runs the same kernel version. Which risk does spreading remove, and which does it not?
    Spreading removes the per-machine risks: a panic, a driver fault or failing hardware ends one machine and its containers, and the surviving replicas carry on elsewhere. It does nothing about a defect in the kernel version itself, because every machine runs the same code — that risk is correlated across the entire fleet and is retired by changing the version, not by placement.

saying these in an interview costs you the question

  • Believes tight resource ceilings contain a kernel-level fault.
  • Treats each container as its own failure domain.
  • Says a panic ends only the container that triggered it.
  • Thinks a call-boundary flaw only matters for privileged containers.
  • Spreads replicas across containers on one machine and calls that redundancy.