Why can a container be killed for running out of memory while its own usage stays under its declared memory ceiling?
answer
- two shortages, not one
- your ceiling is not the node's
- host-scope kill ignores your compliance
- victim chosen for relief, not blame
- footprint outweighs growth rate
basics
~20 sTwo different shortages end in a kill. A ceiling is enforced against one container's own usage; the host's total memory is a separate constraint. When the host runs out, the kernel picks a victim by footprint, and your ceiling grants no immunity.
solid answer
~50 sThere are two independent memory shortages on any node. The first is the ceiling declared for a container: the runtime enforces it against that container's own accounting, and crossing it kills a process in that container regardless of how much memory the machine has. The second is the machine itself running out, which is nobody's ceiling breach — it is the sum of every workload plus the node's own agents. When that happens the kernel's out-of-memory killer selects a victim at host scope, scored mainly by how much memory stopping it would free. A steady, well-behaved workload with a large resident set is therefore an attractive victim, while the workload whose growth caused the shortage may still be small at the instant of the decision and survive. Staying under your ceiling is necessary, not protective.
go deeper
Remember there are two separate memory limits in play: the one declared for your container, and the machine's own total. Being under the first does not protect you from the second.
Explain the mechanics: the runtime enforces the declared ceiling against that container's accounting, while host exhaustion is resolved at machine scope by the kernel scoring candidates for how much memory stopping them frees.
Show how you tell the two apart from evidence — several unrelated workloads stopping together versus one container dying at a repeatable figure — and say which team the investigation belongs to in each case.
Frame it as who gets to choose the victim. Whether the kernel or the platform makes that call is a property of how the estate is configured, and unattributed kills are a signal about the node, not the service.
## Two shortages, not one Every node carries two memory constraints that are enforced separately, and confusing them is why this symptom looks impossible. The first is the **ceiling** declared for a container in its workload spec. The runtime enforces it through the kernel's resource-accounting mechanism, against that one container's own usage. Cross it and a process inside that container is killed — even if the machine has gigabytes free. The second is the **machine's physical memory**, shared by every container on it plus the node agent, the runtime itself and whatever else the operating system is holding. Exhausting that is not a breach of anyone's declared figure. It is an arithmetic outcome: the workloads placed there were allowed to use more, in total, than the node had. | | Ceiling breach | Host exhaustion | |---|---|---| | Scope of the accounting | one container | the whole machine | | What triggers it | that container's usage crossing its declared figure | total free memory on the node approaching zero | | Who is stopped | a process inside that container | a process chosen at host scope, in any container | | Relation to fault | the victim is the cause | the victim need not be the cause | | Node's free memory at the time | may be mostly free | by definition nearly gone | ## How the victim is chosen at host scope The kernel's out-of-memory killer is not looking for a culprit. It is looking for **relief**, and it scores candidates mainly by how much memory stopping each one would return. Platforms then bias that score using what the workload declared — a workload sitting inside its reservation is made a less attractive victim, a workload that declared nothing a more attractive one — and platforms differ in how strongly they apply that bias. The consequences are counter-intuitive in exactly the way interviews probe: - The **largest resident set wins the lottery**, not the fastest grower. A report renderer holding a big working set entirely within its ceiling is cheap, effective relief. - The workload that actually consumed the node may be **small at the instant of the decision** — it allocated in a burst, the node hit zero, and the scoring saw a modest process. - **Nothing about your compliance is consulted directly.** The ceiling you honoured describes the boundary you were told not to cross; it says nothing about what happens when someone else exhausts the shared pool. - The kill is **abrupt**. There is no notice, no grace, and no negotiation with the process. ## The orderly path the platform would rather take The kernel's kill is the last resort, and a healthy platform tries to act before it. The node agent watches free memory and, at a line drawn **above** the point at which the kernel would act, starts reclaiming and evicting by rank: 1. Free memory falls past the agent's line. The agent reclaims what it can and, if that is not enough, stops workloads in a deliberate order, recording a reason. 2. If free memory keeps falling faster than reclaim can work, the line is overrun. 3. The kernel acts at host scope, by footprint, with no ranking and no record beyond the fact of the kill. So the *presence* of a host-scope kill is itself evidence: the orderly path did not get there in time. ## Reading the evidence at 02:00 - **Several unrelated workloads stopping within the same few seconds** points at the node, not at any one of them. - **One container dying repeatedly at a reproducible usage figure, on a node with memory to spare**, points at its own ceiling. - **A sharp slowdown with no stop at all** is not this at all — that is a CPU ceiling, a different resource with the opposite consequence. - **An abrupt stop with no recorded reason** points at the kernel; **a recorded rank and reason** points at the platform's own eviction. ## Why memory behaves differently from CPU CPU is compressible. Demand above a CPU ceiling is throttled: the process runs slower and stays alive, and when demand falls it recovers on its own. Memory a running process already holds cannot be taken back by asking. There is no throttle for it, so the only lever left is to stop something — which is why every path through a memory shortage ends in a kill or an eviction, and why the interesting question is only ever *which one*. ## What this actually changes for your workload - A declared ceiling protects the **node from your workload**, not your workload from the node. - The lever that affects your odds under host-scope pressure is the **reservation** you declared, because that is what the platform's ranking and victim-score bias are drawn against. - Sizing the working set close to what was declared, rather than leaving a declared figure that no longer reflects reality, is the practical form of that. - A kill with no recorded reason is worth escalating as a **node** problem, not a workload bug — treating it as a leak in your own service sends the investigation to the wrong place.
- What does the platform's own eviction give you that the kernel's host-scope kill does not?An order and a record. The node agent selects by declared rank rather than footprint, stops the workload as a termination rather than an abrupt disappearance, and writes down why. The kernel's kill is scored for relief, leaves no ranking, and can land inside a container that was entirely well behaved.
- If the same workload had exceeded its own memory ceiling instead, what would look different?The kill would be scoped to that one container, triggered by its own accounting, and would repeat at roughly the same usage figure. The node could be nearly idle at the time. Most importantly the victim would be the cause, so the investigation belongs in the workload rather than on the node.
- Why does a CPU shortage never produce this symptom?CPU is compressible. Demand above a CPU ceiling is throttled — the process gets less time, runs slower and stays alive, and recovers when demand falls. Memory already held by a running process cannot be reclaimed from it, so a memory shortage has no equivalent of slowing down; something has to stop.
An overbooked flight: you can hold a perfectly valid ticket and still be the passenger taken off, because the airline picks whoever is cheapest to bump, not whoever booked last.
saying these in an interview costs you the question
- Says a declared memory ceiling guarantees the container will not be killed.
- Assumes the process that caused the shortage is the one the kernel kills.
- Says exceeding a memory ceiling throttles the process instead of killing it.
- Treats every memory kill as a leak in their own service.
- Believes a workload always gets notice before a host-scope kill.