How much of every node would you hold back so the platform's ordered eviction, rather than the kernel's kill, chooses the victim?
answer
- two mechanisms can take the workload
- ranked stop versus abrupt kill
- reclaim needs room to finish
- the line is sized against a rate
- buys predictability, not memory
basics
~20 sHold back enough that reclaim finishes before free memory reaches zero. That headroom buys a victim chosen by declared rank with a recorded reason, instead of one chosen abruptly by footprint. The cost is capacity you never sell.
solid answer
~40 sTwo mechanisms can stop a workload on a node that is short of memory, and headroom decides which one wins. The platform's node agent reclaims and evicts by declared rank, leaving a termination and a recorded reason. The kernel's out-of-memory killer acts at host scope, scored for relief rather than intent, abruptly, and can take a process inside a perfectly compliant container — or one of the node's own agents. Headroom is the capacity you refuse to publish to the scheduler plus the free-memory line at which reclaim starts, and it has to be wide enough that reclaim finishes under the fastest growth you expect. Set it too tight and the kernel decides; too loose and you pay for capacity nothing ever uses. It buys predictable blast radius, not more memory.
go deeper
The takeaway is that a node deliberately keeps some memory unsold. That spare room is what lets the platform stop workloads in a chosen order rather than the kernel stopping one at random.
Explain the window: capacity that is never published to the scheduler, plus a free-memory line above the point the kernel acts, giving reclaim time to run before the fallback takes over.
Show how you would size and verify it — the growth rate against the reclaim cycle time, and unattributed kills as the signal that the window is closing too early.
Own the trade explicitly: a permanent density tax on every node, bought to make blast radius predictable and attributable, set per node class rather than once for the estate.
## The two mechanisms competing to stop your workload On a node that runs short of memory, exactly one of two things stops something, and they are not interchangeable. | | The platform's eviction | The kernel's host-scope kill | |---|---|---| | Who is chosen | ranked by usage against what was declared | scored mainly for how much memory it frees | | How it happens | a termination, in a chosen order | abruptly, with no ordering | | What is left behind | a recorded reason you can read at 02:00 | the fact of the kill and little else | | Can it hit the node's own agents | no — it stops workloads | yes, and then the node stops reporting | | Predictable in advance | yes, from the declarations | only statistically, by footprint | The first is a policy you wrote. The second is a fallback that exists so the machine survives. The whole question is which one gets to act. ## What headroom actually is Headroom is two settings that work together: - **Capacity you do not publish.** The node's real memory minus a reserve for its own agents, the runtime and the operating system. The scheduler only ever places against the published figure, so the reserve is never promised to a workload. - **The reclaim line.** The free-memory level at which the node agent starts reclaiming and, if needed, evicting — set above the level at which the kernel would act. Together they create a window between "the agent starts working" and "the kernel takes over". Everything you want — the ranking, the record, the graceful stop — happens inside that window. ## Sizing the line: reclaim needs time Reclaim is not instantaneous. It has to notice, rank, choose and stop, and the memory only comes back when the stopped workload's processes are actually gone. So the line is a **rate** question, not a size question: it has to sit far enough above zero that reclaim completes under the fastest allocation burst you expect on that class of node. A workable way to reason about it: 1. Estimate the fastest realistic growth in used memory per second on that node class, from the burstiest workload you allow there. 2. Estimate how long a reclaim cycle takes end to end, from crossing the line to memory actually returning. 3. The line needs to be at least the product of the two, with margin — anything tighter means the kernel routinely wins the race. Platforms differ in whether the line is expressed as an absolute amount or a percentage, and in whether the eviction that follows gets a grace period or is immediate; the arithmetic above is the same either way. ## The cost, stated honestly - Every unit held back is **capacity you never sell**, on every node, forever. Across a fleet it is a fixed tax on density. - It compounds with reservations: those already reduce how much the scheduler will place, and headroom reduces what it may place against. - Too tight and you lose both the choice of victim and the record of why — and the worst case is the kernel taking the node's own agent, which turns one workload's problem into the whole node going quiet. - Too loose and you are running more nodes than the work requires, to buy an outcome that was already comfortable. ## Where the figure should differ One number for the whole estate is the easy answer and usually the wrong one: - A node class running **latency-critical, hard-to-replace** workloads justifies more headroom: predictability is worth more there than density. - A node class running **retryable batch** work justifies less: an abrupt kill costs a retry, which the work already tolerates. - Nodes whose workloads **allocate in bursts** need a wider line than nodes whose workloads have flat working sets, because the line is sized against the rate. - The figure should be revisited when the workload mix changes, since it was derived from that mix. ## What it does not buy Headroom does not create memory and does not stop workloads being stopped. Pressure arrives less often on a node with headroom, but that is the placement arithmetic doing the work, not the reclaim line. What the line buys is that **when** something is stopped, it is the thing you nominated, stopped in a way that leaves evidence. It is also not a substitute for workloads declaring reservations: without those, the ranking inside your carefully bought window has nothing to rank on, and you have paid for an orderly process that is choosing arbitrarily.
- What evidence tells you the headroom figure is currently too tight?Host-scope kills appearing at all: abrupt stops with no recorded eviction reason, workloads inside their reservations disappearing, and worst of all the node's own agents being taken so the node drops out of reporting. Ordered evictions are the system working as designed; unattributed kills mean the window closed before reclaim finished.
- Does more headroom reduce how often workloads are stopped?Not directly. It changes which mechanism stops them and how predictably. It does mean the scheduler publishes less capacity and so places less work per node, which makes pressure arrive less often — but that is the placement arithmetic doing the work, not the reclaim line.
- Is headroom an alternative to making teams declare reservations?No — they solve different halves. Headroom buys the window in which the platform, not the kernel, chooses. Reservations are what the choosing is based on. With headroom and no declarations you have an orderly process picking effectively at random among workloads that all look equally unprotected.
saying these in an interview costs you the question
- Treats headroom as extra capacity rather than capacity deliberately not sold.
- Assumes the kernel's host-scope kill is as orderly as a platform eviction.
- Sets the reclaim line so tight that reclaim cannot finish in time.
- Thinks headroom removes the need for workloads to declare reservations.
- Sets one figure for the whole estate and never revisits it.
- Believes holding capacity back stops workloads being stopped.