skip to content

When a node runs short of memory, what decides which workload the node agent evicts first?

level: seniorimportance: must knowfreq 55%

answer

  1. rank, not blame
  2. the line is the reservation
  3. no reservation means first in line
  4. at or under reserved ranks last
  5. placement arithmetic, not a cap

basics

~20 s

Rank, not blame, decides. The node agent evicts by how far each workload's usage sits above what it reserved — a workload that declared no reservation is first in line, and one at or under its reservation is last.

solid answer

~40 s

Eviction is the platform's orderly reclaim, and it ranks rather than investigates. The line it measures against is the **reservation** each workload declared: the figure the scheduler subtracted from the node's capacity when it placed the workload. A workload that declared nothing is treated as having reserved zero, so all of its usage is overage and it goes first. Next come workloads above their reservation, worst overage first, with many platforms consulting a declared importance before the overage. A workload whose usage is at or under its reservation ranks last and is touched only if stopping the others was not enough. Note what a reservation is not: it is not a ceiling, the runtime does not enforce it, and exceeding it is perfectly legal right up until the node is short.

code

pseudocode · 12 lines
pseudocode
function evictionOrder(workloads):
    candidates = []
    for each w in workloads:
        # a workload that declared nothing reserves 0, so all of its usage is overage
        reserved = w.memoryReservation or 0
        overage  = w.memoryUsed - reserved
        candidates.add({ workload: w, overage: overage, priority: w.priority })

    # usage at or under the reservation gives a negative overage, so it sorts to the end
    sort candidates by (priority ascending, overage descending)

    return candidates    # the first element is stopped first

go deeper

for a junior

The key fact to hold: declaring a reservation is what puts your workload at the back of the queue when its node runs short. Declaring nothing puts it at the front.

for a middle

Explain the three bands and what separates them, and be precise that a reservation is what the scheduler subtracts at placement while a ceiling is what the runtime enforces on the running process.

for a senior

Show the diagnosis: reconstruct why the evicted workload was not the cause, reading each victim's declared reservation alongside its usage, and say what you would change so the pressure falls where it belongs.

for a principal

Treat the reservation as the estate's contract for blast radius. Generous reservations buy predictable victims and cost density, and a fleet where nothing declares one has no eviction policy at all.

## What eviction is, and what it is not Eviction is the **node agent's deliberate reclaim** under pressure: it notices free memory falling, frees what it can, and if that is not enough it stops workloads in a chosen order, recording why. It is not the kernel's last-resort kill, which happens at host scope, is scored purely for relief, and leaves no ranking behind. Two different mechanisms can take your workload off a node, and this is the one you can actually reason about in advance. Eviction also ranks rather than investigates. It has no notion of which workload *caused* the shortage; it has a line and a sort. ## The line the ranking is drawn against The line is the **reservation** — the figure declared for the workload that the scheduler subtracted from the node's capacity when it decided to place it there. Say precisely what each of the two declared figures does, because getting them backwards is the most common error in this material: - A **reservation** is placement arithmetic. It is what the scheduler treats as consumed on the node the moment the workload lands, whether or not the workload ever uses it. Under pressure it becomes the line the eviction ranking measures against. - A **ceiling** is runtime enforcement. It is what the runtime holds the running process to, and crossing it ends the process regardless of how much memory the node has. So a reservation is never enforced against you while things are healthy — you may sit well above it all day — and a ceiling is never consulted by the eviction ranking. ## The order, in practice | Band | Usage against its reservation | Where it ranks | |---|---|---| | Declared no reservation | reservation counts as zero, so all usage is overage | first out | | Above its reservation | ranked by how far over | next, worst overage first | | At or under its reservation | no overage at all | last, and only if stopping the others was not enough | Within a band, platforms differ in the tie-break. Most let the workload spec declare a relative importance and consult that before the overage; some rank on the overage alone; some also weigh how much memory stopping each candidate would actually return. Where the designs differ, the three bands themselves do not. The sequence on a node in trouble: 1. Free memory crosses the agent's reclaim line. The agent frees what it can without stopping anything. 2. Still short: it builds the ranking above and stops candidates from the top until free memory recovers. 3. Still short, or falling faster than the agent can work: the kernel acts at host scope instead, and the ranking no longer applies. ## What a reservation does not buy - **Not immunity.** It buys a place at the back of the queue, and queues get worked through. - **Not a cap.** Nothing stops you exceeding it; the consequence only arrives when the node is short. - **Not memory held aside for you.** It is an accounting figure, not physical pages pinned to your name. Another workload's pages occupy the machine whether or not you reserved room for yours. - **Not free.** The scheduler subtracted it from the node's capacity at placement, so a generous reservation you never use makes the node look fuller than it is and pushes other work elsewhere. That is the price of the rank. ## When everything is inside its reservation If the node is short and every workload on it is at or under what it reserved, the ranking has no comfortable candidate, and something is stopped anyway. That situation is a statement about the node: the reservations promised on it, plus the node's own agents, runtime and operating system, exceeded what it actually had to give. The fix is upstream of the ranking, not inside it. ## Why the victim is so often not the culprit This is the diagnosis that catches people out. A report renderer sharing a node with unrelated workloads is evicted; the workload whose growth exhausted the node keeps serving. Both facts follow from the ranking: - The renderer declared nothing, so every byte it used counted as overage and put it at the top of the queue. - The grower declared a generous reservation and, at the instant of the decision, was still inside it — so it sat in the protected band. Nothing malfunctioned. The ranking did exactly what it promises, and what it promises is not fairness by cause. ## Reading it at 02:00 - A **recorded eviction reason** means the orderly path ran; an abrupt stop with no record means it did not. - Check the victim's **declared reservation before its usage graph** — the graph tells you what it used, the declaration tells you where it ranked. - If unrelated workloads were stopped in quick succession, read the order they went in: that order *is* the ranking, and it names the bands your workloads were in.

  • Two workloads are both over their reservation by the same amount. What separates them?
    Platforms differ here. Most let the workload spec declare a relative importance and consult it before the overage, so a low-importance workload goes first even at equal overage. Others rank on overage alone, or weigh how much memory stopping each would actually return. The three bands are common; the tie-break is not.
  • A workload comfortably inside its reservation is stopped anyway. What happened?
    Either the reservations on that node, plus the node's own agents and operating system, promised more than it had — so the protected band still had to be worked through — or free memory fell faster than reclaim could act and the kernel made the choice at host scope instead, where the reservation is only a bias on the score.
  • Does a reservation you never use cost anything?
    Yes. The scheduler subtracts it from the node's capacity at placement, so the node counts as fuller than it is and other work is placed elsewhere. You are buying an eviction rank with density, which is a real trade rather than a free setting.

saying these in an interview costs you the question

  • Says the workload using the most memory overall is always evicted first.
  • Treats a reservation as a cap the runtime enforces on the process.
  • Believes a reservation makes a workload immune to eviction.
  • Thinks eviction targets whichever workload caused the shortage.
  • Says declaring no reservation is much the same as reserving a little.
  • Confuses the platform's ranked eviction with the kernel's host-scope kill.