skip to content

A Kubernetes pod cannot be scheduled because no node has room for it. Walk through, step by step, what kube-scheduler's preemption logic does and what happens to the pods it removes.

level: seniorimportance: must knowfreq 52%

answer

  1. PostFilter plugin, only when Filter found zero nodes
  2. Strictly lower priority victims; simulate removal, re-run filters
  3. Minimise victims, prefer fewer PDB violations (soft, not hard)
  4. status.nominatedNodeName = soft reservation, not a binding
  5. Victims deleted gracefully — grace period delays the freed room

basics

~20 s

After all nodes fail filtering, the scheduler's PostFilter preemption plugin looks for one node where deleting some strictly-lower-priority pods would let the pending pod fit. It picks the least disruptive victim set, sets nominatedNodeName on the pending pod, and deletes victims gracefully. The freed node is not guaranteed to the pod.

solid answer

~60 s

Preemption runs in the **PostFilter** extension point, only after the normal filter phase found zero feasible nodes and only if the pod's `preemptionPolicy` is `PreemptLowerPriority`. The `DefaultPreemption` plugin: 1. **Finds candidate nodes** — simulates removing all pods of *strictly lower* priority on each node and re-runs the filters; nodes that still fail (a taint, an affinity rule, a missing volume) are discarded. 2. **Minimises victims** per candidate — it adds back the highest-priority removable pods while the pod still fits, and prefers sets that violate fewer PodDisruptionBudgets. 3. **Picks one node** by heuristics: fewest PDB violations, then lowest highest-priority victim, then smallest priority sum, then fewest victims, then latest-starting pods. 4. **Records `status.nominatedNodeName`** on the pending pod so other pods account for the reservation, then **deletes the victims via the API with their graceful termination period**. The pod then goes back to the queue. It is *not* bound to that node — another pod, or the pod's own changed conditions, can take the room. Victims get SIGTERM and their normal grace period, so freeing capacity is not instantaneous.

code

bash · 3 lines
bash
kubectl get pod urgent-job -o jsonpath='{.status.nominatedNodeName}{"\n"}'
kubectl get events -A --field-selector reason=Preempted
kubectl describe pod urgent-job | sed -n '/Events/,$p'

go deeper

for a junior

Know the headline: if nothing fits, the scheduler may delete lower-priority pods to make room, and those pods are terminated normally.

for a middle

Describe the sequence — filters fail, PostFilter simulates removing lower-priority pods, victims are deleted gracefully, the pod is requeued — and that priority must be strictly lower.

for a senior

Add the operational truth: PDBs are only a preference, nomination is soft, grace periods delay real capacity, and constraint-based failures produce 'no preemption victims found'.

for a principal

Treat sustained preemption as a capacity and tiering signal: decide which workloads are designed to be victims, cap grace periods there, and weigh preemption against adding nodes via the autoscaler.

## Where preemption sits kube-scheduler handles one pending pod at a time in its **scheduling cycle**: PreFilter → Filter → **PostFilter** → PreScore → Score → Reserve → Permit. If the Filter phase returns at least one feasible node, PostFilter never runs. Preemption is therefore strictly a *last resort* for a pod that would otherwise sit in `Pending` forever, and it is implemented by a normal framework plugin called `DefaultPreemption`. Two gates must pass before anything is evicted: - The pending pod's `spec.priority` must be **higher than** the priority of the pods it wants to remove. Equal priority is not enough — victims must be *strictly* lower. - The PriorityClass's `preemptionPolicy` must be `PreemptLowerPriority` (the default). `Never` disables this whole path for that pod. ## Step 1 — candidate nodes For each node the plugin builds a hypothetical node state with every strictly-lower-priority pod removed, then re-runs the filter plugins. This matters: preemption only helps when the *reason* the pod did not fit is something that removing pods can fix. If the node failed because of an untolerated taint, an unsatisfied node affinity, a node-level port conflict with a higher-priority pod, or an unattachable volume, it stays infeasible and is dropped. That is why you often see the event `preemption: 0/40 nodes are available: 40 No preemption victims found for incoming pod` — the pod is blocked by constraints, not by capacity. ## Step 2 — minimising the victim set Removing *all* lower-priority pods is legal but wasteful, so the plugin computes a smaller set. It first checks PodDisruptionBudgets: victims are split into those whose removal would violate a PDB and those that would not. It then tries to **add pods back** in decreasing priority order — keeping the most important ones — as long as the pending pod still fits. The result is roughly "the cheapest set of the least important pods that frees enough room". A crucial caveat: **preemption ignores PDB as a hard rule**. A PDB makes a victim less attractive, but if there is no other way to schedule the pod, the scheduler will still delete a pod whose budget is exhausted. PDBs are honoured by the *eviction API* (drains); preemption deletes pods directly. ## Step 3 — choosing the node Among candidate nodes the plugin applies an ordered tiebreak: fewest PDB violations → the node whose *highest-priority victim* is lowest → smallest sum of victim priorities → fewest victims → the node where victims started most recently. The intent is to disturb the least important, least established work. ## Step 4 — nomination and deletion The scheduler patches the pending pod with `status.nominatedNodeName: <node>`. This is a soft reservation with two effects: humans can see where the pod is headed, and subsequent scheduling cycles account for nominated pods' resources on that node so the scheduler does not hand the same room away in the very next cycle. It then **deletes** each victim through the API server. Deletion is graceful: each victim receives SIGTERM, runs its `preStop` hook, and has until its `terminationGracePeriodSeconds` before SIGKILL. Capacity is only really free once the containers exit and kubelet reports the pod gone. For a pod with a 300-second grace period, the "urgent" pod can wait five minutes. The pending pod is put back in the scheduling queue. **Nomination is not a binding.** Between nomination and the next successful cycle: - another, even higher-priority pod may take the freed room; - the victims' controllers immediately recreate replacement pods, which may land elsewhere or compete; - the Cluster Autoscaler may add a node and the pod may land there instead; - if the pod's own conditions change (it is deleted, or now fits elsewhere), the nomination is dropped. This is why preemption is best-effort and why repeated preemption can look like churn. ## Known limits worth naming - **No cross-node preemption for affinity.** If the pending pod requires inter-pod affinity to a pod on node A, the scheduler will not preempt on node B to satisfy it; it only considers victims on the node it is evaluating. - **Victims that are themselves preemptors** (holding a nomination) complicate matters and are handled conservatively. - **Priority does not evict later.** Once a pod is running, a newly created higher-priority pod only triggers preemption if it cannot schedule anywhere; a cluster with spare capacity never preempts. - **Kubelet is a separate actor.** Under node memory or disk pressure, kubelet performs *node-pressure eviction*, ranking first by whether usage exceeds requests, then by priority. That is a different mechanism from scheduler preemption, though both consult priority. ## Operating it Watch for `Preempted` / `Preempting` events, alert on preemption rate, keep grace periods short on preemptible tiers, and remember that a permanently preempting cluster is a capacity problem wearing a scheduling costume.

  • Does preemption respect PodDisruptionBudgets?
    Only as a preference. The scheduler prefers victim sets that violate no PDB and ranks candidate nodes by fewest violations, but if the only way to schedule the pod is to break a budget, it deletes the pod anyway. PDBs are enforced by the Eviction API used during node drains, whereas preemption issues plain deletes. So a PDB is not a guarantee against preemption.
  • A pod has nominatedNodeName set but is still Pending minutes later. What are the likely causes?
    Victims are still terminating because of long terminationGracePeriodSeconds or a slow preStop hook, so the room is not actually free yet. Or the freed capacity was taken by another pod, since nomination is a soft reservation rather than a binding. It can also mean the victims' controllers recreated replacements that grabbed the space, or that the pod is blocked by a constraint preemption cannot fix and the nomination is stale.
  • Why does a pod sometimes report 'No preemption victims found' even though the node is full of low-priority pods?
    Because the node failed filtering for a reason that removing pods cannot fix — an untolerated taint, unsatisfied node affinity or nodeSelector, a volume that cannot attach, or a topology-spread constraint. The plugin simulates removing lower-priority pods and re-runs the filters; if the node still fails, it is not a candidate. It also reports this when the pods present are of equal or higher priority.

An overbooked restaurant bumping the smallest, most recently seated party rather than the big reservation — and even then, the table isn't yours until you actually sit down; someone else may reach it first.

saying these in an interview costs you the question

  • Saying preemption evicts equal-priority pods — victims must be strictly lower priority
  • Believing a PodDisruptionBudget hard-blocks preemption
  • Assuming nominatedNodeName guarantees the pod gets that node
  • Thinking capacity is freed instantly, ignoring terminationGracePeriodSeconds and preStop hooks
  • Claiming a high-priority pod evicts running pods even when it could schedule on a node with free capacity

context