What does setting preemptionPolicy: Never on a Kubernetes PriorityClass change about how pods using that class are scheduled, and when would you want it?
answer
- Field on PriorityClass: PreemptLowerPriority (default) | Never
- Keeps queue order, drops the PostFilter victim search
- Never = 'I won't evict others', NOT 'I can't be evicted'
- Classic use: big batch/ML jobs that must not kill services
- Saturated cluster → such a pod can stay Pending indefinitely
basics
~20 sPods in that class keep their high priority for queue ordering — they get scheduling attempts before lower-priority pods — but they never evict anyone. If nothing fits they stay Pending. Use it for important-but-not-urgent work that should jump the queue without disrupting running workloads.
solid answer
~50 s`preemptionPolicy` is a field on **PriorityClass** with two values: `PreemptLowerPriority` (the default) and `Never`. It splits the two things priority normally does. With `Never`, the pod's `spec.priority` still controls its **position in the scheduler's pending queue**, so it is attempted before lower-priority pods and gets first refusal on capacity that frees up. What it loses is the **PostFilter preemption step**: if the filter phase finds no feasible node, the scheduler does not look for victims — the pod simply stays `Pending` and is retried. The classic use is high-priority batch or data-science jobs: you want them to be next in line whenever a node frees up, but you do not want them killing running services. It also suits workloads that would be expensive to interrupt others for, or tiers in a shared cluster where a tenant may be *ranked* high but must not be allowed to evict another tenant's pods. The policy is fixed once the pod is admitted, like the priority value itself.
code
yaml · 24 linesapiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: research-high
value: 900000
preemptionPolicy: Never
description: "First in queue, never evicts running pods"
---
apiVersion: batch/v1
kind: Job
metadata:
name: train-ranker
spec:
template:
spec:
priorityClassName: research-high
restartPolicy: Never
containers:
- name: trainer
image: registry.example.com/trainer:0.9
resources:
requests:
cpu: "16"
memory: 64Gigo deeper
Know the one-liner: the pod still gets attempted first, but it never kills other pods to make room.
Separate the two effects of priority — queue ordering versus preemption rights — and name the field's two values and their default.
Discuss when to choose it: big batch jobs, multi-tenant fairness, PDB-sensitive neighbours, and the risk of indefinite Pending on fixed capacity.
Position it as the safe default when rolling out a priority scheme, with preemption granted only to the tiers that genuinely need bounded start time, and pair it with autoscaling or dedicated capacity.
## The two jobs of priority, separated A pod's integer priority normally buys two privileges: 1. **Queue position.** kube-scheduler's default QueueSort plugin orders pending pods by priority descending. Higher-priority pods get a scheduling attempt first, so when a node frees up they have first claim. 2. **Preemption rights.** If no node can host the pod, the PostFilter stage may delete strictly-lower-priority pods to make room. `preemptionPolicy: Never` keeps privilege 1 and removes privilege 2. ```yaml apiVersion: scheduling.k8s.io/v1 kind: PriorityClass metadata: name: high-priority-nonpreempting value: 1000000 preemptionPolicy: Never globalDefault: false description: "Front of the queue, but never evicts running pods" ``` The field lives on the **PriorityClass**, not on the pod, and like the priority value it is resolved at admission and stamped onto `pod.spec.preemptionPolicy`. Editing the class later does not change pods already created. ## What actually differs at runtime When a `Never` pod fails the filter phase on every node, the scheduler records the failure and requeues the pod; there is no search for victims, no `nominatedNodeName`, no deletions. The pod's events show only the normal "0/N nodes are available…" message, without the trailing preemption verdict. Meanwhile, because the pod sits near the front of the queue, the *next* time a node has room — a job finished, a deployment scaled down, the Cluster Autoscaler added a node — this pod is evaluated before the lower-priority pods waiting behind it. In a busy cluster that difference is substantial: without the high priority, a stream of small low-priority pods can starve a large pod indefinitely. ## When to reach for it - **Large batch or ML training jobs.** They need a big contiguous slice of a node and would starve at low priority, but killing production services to start a training run is not acceptable. - **Multi-tenant clusters.** You may be happy to say "tenant A's jobs go first in the queue" while refusing to let tenant A destroy tenant B's running pods. Non-preempting classes make priority a *fairness* knob rather than a *violence* knob. - **Workloads whose neighbours are expensive to restart** — stateful pods with long recovery, or anything under a tight PodDisruptionBudget that preemption would ignore anyway (preemption treats PDBs as a preference only). - **Incremental rollout of a priority scheme.** Introducing tiers with `Never` everywhere first lets you observe queue behaviour without any eviction blast radius, then enable preemption on the one tier that truly needs it. ## Interactions worth knowing - **It does not protect the pod.** `Never` says "I will not preempt others"; it says nothing about whether *this* pod can be preempted. A pod in a non-preempting class is still a valid victim for any pod of higher priority whose policy is `PreemptLowerPriority`. If you want a pod hard to evict, give it a *high value*, not a `Never` policy. - **It does not disable kubelet eviction.** Under node memory or disk pressure, kubelet still evicts pods, ranking by whether usage exceeds requests and then by priority. Scheduler preemption and node-pressure eviction are different systems. - **It plays well with the Cluster Autoscaler.** An unschedulable non-preempting pod is still an unschedulable pod, so it triggers scale-up like any other. Combining a high non-preempting priority with an autoscaled node group is a common way to get "important jobs start soon, nobody gets killed" — at the cost of node spend and scale-up latency. - **Sub-priority ordering still applies.** Among several `Never` pods, the higher value is still attempted first. ## The failure mode to expect The honest downside is that a `Never` pod can wait forever in a saturated cluster: nothing frees room, nothing is evicted, and the only progress comes from workload churn or new nodes. Operationally you should alert on pods pending beyond an SLO rather than assume high priority means "will run soon". If a class of work genuinely must run within a bounded time on fixed capacity, it needs preemption rights or dedicated capacity — a non-preempting priority is a preference, not a guarantee.
- Does preemptionPolicy: Never protect a pod from being preempted by others?No. The field only controls whether this pod may evict others. A pod in a non-preempting class is still a legitimate victim for any pending pod of strictly higher priority whose policy is PreemptLowerPriority. To make a pod hard to evict you raise its priority value, since victims must always be strictly lower priority than the preemptor.
- How does a non-preempting high-priority pod interact with the Cluster Autoscaler?It stays unschedulable when nothing fits, and unschedulable pods are exactly the signal the Cluster Autoscaler acts on, so it triggers a scale-up like any other pending pod. That combination is a common design: important jobs get queue precedence and new capacity rather than evicting neighbours. The cost is scale-up latency and extra node spend.
saying these in an interview costs you the question
- Thinking Never makes the pod immune to preemption rather than unable to preempt
- Believing Never also removes the pod's queue-order advantage
- Setting preemptionPolicy on the Pod spec directly and expecting it to override the class
- Assuming a high non-preempting priority guarantees the pod runs soon in a saturated cluster
- Confusing it with disabling kubelet node-pressure eviction