skip to content

In Kubernetes, what does a PriorityClass object do, how does a pod end up with a priority value, and what is the globalDefault field for?

level: middleimportance: must knowfreq 58%

answer

  1. Cluster-scoped object: name → int32 value
  2. priorityClassName → admission writes spec.priority
  3. Two effects only: queue order + preemption rights
  4. globalDefault: one class, applies at admission, not retroactive
  5. User ceiling 1e9; above is reserved for system classes

basics

~20 s

A PriorityClass is a cluster-scoped object mapping a name to an integer value. A pod names one in spec.priorityClassName; admission copies the number into spec.priority. Higher priority means earlier scheduling attempts and the right to preempt lower-priority pods. globalDefault marks the one class applied to pods that name none.

solid answer

~50 s

`PriorityClass` is a **cluster-scoped** (non-namespaced) object with a `value` (int32) and optional `globalDefault`, `preemptionPolicy` and `description`. A pod references it by name in `spec.priorityClassName`; the Priority admission controller resolves the name and writes the integer into the pod's `spec.priority` — you don't set that field yourself. The number does exactly two things: it **orders the scheduler's pending queue** (higher value gets a scheduling attempt first) and it **authorises preemption** — a pending pod may evict running pods of *strictly lower* priority when nothing else fits. `globalDefault: true` (allowed on at most one class) supplies the priority for pods that name no class; without such a class those pods get **0**. The default applies only at admission, so it does **not** retroactively change existing pods. User-defined values must be ≤ 1,000,000,000; above that is reserved for the built-in `system-cluster-critical` / `system-node-critical` classes. Editing a class's value later does not re-price already-admitted pods.

code

yaml · 26 lines
yaml
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: production
value: 100000
globalDefault: false
description: "Customer-facing services"
---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: best-effort-batch
value: 0
globalDefault: true
preemptionPolicy: PreemptLowerPriority
description: "Default for pods that name no class"
---
apiVersion: v1
kind: Pod
metadata:
  name: checkout-api
spec:
  priorityClassName: production
  containers:
    - name: app
      image: registry.example.com/checkout:1.4.2

go deeper

for a junior

Know that a PriorityClass is a named integer, that pods reference it by name in spec.priorityClassName, and that higher numbers get scheduled first.

for a middle

Explain admission resolving the name into spec.priority, the globalDefault rule and its at-admission-only nature, the 1e9 user ceiling, and that priority is unrelated to QoS.

for a senior

Add the operational consequences: a missing class blocks pod creation, edits do not re-price running pods, consumption is limited via ResourceQuota scopeSelector, and a high globalDefault flattens the whole scheme.

for a principal

Frame it as a cluster-wide policy surface: how many tiers the organisation can actually explain during an incident, spacing values to allow later insertion, and who owns the classes versus who may reference them.

## What "priority" means here Every pod carries an integer field `spec.priority`. Higher means more important. Kubernetes uses that integer for exactly two purposes: 1. **Queue order in the scheduler.** kube-scheduler keeps pending pods in a priority queue; its default `QueueSort` plugin orders by priority descending, then by the time the pod became unschedulable. A high-priority pod therefore gets a scheduling *attempt* before lower-priority pods that have been waiting longer. 2. **Eligibility to preempt.** If a pending pod cannot fit on any node, the scheduler may evict running pods whose priority is *strictly lower* to make room. Priority is **not** a CPU share, **not** the QoS class (`Guaranteed`/`Burstable`/`BestEffort`, which comes from requests vs limits), and it does not make a running pod faster or protect it from anything other than being outranked. A running pod's priority becomes relevant again only in two places: it can be chosen as a preemption victim, and kubelet's node-pressure eviction uses priority as a tiebreaker when ranking pods that exceed their requests. ## The PriorityClass object ```yaml apiVersion: scheduling.k8s.io/v1 kind: PriorityClass metadata: name: business-critical value: 1000000 globalDefault: false preemptionPolicy: PreemptLowerPriority description: "Customer-facing request path" ``` Key properties: - **Cluster-scoped.** There is no namespace on a PriorityClass; the same names are visible to every namespace. Access is controlled with RBAC on the `scheduling.k8s.io` group, and *consumption* is limited per namespace with a `ResourceQuota` carrying a `scopeSelector` on `PriorityClass`. - **`value` is an int32**, and values for user-created classes must be **≤ 1,000,000,000** (one billion). The range above that is reserved so that system components always outrank anything a tenant can create. - **`description`** is free text shown to humans; use it, because a bare integer tells the next engineer nothing. - **`preemptionPolicy`** (`PreemptLowerPriority`, the default, or `Never`) decides whether pods in this class are allowed to evict others. It does not change their queue position. ## How a pod acquires the value You write a *name*, not a number: ```yaml spec: priorityClassName: business-critical ``` The **Priority admission controller** (enabled by default) looks the class up and sets `spec.priority: 1000000` on the pod object. Consequences worth remembering: - Referencing a class that does not exist makes pod **creation fail** — which matters when a Deployment's template names a class you forgot to install: the ReplicaSet controller keeps retrying and no pods appear. - Because resolution happens at admission, the priority is **frozen into the pod**. Changing the PriorityClass `value` afterwards, or deleting the class, does not change pods that already exist. Deleting a class only prevents *new* pods from using it. - You cannot hand-set `spec.priority` to an arbitrary number to outrank your neighbours; the API rejects a value that disagrees with the named class (only the system may set priority directly). ## globalDefault At most one PriorityClass in the cluster may set `globalDefault: true`. Pods admitted **without** a `priorityClassName` then receive that class's value. If no such class exists, priority defaults to **0**. Two details bite people: - The default is applied **at admission only**. Creating a `globalDefault` class today does not raise yesterday's pods; they stay at whatever they were admitted with. - Setting a high `globalDefault` is a trap: it lifts the floor for everything unlabeled, so your "important" tier stops being distinguishable and *every* new pod becomes able to preempt older ones. Most clusters either leave the default at 0 or point `globalDefault` at a deliberately **low** class (e.g. value 0 or negative — negative values are legal) so that opportunistic work sinks below labeled workloads. If two classes claim `globalDefault: true` (possible if they race), the behaviour is undefined-ish: the smallest value wins in practice, but treat it as a misconfiguration and fix it. ## Built-in classes Every cluster ships `system-cluster-critical` (value 2000000000) and `system-node-critical` (value 2000001000) for control-plane and node-level add-ons. They sit above the user ceiling by design. Do not attach them to application workloads: they can preempt essentially anything, and `system-node-critical` also changes kubelet eviction behaviour. ## Practical shape of a scheme A workable ladder is small and spaced out — for example `platform-critical: 1000000`, `production: 100000`, `staging: 10000`, `batch: 0`, `preemptible: -10`. Wide gaps leave room to insert a tier later without renumbering, and few tiers keep the eviction story explainable during an incident.

  • If you raise the value of an existing PriorityClass, what happens to pods already running under it?
    Nothing. The Priority admission controller resolved the name to an integer when each pod was created and stamped it into spec.priority, so running pods keep the old number. Only pods admitted after the edit get the new value, which means a cluster can hold two generations of the same class name at different priorities. If you need uniform behaviour, roll the workloads so their pods are recreated.
  • How do you stop a tenant namespace from using your highest PriorityClass?
    PriorityClass is cluster-scoped, so RBAC on the object only controls who can create or edit classes, not who can reference them. Restrict consumption with a ResourceQuota in the namespace that carries a scopeSelector on PriorityClass — either an In list of allowed classes with quota, or a zero-pod quota scoped to the class you want to forbid. Admission then rejects pods referencing it. A validating admission policy or webhook is the alternative when you want a message rather than a quota error.
  • What does a pod get if there is no globalDefault class and it names none?
    Priority 0. That is the neutral middle of the range, since negative values are legal, so pods in a class with a negative value sit below unlabeled pods and can be preempted by them.

Think of an airline boarding group printed on your ticket at check-in: it decides who is called first and, in an overbooking, who gets bumped. Changing the group rules tomorrow does not reprint tickets already issued.

saying these in an interview costs you the question

  • Saying PriorityClass changes CPU or memory shares — it affects scheduling order and preemption only, never runtime resource allocation
  • Confusing priority with QoS class; QoS comes from requests vs limits and drives kubelet eviction, priority is a separate integer
  • Believing a higher priority protects a running pod from being descheduled or restarted for any reason
  • Thinking PriorityClass is namespaced, or that RBAC on the object stops tenants from referencing it (that needs a scoped ResourceQuota)
  • Claiming editing a class's value re-prices existing pods

context