skip to content

You are defining the pod priority tiers for a shared Kubernetes cluster running production services, internal tooling and opportunistic batch jobs. How do you choose the tiers and their values, and what failure modes do you design against?

level: principalimportance: nice to knowfreq 28%

answer

  1. Ladder answers one question: whose work stops when capacity runs out
  2. 4-5 tiers, 10× gaps, negative bottom tier, low globalDefault
  3. Default every class to Never; grant preemption as an exception
  4. Victim tier engineered cheap: short grace period, checkpointing
  5. Guard against inflation, PDB-as-guarantee, thrash, priority-instead-of-capacity

basics

~20 s

Use few, widely spaced tiers you can explain during an incident — roughly platform, production, internal, batch, preemptible — set the global default low, grant preemption rights only where start time must be bounded, and control who may use each tier with scoped quotas. Alert on preemption rate as a capacity signal.

solid answer

~60 s

**Keep the ladder small and spaced**: e.g. `platform: 1000000000`, `production: 1000000`, `internal: 100000`, `batch: 1000`, `preemptible: -10`. Wide gaps let you insert a tier later without renumbering; four or five tiers is what humans can reason about at 3am. **Set `globalDefault` low or leave it at 0** so unlabeled work sinks rather than floats. A high default silently flattens the scheme. **Grant preemption selectively.** Most tiers should carry `preemptionPolicy: Never` — they get queue precedence without eviction power. Only the tier whose start time must be bounded gets `PreemptLowerPriority`, and it should mostly find victims in the `preemptible` tier, which you design to be cheap: short `terminationGracePeriodSeconds`, checkpointing, controllers that requeue. **Enforce who may use what** with `ResourceQuota` scoped on `PriorityClass` per namespace, not by convention. **Design against**: priority inflation (everything becomes production), PDBs treated as a guarantee (preemption ignores them), preemption thrash where victims are recreated and preempted again, and using priority as a substitute for capacity. Alert on preemption rate and on pending-time SLOs, and remember preemption competes with the Cluster Autoscaler: sometimes buying a node is the right answer.

code

yaml · 26 lines
yaml
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata: {name: platform-critical}
value: 1000000000
preemptionPolicy: PreemptLowerPriority
description: "Cluster-wide blast radius if unavailable"
---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata: {name: production}
value: 1000000
preemptionPolicy: Never
---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata: {name: batch}
value: 1000
preemptionPolicy: Never
globalDefault: true
---
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata: {name: preemptible}
value: -10
preemptionPolicy: Never
description: "Disposable; short grace period required"

go deeper

for a junior

Focus on the recall: a small ordered list of named classes, the default should be low, and application pods never use system classes.

for a middle

Propose a concrete four-tier ladder with spaced values, explain globalDefault, and note that preemptionPolicy can be set per class.

for a senior

Add enforcement (scoped ResourceQuota), victim-tier engineering with short grace periods, and monitoring of preemption rate and pending duration.

for a principal

Frame the whole thing as policy: who owns each class, how inflation is governed, the explicit preemption-versus-autoscaling economics, and a rollout that starts non-preempting and grants eviction rights only after observation.

## Start from the decision the scheme has to make A priority scheme answers one question: *when the cluster is short of capacity, whose work stops?* Everything else — the numbers, the names, the policies — is bookkeeping. So begin by writing down, in plain language, the ordering you would defend to the affected teams. If you cannot state it in one sentence per tier, the scheme is too complicated to hold during an incident. A typical defensible ladder: | Tier | Value | Preemption | Meaning | |---|---|---|---| | `platform-critical` | 1000000000 | PreemptLowerPriority | Ingress, service mesh control, logging pipeline — cluster-wide blast radius if down | | `production` | 1000000 | Never | Customer-facing services | | `internal` | 100000 | Never | Dashboards, dev tooling, CI runners | | `batch` | 1000 | Never | Scheduled ETL, reports | | `preemptible` | -10 | Never | Opportunistic work, explicitly disposable | Note the shape: only the top tier can evict, values are orders of magnitude apart, and the bottom tier is *negative* so it sits below anything unlabeled (priority 0). ## Design rules **Few tiers, wide gaps.** Gaps of 10× or more mean you can insert `production-tier-2` later without touching existing classes — important because edits do not re-price running pods, so renumbering produces a cluster with two generations of the same name at different values. **Low or zero globalDefault.** The default class is the one nobody chose, which usually means the pod that nobody thought about. Letting it inherit a high number gives every unreviewed workload eviction rights. Point `globalDefault` at `batch`, or leave it unset so unlabeled pods are 0. **Preemption is a separate grant from rank.** `preemptionPolicy: Never` keeps queue precedence while removing eviction power. Default every tier to `Never`, then justify each exception. In most clusters exactly one or two tiers need real preemption: the platform tier (so a DaemonSet-adjacent component can always start) and, sometimes, a bounded-latency job tier whose victims are all `preemptible`. **Make the victim tier cheap.** If a tier exists to be preempted, engineer it that way: `terminationGracePeriodSeconds: 10–30`, checkpointing or idempotent restart, a controller that requeues rather than failing the job, and no PodDisruptionBudget illusions. A 300-second grace period on a victim means the "urgent" preemptor waits five minutes. **Enforce eligibility mechanically.** PriorityClasses are cluster-scoped and referencing them is not gated by RBAC. Per-namespace `ResourceQuota` with `scopeSelector` on `PriorityClass` — either an allowlist with real pod counts, or a zero quota on forbidden classes — is the enforcement point. Convention plus code review is not enforcement. ## Failure modes to design against **Priority inflation.** Left ungoverned, every team labels itself `production` and the ladder collapses into one rung. Counter it with quota-based eligibility, an approval path with a named owner per class, and periodic audits (`kubectl get pods -A -o custom-columns=...,PRIO:.spec.priority --sort-by=.spec.priority`). **PDB mistaken for a guarantee.** Preemption ranks candidate nodes by fewest PodDisruptionBudget violations, but it will break a budget if that is the only way to schedule. Teams that treat a PDB as protection will be surprised. Document that PDBs bind drains (the Eviction API), not preemption. **Preemption thrash.** Victims belong to controllers that recreate them, which then get preempted again — CPU burned, nothing progresses, logs full of Preempted events. Detect it by alerting on preemption *rate* per namespace, and fix it either by capacity or by making the low tier back off. **Nomination is not a booking.** `status.nominatedNodeName` is a soft reservation; the freed room can be taken by someone else, and grace periods delay it. So "high priority" never means "starts immediately". Alert on pending duration against an SLO instead of assuming the tier delivers. **Priority substituting for capacity.** A cluster that preempts continuously is under-provisioned. Preemption is a redistribution mechanism, not a source of resources. The genuinely comparative question is preemption versus Cluster Autoscaler scale-up: preemption is instant-ish and free but destructive; scale-up is non-destructive but costs money and takes minutes. Bounded-latency work on a cost-sensitive cluster leans on preemption of a disposable tier; everything else should lean on capacity. **Priority is not QoS, and not isolation.** A `production` pod with no memory requests is still `BestEffort` and can be OOM-killed regardless of its priority. Tiering must be paired with a requests/limits policy, and hard isolation between tenants remains a matter of separate node pools, taints or separate clusters — priority only orders the queue and the executions. ## Rollout Ship the classes with `preemptionPolicy: Never` everywhere first and observe for a few weeks: who ends up in which tier, how long each tier waits, whether the ordering matches intuition. Then enable preemption on the single tier that needs it, with the victim tier already engineered to be disposable. That sequencing gives you the fairness benefit with zero eviction blast radius while the scheme is still wrong.

  • When is adding nodes via the Cluster Autoscaler the better answer than granting a tier preemption rights?
    When the victims are not genuinely disposable — stateful pods, long-running jobs without checkpointing, or another tenant's services — the destruction cost exceeds the node cost. Preemption is fast and free in money but destroys work and can thrash; scale-up is non-destructive but takes minutes and costs. Grant preemption only where you can point at a specific disposable tier that will supply the victims, and rely on capacity otherwise.
  • How would you detect that a priority scheme has degraded in practice?
    Audit the distribution of spec.priority across pods: if most pods sit in the top one or two tiers, inflation has flattened the ladder. Watch the preemption event rate per namespace, since continuous preemption means chronic under-provisioning rather than working policy. Track pending duration per tier against an SLO, because a tier that waits as long as the one below it is not buying anything.

Triage tags in an emergency department: only a few colours, everyone must recognise them under stress, and the tag decides order of care — it does not create more doctors.

saying these in an interview costs you the question

  • Designing ten-plus tiers nobody can recall during an incident
  • Setting a high globalDefault so every unlabeled pod can preempt
  • Assuming PodDisruptionBudgets protect pods from preemption
  • Treating priority as tenant isolation or as a resource guarantee instead of node pools plus requests/limits
  • Using preemption to paper over a cluster that is simply too small

context