Kubernetes ships two built-in PriorityClasses named system-cluster-critical and system-node-critical. What are they for, how do they differ, and what rules should govern their use?
answer
- Both above the 1e9 user ceiling — tenants can never outrank them
- cluster-critical 2000000000 = reschedulable add-ons (CoreDNS, metrics-server)
- node-critical 2000001000 = per-node DaemonSets (CNI, kube-proxy, CSI node)
- node-critical also ranks last in kubelet node-pressure eviction
- Restrict with ResourceQuota scopeSelector on PriorityClass
basics
~20 sThey are reserved classes for infrastructure: system-cluster-critical (2000000000) for control-plane-level add-ons like CoreDNS, and system-node-critical (2000001000, the highest) for per-node agents like the CNI and kube-proxy. Both sit above the one-billion user ceiling. Never put application workloads in them.
solid answer
~60 sBoth are created by the cluster automatically and sit **above the 1,000,000,000 ceiling** that user-created PriorityClasses may not exceed, so nothing a tenant can define outranks them. - **`system-cluster-critical` = 2000000000** — for add-ons the *cluster* cannot work without but that can move between nodes: CoreDNS, metrics-server, the CSI controller, the cluster autoscaler. - **`system-node-critical` = 2000001000** — the highest priority in the cluster, for pods a *node* cannot work without: the CNI daemonset, kube-proxy, the CSI node plugin, log/monitoring agents you truly cannot lose. Beyond preemption, kubelet also treats priority in node-pressure eviction ranking, so node-critical pods are the last to be evicted under memory or disk pressure. Rules of use: they are for platform components only, applied through the DaemonSet/Deployment templates the platform team owns. Keep tenant access away with a `ResourceQuota` carrying a `scopeSelector` on `PriorityClass`, or an admission policy. Putting an application in `system-node-critical` means it can preempt CoreDNS and survive pressure eviction ahead of it — a reliable way to make an outage worse.
code
yaml · 21 linesapiVersion: apps/v1
kind: DaemonSet
metadata:
name: cni-agent
namespace: kube-system
spec:
selector:
matchLabels: {app: cni-agent}
template:
metadata:
labels: {app: cni-agent}
spec:
priorityClassName: system-node-critical
tolerations:
- operator: Exists
containers:
- name: agent
image: registry.example.com/cni-agent:1.7
resources:
requests: {cpu: 100m, memory: 128Mi}
limits: {cpu: 100m, memory: 128Mi}go deeper
Know they exist, that they are reserved for cluster components like CoreDNS and the CNI agent, and that application pods should not use them.
Give the values and the cluster-versus-node distinction, and explain why they sit above the one-billion user ceiling.
Add kubelet's node-pressure eviction ranking, the ResourceQuota scopeSelector control, and the outage pattern where an app in a system class preempts DNS.
Define the governance model: who may grant these classes, a platform-owned high user class for platform-but-not-system components, and enforcement via admission policy plus auditing.
## Why reserved classes exist User-created PriorityClasses are capped at **1,000,000,000**. That cap exists precisely so the cluster can reserve a band above it for its own components. Two classes ship in that band by default: | Class | Value | Intended for | |---|---|---| | `system-cluster-critical` | 2000000000 | Cluster-level add-ons that must run *somewhere* | | `system-node-critical` | 2000001000 | Pods that must run on *every* node for it to function | Because preemption requires the preemptor to have strictly higher priority than the victim, nothing a tenant creates can ever evict a pod in either class, while a node-critical pod can evict everything — including cluster-critical pods. ## The distinction in practice **Cluster-critical** describes reschedulable infrastructure. CoreDNS is the canonical example: DNS resolution breaking takes the whole cluster down, but any node will do, so if one node is full CoreDNS can start elsewhere. Metrics-server, the CSI controller deployment, ingress controllers, and the Cluster Autoscaler itself typically live here. **Node-critical** describes per-node agents, almost always DaemonSets, where "schedule it elsewhere" is meaningless. Without the CNI plugin the node cannot give pods networking; without kube-proxy (or its replacement) service routing on that node fails; without the CSI node plugin volumes cannot mount. These pods must be able to displace *any* workload on their node, so they get the highest value in the cluster. The second, less-known effect of `system-node-critical` is on **kubelet node-pressure eviction**. When a node hits a memory or disk eviction threshold, kubelet ranks pods first by whether their usage exceeds their requests, then by **pod priority**, then by how far usage exceeds requests. Node-critical pods therefore end up last on the kill list. Some kubelet configurations additionally give critical pods access to a portion of reserved resources. The upshot: attaching this class to an application does not just affect scheduling — it changes who dies when a node runs out of memory. ## Governance The classes are cluster-scoped, so they are *visible* to every namespace. RBAC on `scheduling.k8s.io` controls who can create or edit PriorityClasses, but **not** who may reference one from a pod. To restrict consumption you use a `ResourceQuota` with a `scopeSelector`: ```yaml apiVersion: v1 kind: ResourceQuota metadata: name: deny-system-critical namespace: team-payments spec: hard: pods: "0" scopeSelector: matchExpressions: - operator: In scopeName: PriorityClass values: ["system-cluster-critical", "system-node-critical"] ``` With that in place, a pod in `team-payments` naming either class is rejected at admission with a quota error. The inverse pattern — an `In` quota listing the classes a namespace *may* use with a real pod count — is also common. Many platforms additionally enforce it with a validating admission policy so the error message is explanatory rather than a confusing quota failure. ## Failure modes from misuse - **Application in a system class.** It can preempt CoreDNS, kube-proxy or the CNI agent. Losing DNS to make room for a batch job converts a capacity problem into a cluster-wide outage, and the blast radius is not limited to the offending team. - **Everything critical.** If every platform DaemonSet is node-critical, the class stops discriminating; during pressure kubelet has no useful ranking left, and nodes cannot shed anything. - **Node-critical on a Deployment with many replicas.** Node-critical is a statement about per-node necessity; a scalable service in that class will simply bulldoze workloads across the fleet. - **Assuming the class implies resource guarantees.** It does not. A node-critical pod with no requests is still a `BestEffort` pod for QoS purposes and can be OOM-killed by the kernel long before kubelet's ranking matters. Critical components need explicit requests — ideally `Guaranteed` QoS — in addition to the priority class. ## Checklist for a real cluster 1. Inventory which pods currently carry either class (`kubectl get pods -A -o custom-columns=...,PRIO:.spec.priorityClassName`). 2. Confirm every node-critical pod is genuinely a per-node DaemonSet. 3. Ensure each carries CPU/memory requests, not just a priority. 4. Block tenant namespaces from both classes with a scoped quota or admission policy. 5. Give your own platform-but-not-system components a high *user* class (e.g. 1,000,000,000) rather than borrowing a system one.
- How do you prevent a tenant namespace from using system-node-critical?RBAC on PriorityClass objects only governs creating and editing them, not referencing them, so you use a ResourceQuota in the namespace with a scopeSelector matching PriorityClass In [system-node-critical] and hard pods: "0". Any pod naming that class is then rejected at admission. Many platforms pair this with a validating admission policy so the rejection message explains the rule.
- Does system-node-critical guarantee the pod gets the memory it needs?No. Priority influences preemption and kubelet's eviction ranking, but the kernel OOM killer works from cgroup limits and oom_score_adj derived from QoS class, not from PriorityClass. A node-critical pod with no resource requests is still BestEffort and can be killed first by the kernel. Critical components should declare explicit requests and limits, ideally landing in Guaranteed QoS.
saying these in an interview costs you the question
- Using a system class for an application workload to 'make sure it runs'
- Believing the system classes are namespaced or that RBAC alone prevents their use
- Saying system-cluster-critical is higher than system-node-critical
- Assuming the class provides resource guarantees or OOM protection independent of QoS
- Thinking you can create a user PriorityClass above one billion to outrank them