When a Kubernetes kubelet must reclaim memory on a node, in what order does it choose which pods to kill? Explain how a pod's Quality of Service class and its resource requests determine that ranking.
answer
- Guaranteed = limits == requests, every container
- BestEffort = nothing set → always over request → first out
- rank: over-request? → priority → size of overage
- Guaranteed = last in line, not exempt
- oom_score_adj: −997 / scaled / 1000
basics
~20 sThe kubelet ranks pods by whether usage exceeds requests, then by pod priority, then by how much usage exceeds requests. In practice BestEffort pods (no requests) go first, then Burstable pods over their request, and Guaranteed pods (limits equal requests) last.
solid answer
~60 sA pod's **QoS class** is derived from its containers' requests and limits: - **Guaranteed** — every container sets both, and limits equal requests for CPU and memory. - **Burstable** — at least one request set, but not the Guaranteed pattern. - **BestEffort** — no requests or limits anywhere. Under memory pressure the kubelet sorts candidates by: 1. whether the pod's memory usage **exceeds its request** — pods over request are evicted before pods under it; 2. **pod priority** (`priorityClassName`) — lower priority first; 3. how much usage exceeds the request — the biggest overage first. Because BestEffort pods have a request of zero, they are *always* over request, so they are evicted first. Burstable pods that outgrew their request come next; Burstable pods within their request and Guaranteed pods are effectively last and are only touched if reclaiming from everything else was not enough. Guaranteed is not immunity. The practical lesson: setting honest memory **requests** is what protects a pod, and the node kernel OOM score is derived from the same QoS classes.
code
yaml · 12 lines# Guaranteed: limits == requests on every container
resources:
requests: { cpu: "500m", memory: "512Mi" }
limits: { cpu: "500m", memory: "512Mi" }
---
# Burstable: request set, limit higher (or absent)
resources:
requests: { cpu: "100m", memory: "256Mi" }
limits: { memory: "1Gi" }
---
# BestEffort: nothing set — evicted first
resources: {}go deeper
Know the three QoS classes and how requests/limits produce them, and that pods with no requests are evicted first.
Give the real ranking — over-request first, then lower priority, then largest overage — and explain why BestEffort always sorts to the top.
Add operational judgment: right-sizing requests as the primary defence, PriorityClasses for node-critical DaemonSets, that PDBs do not apply, and the oom_score_adj mapping.
Weigh packing density against blast radius: Guaranteed everywhere wastes capacity, BestEffort batch tiers absorb pressure by design; decide per workload class and encode it in platform policy.
## Quality of Service classes Kubernetes assigns each pod one of three **QoS classes** at admission time, purely from its containers' `resources.requests` and `resources.limits`: - **Guaranteed** — *every* container in the pod specifies both CPU and memory requests and limits, and for each resource the limit equals the request. This is the strongest class. - **Burstable** — the pod does not qualify as Guaranteed, but at least one container sets a request or limit. The pod is entitled to its request and may burst above it if the node has spare capacity. - **BestEffort** — no container sets any request or limit. The pod is entitled to nothing. You can read the assigned class with `kubectl get pod X -o jsonpath='{.status.qosClass}'`. ## What "request" means for eviction A **request** is a reservation: the scheduler only places a pod on a node whose unallocated capacity covers it, and the kubelet treats it as the pod's entitlement. A **limit** is a ceiling enforced by cgroups: exceed the memory limit and the container is `OOMKilled` individually. Eviction ranking is about the request, not the limit. The kubelet asks: *is this pod using more memory than it reserved?* ## The ranking algorithm When a memory signal breaches its threshold, the kubelet builds a list of candidate pods and sorts them by, in order: 1. **Usage relative to request.** Pods whose memory working set is *above* their memory request rank ahead of pods below their request. This is the dominant term. 2. **Pod priority.** Among pods in the same over/under bucket, the one with the lower `priorityClassName` value is evicted first. This is how you protect critical infrastructure DaemonSets — give them a high PriorityClass such as `system-node-critical`. 3. **Magnitude of the overage.** Among equal-priority pods that are all over request, the one exceeding its request by the most is evicted first. Because a BestEffort pod has a memory request of zero, *any* memory it uses puts it over request, so BestEffort pods land at the head of the list every time. A Burstable pod that requested 256Mi and is using 900Mi is next. A Burstable pod using 100Mi against a 256Mi request, and a Guaranteed pod (which by construction cannot exceed its request without hitting its equal limit and being OOMKilled first), sit at the bottom. So the familiar shorthand — *BestEffort first, then Burstable, then Guaranteed* — is a **consequence** of the ranking rather than the rule itself. Saying the rule in terms of usage-versus-request is the better answer, because it explains why a Burstable pod comfortably inside its request survives while a greedier one does not. ## Guaranteed is not immunity If reclaiming from every over-request pod still leaves the node below its threshold, the kubelet keeps going into Guaranteed pods. The class buys you an ordering advantage, not an exemption. Similarly, high pod priority delays but does not prevent node-pressure eviction — this is a well-known asymmetry with API-initiated eviction (the `/eviction` subresource used by `kubectl drain`), which *does* respect PodDisruptionBudgets. **Node-pressure eviction ignores PodDisruptionBudgets and does not honour graceful termination on hard thresholds.** ## The kernel side: OOM scores The same classes shape behaviour one level down. The kubelet sets each container's `oom_score_adj`: Guaranteed containers get −997 (essentially never chosen), BestEffort get 1000 (chosen first), and Burstable get a value scaled inversely to their memory request — the smaller the request relative to node capacity, the higher the score. So if the kernel OOM killer fires before the kubelet can act, victim selection follows the same philosophy. ## Practical consequences - **Set memory requests close to real steady-state usage.** That single act moves a pod from "always over request" to "under request" and is the cheapest reliability win available. - **Do not run production workloads BestEffort.** They are the designated victims. - **Use PriorityClasses deliberately** for logging/monitoring agents and other node-critical DaemonSets. - **Guaranteed for latency-critical or stateful pods**, accepting the cost that limit equals request means no bursting and lower packing density. ## Interview framing State the three classes and how they are derived, give the ranking as usage-vs-request → priority → magnitude of overage, note that BestEffort-first falls out of that, and finish with the caveat that Guaranteed is only last in line, not exempt, and that node-pressure eviction ignores PodDisruptionBudgets.
- Does a Guaranteed QoS class make a pod immune from node-pressure eviction?No. Guaranteed only places the pod last in the eviction ranking. If the kubelet has already reclaimed from every BestEffort and over-request Burstable pod and the node is still below its threshold, it will evict Guaranteed pods too. The class is an ordering advantage, not an exemption.
- Does node-pressure eviction respect PodDisruptionBudgets?No. PodDisruptionBudgets constrain *API-initiated* eviction — the `/eviction` subresource used by `kubectl drain` and by cluster autoscalers. Node-pressure eviction is the kubelet acting unilaterally to save the node, so it bypasses PDBs entirely and, on a hard threshold, also bypasses the pod's termination grace period.
saying these in an interview costs you the question
- Stating "BestEffort → Burstable → Guaranteed" as the rule without knowing it derives from usage-versus-request
- Claiming Guaranteed pods can never be evicted
- Thinking limits, not requests, drive eviction ranking
- Assuming PodDisruptionBudgets protect pods from kubelet node-pressure eviction
- Believing a high PriorityClass fully exempts a pod rather than reordering it