skip to content

Kubernetes assigns every Pod a Quality of Service class of Guaranteed, Burstable, or BestEffort. How is that class derived from the Pod spec, and where does it actually change behaviour?

level: middleimportance: must knowfreq 62%

answer

  1. Guaranteed = every container, both resources, request == limit
  2. BestEffort = nothing set anywhere
  3. Burstable = everything else
  4. eviction order: BestEffort → Burstable → Guaranteed
  5. oom_score_adj 1000 vs -997; static CPU pinning needs Guaranteed

basics

~20 s

Guaranteed: every container sets CPU and memory requests and limits, and request equals limit. BestEffort: no container sets any request or limit. Anything in between is Burstable. The class drives eviction order under node pressure and OOM kill priority.

solid answer

~50 s

QoS is derived, not declared — the kubelet computes it from the spec and writes it to `status.qosClass`. - **Guaranteed**: every container (including init containers) sets both CPU and memory requests *and* limits, with request equal to limit for each. - **BestEffort**: no container sets any request or limit at all. - **Burstable**: everything else — at least one request or limit somewhere, but not fully matched. It matters in three places. First, **node-pressure eviction**: the kubelet evicts BestEffort Pods first, then Burstable Pods exceeding their requests, and Guaranteed Pods last. Second, **kernel OOM ranking**: the kubelet writes `oom_score_adj` per class — roughly 1000 for BestEffort, -997 for Guaranteed, and a computed middle value for Burstable — so under node-level memory pressure the kernel kills low-QoS processes first. Third, **exclusive CPU pinning**: the CPU Manager's `static` policy only grants dedicated cores to Guaranteed Pods whose CPU request is a whole integer. It is a per-Pod property, not per-container.

code

yaml · 11 lines
yaml
spec:
  containers:
    - name: app
      image: example/app:2.0
      resources:
        requests:
          cpu: "2"
          memory: "2Gi"
        limits:
          cpu: "2"
          memory: "2Gi"

go deeper

for a junior

Recall the three class names and the rule that derives each one from requests and limits.

for a middle

Explain the derivation precisely, including that it covers every container in the Pod, and name eviction ordering as the main consequence.

for a senior

Connect it to incident behaviour — eviction ranking versus kernel oom_score_adj, and why sidecar injection quietly demotes Pods.

for a principal

Treat QoS as a platform policy lever: which tiers of workload may be BestEffort, how defaults are enforced cluster-wide, and where Guaranteed plus CPU pinning is worth the packing loss.

## QoS is computed, not configured There is no `qosClass` field you can set. The kubelet inspects the Pod's containers when it admits the Pod and stamps the result into `status.qosClass`. Because it is derived, the only way to change a Pod's class is to change its resource fields — and since resource fields on a running Pod are largely immutable (in-place resize being a newer, still-maturing capability), the class is effectively fixed for the Pod's life. ## The three classes **Guaranteed** requires every container in the Pod — regular containers and init containers alike — to specify both CPU and memory in both `requests` and `limits`, with the request equal to the limit for each resource. One container missing a CPU limit, or one where request is 500m and limit is 1, drops the whole Pod to Burstable. Setting only limits still qualifies, because Kubernetes defaults the missing request to the limit. **BestEffort** is the opposite extreme: not a single container declares any request or limit for CPU or memory. These Pods are scheduled as if they were free, and can be placed on already-full nodes. **Burstable** is everything in the middle: at least one container declares at least one request or limit, but the Pod does not meet the Guaranteed bar. This is where most real workloads land — for example memory request equal to memory limit but a CPU request below the CPU limit. ## Where the class changes behaviour **Node-pressure eviction.** When the kubelet detects a resource crossing its eviction threshold — most commonly `memory.available` or disk — it ranks Pods and evicts until the signal clears. The ranking considers whether the Pod's usage exceeds its requests, then Pod priority, then how far above the request usage sits. In practice that means BestEffort Pods (which by definition exceed a zero request the moment they allocate anything) go first, Burstable Pods over their request next, and Guaranteed Pods — which cannot exceed their request without hitting their own limit first — last. Eviction is graceful: the Pod is terminated and, if owned by a controller, rescheduled elsewhere. **Kernel OOM ranking.** Eviction is the kubelet acting ahead of trouble. If memory disappears faster than the kubelet reacts, the kernel OOM killer fires instead, and it chooses by `oom_score_adj`. The kubelet writes that value per container based on QoS: BestEffort gets 1000 (most attractive victim), Guaranteed gets -997 (near-immune, alongside critical system daemons), and Burstable gets a value derived from its memory request relative to node capacity — the smaller the request relative to the node, the more killable. This is why a well-sized Guaranteed Pod usually survives a node-level memory crisis while a BestEffort neighbour disappears. **CPU and memory pinning.** The kubelet's CPU Manager `static` policy hands out exclusive, pinned cores — but only to Guaranteed Pods requesting an integer number of CPUs. A Guaranteed Pod requesting `1500m` does not qualify. The Memory Manager and Topology Manager have similar Guaranteed-only behaviours for NUMA-aligned allocation. If you are chasing low, predictable tail latency, Guaranteed with integer CPU is the entry ticket. ## Common design consequences Because QoS is per Pod and requires *every* container to comply, a sidecar with no resources set silently demotes an otherwise carefully-specified application Pod — a frequent real-world surprise once a service mesh or log shipper is injected. Platform teams usually solve this with a LimitRange supplying defaults so nothing is accidentally BestEffort, or with a mutating admission policy that stamps resources onto injected sidecars. Also note what QoS does *not* do: it does not affect scheduling order, it does not drive preemption (that is PriorityClass, a separate mechanism), and it grants no extra CPU. A Guaranteed Pod is not faster than a Burstable Pod using the same amount of CPU; it is merely harder to evict and more predictable.

  • A Pod is carefully configured as Guaranteed, but after a service mesh sidecar is injected `status.qosClass` reads Burstable. Why?
    QoS is a property of the whole Pod, and every container must satisfy the criteria. The injected sidecar has no requests and limits, or has a request below its limit, so the Pod is demoted. The fix is to configure resources on the injected sidecar, or to supply namespace defaults with a LimitRange so no container is ever resource-less.
  • Does a Guaranteed Pod get more CPU than a Burstable one?
    No. QoS grants no additional CPU time; scheduling and CPU shares still come from the requests. Guaranteed only means the Pod is last in line for eviction, near-immune to the kernel OOM killer, and eligible for exclusive CPU pinning under the static CPU Manager policy.

saying these in an interview costs you the question

  • Believing QoS is a field you set on the Pod rather than something the kubelet derives.
  • Thinking QoS is per container — one non-compliant sidecar changes the whole Pod's class.
  • Confusing QoS with PriorityClass; preemption and scheduling order come from priority, not QoS.
  • Claiming Guaranteed Pods are never evicted — they still go if they exceed their own requests or on disk/PID pressure.
  • Saying a Pod with only limits and no requests must be Burstable; missing requests default to the limits, so it can still be Guaranteed.

context