skip to content

Node Autoscaling

When no node fits a Pending pod, Cluster Autoscaler grows a node group while Karpenter provisions an instance shaped to the pod itself; both shrink again when nodes sit underused. The interview question is usually why the cluster never scales back down.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

5

Cluster Autoscaler grew a 140-node Kubernetes cluster for a burst, but a week later nodes sit mostly idle and none are removed. How do you find what blocks scale-down?

level: seniorimportance: must knowfreq 58%

answer

  1. running, candidate, movable
  2. requests, not usage; the larger ratio
  3. delay after every scale-up
  4. status ConfigMap and unremovable reasons
  5. ALLOWED DISRUPTIONS zero pins nodes

basics

~20 s

Check that scale-down is running, then that each node is really under the 0.5 threshold, which is calculated from requests rather than usage. Then find the pod pinning it: local storage, unbudgeted kube-system pods, bare pods, safe-to-evict false, or a zero-disruption PDB.

solid answer

~40 s

I work down three layers. First, is scale-down running? Look for `--scale-down-enabled`, groups at minimum size, and constant scale-ups resetting the 10-minute `--scale-down-delay-after-add`. The `cluster-autoscaler-status` ConfigMap and the autoscaler's logs say why each candidate was rejected. Second, is the node really under `--scale-down-utilization-threshold` (0.5)? Utilisation is the higher of the CPU-request and memory-request ratios against allocatable, so an idle-looking node with 57% of its memory requested is not a candidate. Third, which pod pins it? The usual culprits are a disk-backed `emptyDir` or `hostPath` volume, an unbudgeted non-DaemonSet `kube-system` pod, a pod with no controller, `safe-to-evict: "false"`, a PDB with zero allowed disruptions, or a pod that fits on no other node. Then I fix that cause rather than lowering the threshold.

code

yaml · 10 lines
yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: coredns
  namespace: kube-system
spec:
  maxUnavailable: 1
  selector:
    matchLabels:
      k8s-app: kube-dns

go deeper

for a junior

Recall that Cluster Autoscaler removes nodes based on requested resources, and only when every pod on the node can move somewhere else.

for a middle

Explain the max-of-ratios utilisation formula, the 10-minute unneeded and after-add delays, and each default drainability rule.

for a senior

Show a methodical diagnosis: read the status ConfigMap and the unremovable reasons, check PDB allowed disruptions, then fix the pinning pod instead of tuning thresholds.

for a principal

Turn recurring blockers into platform policy, such as mandatory budgets and request right-sizing, and report stranded node-hours per team.

## First principle: scale-down is a simulation over requests **Cluster Autoscaler** removes a node only when three things hold for long enough: 1. The node's **utilisation is below the threshold**. Utilisation is the *higher* of two ratios, the sum of the pods' CPU **requests** over the node's allocatable CPU and the sum of their memory requests over allocatable memory. The default `--scale-down-utilization-threshold` is `0.5`. Live usage never enters this calculation. 2. **Every pod on the node can move**: the drainability rules allow it, and a simulated reschedule finds another node that satisfies the pod's requests, affinity, taints and topology rules. 3. The node has been unneeded for `--scale-down-unneeded-time` (10 minutes by default), and no global delay is active. A cluster that "never scales down" is failing one of these, and the autoscaler records which one. ## Step 1: is scale-down running at all? - **Flags.** `--scale-down-enabled=false` turns scale-down off; the flag is deprecated, but current source still honours it. The node-count floor also matters: a node group already at its minimum size never shrinks. - **Delays.** Every scale-up pauses scale-down evaluation for `--scale-down-delay-after-add` (10 minutes by default), and a failed scale-down pauses it for `--scale-down-delay-after-failure` (3 minutes). A cluster that grows every few minutes during a long burst never reaches a scale-down pass. - **Status.** The autoscaler writes a summary to the `cluster-autoscaler-status` ConfigMap in its own namespace, usually `kube-system`, and logs why each candidate node cannot be removed. ## Step 2: is the node really under the threshold? "20% utilised" on a dashboard usually means CPU usage. The autoscaler looks at requests, and it takes the larger ratio. Take one node from the 140-node cluster: | Resource | Requested | Allocatable | Ratio | |---|---|---|---| | CPU | 1,582m | 7,910m | 0.20 | | Memory | 16,245Mi | 28,500Mi | 0.57 | The utilisation is `max(0.20, 0.57) = 0.57`, which is above `0.5`, so the node is not a candidate at all. The fix is right-sizing memory requests, not tuning the autoscaler. DaemonSet pods count toward utilisation too, unless `--ignore-daemonsets-utilization` is set. ## Step 3: which pod pins the node? The default drainability rules block a node when any of its pods: - uses **local storage**, meaning `hostPath` or a disk-backed `emptyDir` (`--skip-nodes-with-local-storage`, default `true`); - is a non-DaemonSet, non-mirror pod in **`kube-system`** that no PodDisruptionBudget covers (`--skip-nodes-with-system-pods`, default `true`). Current Cluster Autoscaler source releases such a pod once it is older than `--blocking-system-pod-distruption-timeout` (one hour by default, and the flag really is spelled that way); older releases treated it as a permanent block; - has **no controller**, or a controller kind the autoscaler does not recognise; - carries `cluster-autoscaler.kubernetes.io/safe-to-evict: "false"`; - is covered by a **PodDisruptionBudget** whose `status.disruptionsAllowed` is below 1; - **fits nowhere else**: required anti-affinity, a `nodeSelector` that only this node satisfies, or a PersistentVolume bound to this node's zone. A node can also be excluded outright with the node annotation `cluster-autoscaler.kubernetes.io/scale-down-disabled: "true"`. ```bash kubectl -n kube-system get configmap cluster-autoscaler-status -o yaml kubectl get pdb -A kubectl get pods -A -o wide --field-selector spec.nodeName=general-b-7x2kq kubectl describe node general-b-7x2kq | grep -A8 'Allocated resources' ``` In `kubectl get pdb -A`, look for an `ALLOWED DISRUPTIONS` of `0` on budgets whose pods are spread across the stuck nodes. One `maxUnavailable: 0` budget on a 1,180-pod namespace can pin dozens of nodes. ## Step 4: fix the cause, not the symptom | Cause | Fix | |---|---| | Requests far above usage | Right-size requests; the autoscaler follows requests | | Disposable local scratch | `safe-to-evict-local-volumes` naming the volume | | Unbudgeted `kube-system` add-on | Give it a PDB that allows one disruption | | Budget with zero allowed disruptions | Allow at least one, or add replicas | | Pods that fit nowhere else | Relax required rules, or accept the node | | Constant scale-ups | Smooth the burst, or accept the delay | Lowering `--scale-down-utilization-threshold` or shortening `--scale-down-unneeded-time` does not help when a pod blocks the node. It only makes the nodes that can already move churn faster. After the fix, watch the next few scale-down passes. Pods on a removed node get `ScaleDown` events, and a failed attempt leaves a `ScaleDownFailed` warning on the node or pod. The node count should fall in steps as the unneeded time passes for each node, rather than all at once.

  • Why does adding a PodDisruptionBudget to a `kube-system` Deployment make its nodes removable by Cluster Autoscaler?
    The system-pod rule only blocks non-DaemonSet, non-mirror `kube-system` pods that no budget covers. An unbudgeted system pod gives the autoscaler no signal about how much disruption the add-on tolerates, so it plays safe. Once a PDB covers the pod, the autoscaler moves it within that budget, provided the budget allows at least one disruption.
  • Would lowering `--scale-down-utilization-threshold` from 0.5 to 0.3 fix a Cluster Autoscaler that never removes nodes?
    Only if the nodes are above the threshold and nothing pins them. Lowering it makes *fewer* nodes candidates, so it is the wrong direction for requests that look too high; raising it widens the net. It does nothing for nodes pinned by local storage, budgets or unmovable pods, which is the common case. Find the recorded reason first.
  • In a Kubernetes cluster with heavy DaemonSet agents, nodes with no workload stay above the Cluster Autoscaler threshold. Which setting is relevant?
    `--ignore-daemonsets-utilization`. By default, DaemonSet pods' requests count toward a node's utilisation, so a node carrying heavy agents can stay above the threshold even with no workload on it. Setting the flag to `true` leaves them out of the calculation. DaemonSet pods never block a drain, because they are not rescheduled elsewhere.

saying these in an interview costs you the question

  • Cluster Autoscaler scales down when CPU usage from metrics-server drops
  • Utilisation is the average of the CPU and memory ratios
  • Lowering the utilisation threshold makes more nodes removable
  • DaemonSet pods are what usually block scale-down
  • A PodDisruptionBudget can never block the autoscaler, only kubectl drain
  • Scale-down resumes immediately after each scale-up finishes
open as a page

What does the `cluster-autoscaler.kubernetes.io/safe-to-evict` annotation on a Kubernetes pod do, and when would you set it to `true` or `false`?

level: middleimportance: should knowfreq 44%

basics

~10 s

It is a per-pod override for Cluster Autoscaler's scale-down check: true lets the autoscaler remove the pod's node even if the pod would normally block it, and false pins the node against scale-down.

open as a page

Karpenter keeps replacing nodes under a latency-sensitive Kubernetes workload. How do NodePool consolidation and drift work, and how do you limit the disruption they cause?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Consolidation deletes or replaces nodes whose pods fit elsewhere more cheaply, and drift replaces nodes that no longer match their NodePool or NodeClass. Limit both with consolidateAfter, disruption budgets scoped by reason and schedule, PDBs, and the do-not-disrupt annotation.

open as a page

You own node autoscaling for a 140-node Kubernetes cluster shared by 22 product teams. How do you set scale-down policy: its aggressiveness, who may block it, and what that costs?

level: principalimportance: should knowfreq 29%

basics

~20 s

Set scale-down aggressiveness per pool, not globally. Allow vetoes only as reviewed exceptions: budgets that allow a disruption, and time-limited pins on a tainted pool. Report stranded node-hours per team so each veto's cost is visible.

open as a page

When several Cluster Autoscaler node groups could each fit a Pending Kubernetes pod, how does it pick one to grow, and what do its expanders optimise?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

Cluster Autoscaler simulates the Pending pods against a template node for each node group, then an expander picks among the groups that fit. The default expander, least-waste, picks the group that leaves the least CPU and memory unused.

open as a page