skip to content

Cluster Autoscaler grew a 140-node Kubernetes cluster for a burst, but a week later nodes sit mostly idle and none are removed. How do you find what blocks scale-down?

level: seniorimportance: must knowfreq 58%

answer

  1. running, candidate, movable
  2. requests, not usage; the larger ratio
  3. delay after every scale-up
  4. status ConfigMap and unremovable reasons
  5. ALLOWED DISRUPTIONS zero pins nodes

basics

~20 s

Check that scale-down is running, then that each node is really under the 0.5 threshold, which is calculated from requests rather than usage. Then find the pod pinning it: local storage, unbudgeted kube-system pods, bare pods, safe-to-evict false, or a zero-disruption PDB.

solid answer

~40 s

I work down three layers. First, is scale-down running? Look for `--scale-down-enabled`, groups at minimum size, and constant scale-ups resetting the 10-minute `--scale-down-delay-after-add`. The `cluster-autoscaler-status` ConfigMap and the autoscaler's logs say why each candidate was rejected. Second, is the node really under `--scale-down-utilization-threshold` (0.5)? Utilisation is the higher of the CPU-request and memory-request ratios against allocatable, so an idle-looking node with 57% of its memory requested is not a candidate. Third, which pod pins it? The usual culprits are a disk-backed `emptyDir` or `hostPath` volume, an unbudgeted non-DaemonSet `kube-system` pod, a pod with no controller, `safe-to-evict: "false"`, a PDB with zero allowed disruptions, or a pod that fits on no other node. Then I fix that cause rather than lowering the threshold.

code

yaml · 10 lines
yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: coredns
  namespace: kube-system
spec:
  maxUnavailable: 1
  selector:
    matchLabels:
      k8s-app: kube-dns

go deeper

for a junior

Recall that Cluster Autoscaler removes nodes based on requested resources, and only when every pod on the node can move somewhere else.

for a middle

Explain the max-of-ratios utilisation formula, the 10-minute unneeded and after-add delays, and each default drainability rule.

for a senior

Show a methodical diagnosis: read the status ConfigMap and the unremovable reasons, check PDB allowed disruptions, then fix the pinning pod instead of tuning thresholds.

for a principal

Turn recurring blockers into platform policy, such as mandatory budgets and request right-sizing, and report stranded node-hours per team.

## First principle: scale-down is a simulation over requests **Cluster Autoscaler** removes a node only when three things hold for long enough: 1. The node's **utilisation is below the threshold**. Utilisation is the *higher* of two ratios, the sum of the pods' CPU **requests** over the node's allocatable CPU and the sum of their memory requests over allocatable memory. The default `--scale-down-utilization-threshold` is `0.5`. Live usage never enters this calculation. 2. **Every pod on the node can move**: the drainability rules allow it, and a simulated reschedule finds another node that satisfies the pod's requests, affinity, taints and topology rules. 3. The node has been unneeded for `--scale-down-unneeded-time` (10 minutes by default), and no global delay is active. A cluster that "never scales down" is failing one of these, and the autoscaler records which one. ## Step 1: is scale-down running at all? - **Flags.** `--scale-down-enabled=false` turns scale-down off; the flag is deprecated, but current source still honours it. The node-count floor also matters: a node group already at its minimum size never shrinks. - **Delays.** Every scale-up pauses scale-down evaluation for `--scale-down-delay-after-add` (10 minutes by default), and a failed scale-down pauses it for `--scale-down-delay-after-failure` (3 minutes). A cluster that grows every few minutes during a long burst never reaches a scale-down pass. - **Status.** The autoscaler writes a summary to the `cluster-autoscaler-status` ConfigMap in its own namespace, usually `kube-system`, and logs why each candidate node cannot be removed. ## Step 2: is the node really under the threshold? "20% utilised" on a dashboard usually means CPU usage. The autoscaler looks at requests, and it takes the larger ratio. Take one node from the 140-node cluster: | Resource | Requested | Allocatable | Ratio | |---|---|---|---| | CPU | 1,582m | 7,910m | 0.20 | | Memory | 16,245Mi | 28,500Mi | 0.57 | The utilisation is `max(0.20, 0.57) = 0.57`, which is above `0.5`, so the node is not a candidate at all. The fix is right-sizing memory requests, not tuning the autoscaler. DaemonSet pods count toward utilisation too, unless `--ignore-daemonsets-utilization` is set. ## Step 3: which pod pins the node? The default drainability rules block a node when any of its pods: - uses **local storage**, meaning `hostPath` or a disk-backed `emptyDir` (`--skip-nodes-with-local-storage`, default `true`); - is a non-DaemonSet, non-mirror pod in **`kube-system`** that no PodDisruptionBudget covers (`--skip-nodes-with-system-pods`, default `true`). Current Cluster Autoscaler source releases such a pod once it is older than `--blocking-system-pod-distruption-timeout` (one hour by default, and the flag really is spelled that way); older releases treated it as a permanent block; - has **no controller**, or a controller kind the autoscaler does not recognise; - carries `cluster-autoscaler.kubernetes.io/safe-to-evict: "false"`; - is covered by a **PodDisruptionBudget** whose `status.disruptionsAllowed` is below 1; - **fits nowhere else**: required anti-affinity, a `nodeSelector` that only this node satisfies, or a PersistentVolume bound to this node's zone. A node can also be excluded outright with the node annotation `cluster-autoscaler.kubernetes.io/scale-down-disabled: "true"`. ```bash kubectl -n kube-system get configmap cluster-autoscaler-status -o yaml kubectl get pdb -A kubectl get pods -A -o wide --field-selector spec.nodeName=general-b-7x2kq kubectl describe node general-b-7x2kq | grep -A8 'Allocated resources' ``` In `kubectl get pdb -A`, look for an `ALLOWED DISRUPTIONS` of `0` on budgets whose pods are spread across the stuck nodes. One `maxUnavailable: 0` budget on a 1,180-pod namespace can pin dozens of nodes. ## Step 4: fix the cause, not the symptom | Cause | Fix | |---|---| | Requests far above usage | Right-size requests; the autoscaler follows requests | | Disposable local scratch | `safe-to-evict-local-volumes` naming the volume | | Unbudgeted `kube-system` add-on | Give it a PDB that allows one disruption | | Budget with zero allowed disruptions | Allow at least one, or add replicas | | Pods that fit nowhere else | Relax required rules, or accept the node | | Constant scale-ups | Smooth the burst, or accept the delay | Lowering `--scale-down-utilization-threshold` or shortening `--scale-down-unneeded-time` does not help when a pod blocks the node. It only makes the nodes that can already move churn faster. After the fix, watch the next few scale-down passes. Pods on a removed node get `ScaleDown` events, and a failed attempt leaves a `ScaleDownFailed` warning on the node or pod. The node count should fall in steps as the unneeded time passes for each node, rather than all at once.

  • Why does adding a PodDisruptionBudget to a `kube-system` Deployment make its nodes removable by Cluster Autoscaler?
    The system-pod rule only blocks non-DaemonSet, non-mirror `kube-system` pods that no budget covers. An unbudgeted system pod gives the autoscaler no signal about how much disruption the add-on tolerates, so it plays safe. Once a PDB covers the pod, the autoscaler moves it within that budget, provided the budget allows at least one disruption.
  • Would lowering `--scale-down-utilization-threshold` from 0.5 to 0.3 fix a Cluster Autoscaler that never removes nodes?
    Only if the nodes are above the threshold and nothing pins them. Lowering it makes *fewer* nodes candidates, so it is the wrong direction for requests that look too high; raising it widens the net. It does nothing for nodes pinned by local storage, budgets or unmovable pods, which is the common case. Find the recorded reason first.
  • In a Kubernetes cluster with heavy DaemonSet agents, nodes with no workload stay above the Cluster Autoscaler threshold. Which setting is relevant?
    `--ignore-daemonsets-utilization`. By default, DaemonSet pods' requests count toward a node's utilisation, so a node carrying heavy agents can stay above the threshold even with no workload on it. Setting the flag to `true` leaves them out of the calculation. DaemonSet pods never block a drain, because they are not rescheduled elsewhere.

saying these in an interview costs you the question

  • Cluster Autoscaler scales down when CPU usage from metrics-server drops
  • Utilisation is the average of the CPU and memory ratios
  • Lowering the utilisation threshold makes more nodes removable
  • DaemonSet pods are what usually block scale-down
  • A PodDisruptionBudget can never block the autoscaler, only kubectl drain
  • Scale-down resumes immediately after each scale-up finishes