Cluster Autoscaler grew a 140-node Kubernetes cluster for a burst, but a week later nodes sit mostly idle and none are removed. How do you find what blocks scale-down?
answer
- running, candidate, movable
- requests, not usage; the larger ratio
- delay after every scale-up
- status ConfigMap and unremovable reasons
- ALLOWED DISRUPTIONS zero pins nodes
basics
~20 sCheck that scale-down is running, then that each node is really under the 0.5 threshold, which is calculated from requests rather than usage. Then find the pod pinning it: local storage, unbudgeted kube-system pods, bare pods, safe-to-evict false, or a zero-disruption PDB.
solid answer
~40 sI work down three layers. First, is scale-down running? Look for `--scale-down-enabled`, groups at minimum size, and constant scale-ups resetting the 10-minute `--scale-down-delay-after-add`. The `cluster-autoscaler-status` ConfigMap and the autoscaler's logs say why each candidate was rejected. Second, is the node really under `--scale-down-utilization-threshold` (0.5)? Utilisation is the higher of the CPU-request and memory-request ratios against allocatable, so an idle-looking node with 57% of its memory requested is not a candidate. Third, which pod pins it? The usual culprits are a disk-backed `emptyDir` or `hostPath` volume, an unbudgeted non-DaemonSet `kube-system` pod, a pod with no controller, `safe-to-evict: "false"`, a PDB with zero allowed disruptions, or a pod that fits on no other node. Then I fix that cause rather than lowering the threshold.
code
yaml · 10 linesapiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: coredns
namespace: kube-system
spec:
maxUnavailable: 1
selector:
matchLabels:
k8s-app: kube-dnsgo deeper
Recall that Cluster Autoscaler removes nodes based on requested resources, and only when every pod on the node can move somewhere else.
Explain the max-of-ratios utilisation formula, the 10-minute unneeded and after-add delays, and each default drainability rule.
Show a methodical diagnosis: read the status ConfigMap and the unremovable reasons, check PDB allowed disruptions, then fix the pinning pod instead of tuning thresholds.
Turn recurring blockers into platform policy, such as mandatory budgets and request right-sizing, and report stranded node-hours per team.
## First principle: scale-down is a simulation over requests **Cluster Autoscaler** removes a node only when three things hold for long enough: 1. The node's **utilisation is below the threshold**. Utilisation is the *higher* of two ratios, the sum of the pods' CPU **requests** over the node's allocatable CPU and the sum of their memory requests over allocatable memory. The default `--scale-down-utilization-threshold` is `0.5`. Live usage never enters this calculation. 2. **Every pod on the node can move**: the drainability rules allow it, and a simulated reschedule finds another node that satisfies the pod's requests, affinity, taints and topology rules. 3. The node has been unneeded for `--scale-down-unneeded-time` (10 minutes by default), and no global delay is active. A cluster that "never scales down" is failing one of these, and the autoscaler records which one. ## Step 1: is scale-down running at all? - **Flags.** `--scale-down-enabled=false` turns scale-down off; the flag is deprecated, but current source still honours it. The node-count floor also matters: a node group already at its minimum size never shrinks. - **Delays.** Every scale-up pauses scale-down evaluation for `--scale-down-delay-after-add` (10 minutes by default), and a failed scale-down pauses it for `--scale-down-delay-after-failure` (3 minutes). A cluster that grows every few minutes during a long burst never reaches a scale-down pass. - **Status.** The autoscaler writes a summary to the `cluster-autoscaler-status` ConfigMap in its own namespace, usually `kube-system`, and logs why each candidate node cannot be removed. ## Step 2: is the node really under the threshold? "20% utilised" on a dashboard usually means CPU usage. The autoscaler looks at requests, and it takes the larger ratio. Take one node from the 140-node cluster: | Resource | Requested | Allocatable | Ratio | |---|---|---|---| | CPU | 1,582m | 7,910m | 0.20 | | Memory | 16,245Mi | 28,500Mi | 0.57 | The utilisation is `max(0.20, 0.57) = 0.57`, which is above `0.5`, so the node is not a candidate at all. The fix is right-sizing memory requests, not tuning the autoscaler. DaemonSet pods count toward utilisation too, unless `--ignore-daemonsets-utilization` is set. ## Step 3: which pod pins the node? The default drainability rules block a node when any of its pods: - uses **local storage**, meaning `hostPath` or a disk-backed `emptyDir` (`--skip-nodes-with-local-storage`, default `true`); - is a non-DaemonSet, non-mirror pod in **`kube-system`** that no PodDisruptionBudget covers (`--skip-nodes-with-system-pods`, default `true`). Current Cluster Autoscaler source releases such a pod once it is older than `--blocking-system-pod-distruption-timeout` (one hour by default, and the flag really is spelled that way); older releases treated it as a permanent block; - has **no controller**, or a controller kind the autoscaler does not recognise; - carries `cluster-autoscaler.kubernetes.io/safe-to-evict: "false"`; - is covered by a **PodDisruptionBudget** whose `status.disruptionsAllowed` is below 1; - **fits nowhere else**: required anti-affinity, a `nodeSelector` that only this node satisfies, or a PersistentVolume bound to this node's zone. A node can also be excluded outright with the node annotation `cluster-autoscaler.kubernetes.io/scale-down-disabled: "true"`. ```bash kubectl -n kube-system get configmap cluster-autoscaler-status -o yaml kubectl get pdb -A kubectl get pods -A -o wide --field-selector spec.nodeName=general-b-7x2kq kubectl describe node general-b-7x2kq | grep -A8 'Allocated resources' ``` In `kubectl get pdb -A`, look for an `ALLOWED DISRUPTIONS` of `0` on budgets whose pods are spread across the stuck nodes. One `maxUnavailable: 0` budget on a 1,180-pod namespace can pin dozens of nodes. ## Step 4: fix the cause, not the symptom | Cause | Fix | |---|---| | Requests far above usage | Right-size requests; the autoscaler follows requests | | Disposable local scratch | `safe-to-evict-local-volumes` naming the volume | | Unbudgeted `kube-system` add-on | Give it a PDB that allows one disruption | | Budget with zero allowed disruptions | Allow at least one, or add replicas | | Pods that fit nowhere else | Relax required rules, or accept the node | | Constant scale-ups | Smooth the burst, or accept the delay | Lowering `--scale-down-utilization-threshold` or shortening `--scale-down-unneeded-time` does not help when a pod blocks the node. It only makes the nodes that can already move churn faster. After the fix, watch the next few scale-down passes. Pods on a removed node get `ScaleDown` events, and a failed attempt leaves a `ScaleDownFailed` warning on the node or pod. The node count should fall in steps as the unneeded time passes for each node, rather than all at once.
- Why does adding a PodDisruptionBudget to a `kube-system` Deployment make its nodes removable by Cluster Autoscaler?The system-pod rule only blocks non-DaemonSet, non-mirror `kube-system` pods that no budget covers. An unbudgeted system pod gives the autoscaler no signal about how much disruption the add-on tolerates, so it plays safe. Once a PDB covers the pod, the autoscaler moves it within that budget, provided the budget allows at least one disruption.
- Would lowering `--scale-down-utilization-threshold` from 0.5 to 0.3 fix a Cluster Autoscaler that never removes nodes?Only if the nodes are above the threshold and nothing pins them. Lowering it makes *fewer* nodes candidates, so it is the wrong direction for requests that look too high; raising it widens the net. It does nothing for nodes pinned by local storage, budgets or unmovable pods, which is the common case. Find the recorded reason first.
- In a Kubernetes cluster with heavy DaemonSet agents, nodes with no workload stay above the Cluster Autoscaler threshold. Which setting is relevant?`--ignore-daemonsets-utilization`. By default, DaemonSet pods' requests count toward a node's utilisation, so a node carrying heavy agents can stay above the threshold even with no workload on it. Setting the flag to `true` leaves them out of the calculation. DaemonSet pods never block a drain, because they are not rescheduled elsewhere.
saying these in an interview costs you the question
- Cluster Autoscaler scales down when CPU usage from metrics-server drops
- Utilisation is the average of the CPU and memory ratios
- Lowering the utilisation threshold makes more nodes removable
- DaemonSet pods are what usually block scale-down
- A PodDisruptionBudget can never block the autoscaler, only kubectl drain
- Scale-down resumes immediately after each scale-up finishes