skip to content

VPA and Cluster Autoscaler Overview

The Vertical Pod Autoscaler watches real usage and rewrites a pod's requests - recommending only, or evicting the pod to resize it. Interviewers pair it with HPA to ask why both on CPU fight each other, and when recommender-only mode is the safer answer.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

4

How does the Kubernetes Cluster Autoscaler decide to add a node, and what conditions must hold before it removes one?

level: middleimportance: must knowfreq 56%

answer

  1. requests, never usage
  2. scale-up trigger = Pending unschedulable Pods
  3. template node simulation -> expander picks the group
  4. unneeded 10m + utilization < 0.5 -> cordon, drain, delete
  5. blockers: bare Pods, emptyDir, kube-system without PDB, restrictive PDB

basics

~20 s

Scale-up: it watches for Pending Pods the scheduler could not place, simulates whether a new node from some node group would fit them, and grows that group. Scale-down: a node whose requested resources stay under a threshold for a sustained period, and all of whose Pods can move elsewhere, is cordoned, drained, and deleted.

solid answer

~50 s

Cluster Autoscaler runs as a Deployment and works on **resource requests, never actual usage**. **Scale-up** is triggered by unschedulable Pods. When the scheduler cannot place a Pod it stays `Pending` with a `FailedScheduling` event. Each scan interval (default 10s) the autoscaler takes those Pods and, for each node group, simulates a template node from that group: would the Pod become schedulable, respecting its requests, node selectors, affinity, taints and topology? If yes the group is a candidate; an *expander* strategy (random, most-pods, least-waste, price, priority) picks between candidates, and the cloud provider's group size is increased. **Scale-down** is independent. A node is a candidate when the sum of its Pods' requests falls below `--scale-down-utilization-threshold` (default 0.5) and stays there for `--scale-down-unneeded-time` (default 10 minutes), **and** every Pod on it could be rescheduled elsewhere. The node is then cordoned and drained through the Eviction API, honouring disruption budgets and grace periods, and only then deleted.

code

bash · 4 lines
bash
kubectl -n kube-system logs deploy/cluster-autoscaler | grep -E 'scale_down|scale_up|unremovable'
kubectl get pods --field-selector=status.phase=Pending -A
kubectl describe pod <pending-pod> | grep -A10 Events
kubectl get configmap -n kube-system cluster-autoscaler-status -o yaml

go deeper

for a junior

Know that it adds nodes when Pods cannot be scheduled and removes nodes that are not needed, and that it is different from the HPA.

for a middle

Explain the requests-not-usage model, the template-node simulation, the utilization threshold and unneeded window, and the cordon-drain-delete sequence.

for a senior

Diagnose real cases: reading autoscaler logs, the unremovable-node blockers, template label and taint mismatches, expander choice, and interaction with disruption budgets.

for a principal

Reason about node-group topology and instance shapes, headroom via placeholder Pods versus provisioning latency, cost per unit of scheduled capacity, and whether group-based autoscaling or just-in-time provisioning fits the workload mix.

## What it is Cluster Autoscaler (CA) changes the **number of nodes**. It is a separate component from the Horizontal Pod Autoscaler, which changes the number of Pods, and it does not run in the control plane by default — it is a Deployment you install, wired to a cloud provider so it can resize node groups (AWS Auto Scaling Groups, GCP managed instance groups, Azure VM scale sets, and so on). ## The single most important fact CA reasons about **requests**, not usage. A node whose Pods request 100% of its CPU but idle at 3% is "full" and will never be scaled down. A node whose Pods request nothing but are saturating the CPU is "empty" and is a scale-down candidate. Every confusing CA behaviour traces back to this. It follows that CA is only as good as your requests: the fix for "nodes are full but idle" is right-sizing requests, which is exactly the job of vertical pod autoscaling. ## Scale-up in detail 1. The scheduler fails to place a Pod; the Pod sits `Pending` with `PodScheduled=False, reason=Unschedulable`. 2. Every scan interval (`--scan-interval`, default 10s) CA collects unschedulable Pods. 3. For each node group it builds a **template node** — a synthetic Node object from the group's instance type, labels and taints — and runs scheduler predicates to ask whether the Pod would fit. This is why a node group's template must carry the labels and taints the Pod requires: labels applied by a boot script after registration are invisible to the simulation, and the Pod stays Pending forever while CA reports no expansion option. 4. It bin-packs multiple Pending Pods into the simulated node so it can decide how many nodes to add, up to the group's configured maximum. 5. If several groups qualify, the **expander** chooses: `random`, `most-pods`, `least-waste` (smallest leftover capacity), `price`, `priority` (an explicit ranked list you configure). 6. CA increases the group size and waits; `--max-node-provision-time` (default 15 minutes) bounds how long a promised node may take before the attempt is abandoned. A Pod that no template node can satisfy — because it requests more CPU than any instance type has, or has an affinity rule no new empty node can fix — produces the event "pod didn't trigger scale-up", which is CA telling you the problem is the Pod, not the capacity. ## Scale-down in detail Scale-down is deliberately conservative because removing a node is disruptive. A node becomes *unneeded* when its total requested CPU and memory is below `--scale-down-utilization-threshold` (default 0.5) **and** CA can simulate every one of its Pods being placed on other existing nodes. It must remain unneeded continuously for `--scale-down-unneeded-time` (default 10 minutes). Additional cooldowns prevent flapping: `--scale-down-delay-after-add` (default 10 minutes), plus separate delays after a delete or a failure. Then CA cordons the node (marks it unschedulable), drains it by creating **Eviction** objects for each Pod — which means PodDisruptionBudgets are honoured and a budget with zero allowed disruptions blocks the drain — waits for termination grace periods, and finally asks the cloud provider to delete the instance. `--max-graceful-termination-sec` (default 600) bounds the wait. ## Why a node refuses to go away This is the most common operational question. CA will not remove a node hosting: - Pods with **no controller** (bare Pods created directly), since nothing would recreate them. - Pods using **local storage** — `emptyDir` or `hostPath` — unless annotated `cluster-autoscaler.kubernetes.io/safe-to-evict: "true"`. - **kube-system** Pods that have no PodDisruptionBudget. - Pods whose **PodDisruptionBudget** would be violated by the eviction. - Pods that **cannot be rescheduled elsewhere** because of node selectors, affinity, taints, or simply no free requested capacity. - Anything on a node annotated `cluster-autoscaler.kubernetes.io/scale-down-disabled: "true"`. CA logs its reasoning per node, and that log is the fastest route to the answer. A single DaemonSet-like Pod misconfigured as a plain Deployment with an `emptyDir` can pin an entire underutilised node group indefinitely. ## Operational notes - Node groups should contain **identically shaped** instances; CA assumes any node in a group is interchangeable when it builds templates. - CA respects each group's min and max size; hitting max is a silent ceiling worth alerting on. - It interacts with Pod priority: by default Pods below a configurable priority cutoff do not trigger scale-up, which is how "balloon" placeholder Pods are used to keep warm spare capacity. - An alternative model — provisioning individual right-sized nodes on demand rather than resizing fixed groups — exists and is worth knowing about, but the group-based model above is what "Cluster Autoscaler" means.

  • Every node in a cluster sits at 15% CPU usage, yet Cluster Autoscaler removes nothing. Why?
    Because it measures requested resources, not usage. If the Pods on each node request more than the utilization threshold, typically 50%, the nodes count as needed regardless of how idle they are. The fix is to lower the requests so they reflect real consumption, at which point nodes become genuine scale-down candidates.
  • A Pod stays Pending and the events say it did not trigger a scale-up. What are the likely causes?
    Either no node group's template node could satisfy the Pod, or the matching group is already at its maximum size. Template mismatches are common: the Pod requires a label or tolerates a taint that the group template does not advertise, or it requests more CPU or memory than any instance type in the group provides. Check the autoscaler logs, which name the expansion options it rejected and why.

It is a car-park attendant who counts reserved spaces, not cars. A row where every space is booked stays open even if nobody parked; a row with no bookings is closed even if cars are circling.

saying these in an interview costs you the question

  • Saying Cluster Autoscaler scales on CPU utilization like the HPA does — it scales on Pending Pods and requested resources.
  • Believing it adds Pods, or confusing it with the Horizontal Pod Autoscaler.
  • Assuming an empty-looking node is removed immediately, ignoring the sustained-unneeded window and cooldowns.
  • Forgetting that drain goes through the Eviction API, so a PodDisruptionBudget can block node removal entirely.
  • Expecting a node group to scale up for a Pod whose required labels or taints are not on the group's node template.

context

open as a page

What are the update modes of the Kubernetes Vertical Pod Autoscaler, and what does each one actually do to a running Pod?

level: middleimportance: should knowfreq 44%

basics

~20 s

Off only publishes recommendations for humans to read. Initial applies them when a Pod is created and never again. Auto (and Recreate) additionally evicts running Pods so they are recreated with new resource requests. VPA changes requests per Pod; it never changes replica count.

open as a page

Why is running the Kubernetes Vertical Pod Autoscaler in Auto mode alongside a Horizontal Pod Autoscaler that targets CPU utilization considered unsafe, and what can you do instead?

level: seniorimportance: should knowfreq 40%

basics

~20 s

CPU-based horizontal scaling measures usage divided by the request. VPA moves the request, so the same real load changes the measured utilization and makes the horizontal controller scale in or out for no workload reason. The two form a feedback loop. Split them: VPA on memory only, horizontal scaling on CPU or on custom metrics.

open as a page

When would you choose just-in-time node provisioning such as Karpenter over the node-group-based Kubernetes Cluster Autoscaler, and what do you give up by doing so?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Cluster Autoscaler resizes pre-defined, fixed-shape node groups. Karpenter reads Pending Pods and launches individually chosen instances from a broad type list, then consolidates. Choose it for heterogeneous workloads, faster provisioning and better bin-packing; you give up predictability, a cloud-agnostic component, and stable long-lived nodes.

open as a page