skip to content

When several Cluster Autoscaler node groups could each fit a Pending Kubernetes pod, how does it pick one to grow, and what do its expanders optimise?

level: middleimportance: nice to knowfreq 31%

answer

  1. a template node per group
  2. options, then one filter picks
  3. default changed from random
  4. comma-separated chain, random tiebreak
  5. unmatched regex means never grown

basics

~20 s

Cluster Autoscaler simulates the Pending pods against a template node for each node group, then an expander picks among the groups that fit. The default expander, least-waste, picks the group that leaves the least CPU and memory unused.

solid answer

~40 s

For each node group, Cluster Autoscaler builds a template node and runs the scheduler's filters to see which Pending pods would fit and how many nodes it would need. The groups that pass are the expansion options, and `--expander` picks one. The default is `least-waste`, which minimises the unused fraction of requested CPU and memory. `most-pods` maximises pods placed, `least-nodes` minimises nodes added, `price` uses provider pricing where it exists, and `priority` ranks groups by regex from the `cluster-autoscaler-priority-expander` ConfigMap. The flag takes a comma-separated chain, such as `priority,least-waste`, where each expander filters the survivors of the one before and any final tie is broken at random. With `priority`, a group that matches no entry is not used at all.

code

bash · 5 lines
bash
cluster-autoscaler \
  --cloud-provider=clusterapi \
  --expander=priority,least-waste \
  --balance-similar-node-groups=true \
  --max-node-provision-time=15m

go deeper

for a junior

Recall that an expander only chooses which node group grows once Cluster Autoscaler has already decided to add nodes, and that least-waste is the default.

for a middle

Explain template-node simulation, how least-waste scores options from requests, and how a comma-separated chain filters options down to one.

for a senior

Know the priority ConfigMap's traps, especially unmatched groups being skipped, and pair it with per-zone node groups and similar-group balancing.

for a principal

Use expander ordering to encode cost policy, such as spot before on-demand, and weigh that against the capacity you lose when spot is unavailable.

## Where the choice happens **Cluster Autoscaler** reacts to pods that kube-scheduler has marked unschedulable. It does not watch CPU or memory usage. Every scan (`--scan-interval`, 10 seconds by default) it takes the Pending pods and asks, for each **node group** (a set of identical machines that the cloud provider scales as one unit), "if I added nodes to this group, which of these pods would fit, and how many nodes would I need?" It cannot ask a node that does not exist yet, so it builds a **template node** for each group, taken from an existing node in the group or from the group's configuration. It then runs the scheduler's filtering logic against that template. Groups where no pending pod would fit, or that are already at their maximum size, drop out. What remains is a list of **expansion options**, each one being "grow group X by N nodes to place pods P". The **expander** picks one option from that list. ## The expanders The `--expander` flag names the strategy. Its default is **`least-waste`**; older releases defaulted to `random`, which is why older write-ups still say so. | Expander | Picks the option that... | |---|---| | `least-waste` | leaves the smallest fraction of CPU plus memory unused after placing the pods | | `most-pods` | places the most pending pods | | `least-nodes` | needs the fewest new nodes | | `random` | is drawn at random | | `price` | is cheapest, where the cloud provider integration supplies pricing | | `priority` | belongs to the highest-priority group in a ConfigMap | | `grpc` | an external service chooses, over gRPC | `least-waste` works per option: it adds up the pods' CPU and memory **requests**, divides by the capacity of the new nodes, and scores the unused fraction of each. Requests are what count here; actual usage never enters the calculation. ## Chaining expanders The flag accepts a **comma-separated list**, and the expanders run in order as filters. Each one narrows the set of options, and the next one only sees the survivors. If more than one option is still left at the end, the tie is broken at random. A common chain is `--expander=priority,least-waste`: the ConfigMap ranks families of node groups, and `least-waste` chooses within the top family. ## The priority expander The `priority` expander reads a ConfigMap named **`cluster-autoscaler-priority-expander`**, from the key `priorities`. The value maps integer priorities to lists of **regular expressions** that are matched against node group names: 1. Each option's group name is matched against every list. 2. The highest priority with a match wins, and all options at that priority survive. 3. A group that matches no entry is **not used** for that scale-up, and the autoscaler records a warning about it. That third rule catches people out. A group added later with a name nobody put in the ConfigMap will never grow. ```yaml apiVersion: v1 kind: ConfigMap metadata: name: cluster-autoscaler-priority-expander namespace: kube-system data: priorities: |- 60: - .*dispatch-spot-.* 20: - .*dispatch-ondemand-.* 1: - .* ``` The catch-all `.*` at priority 1 keeps unlisted groups usable as a last resort. ## Worked example: a webhook-delivery dispatcher A 140-node cluster shared by 22 teams has three node groups the webhook-delivery dispatcher could use: 8-vCPU general nodes, 16-vCPU general nodes, and 16-vCPU spot nodes. A burst leaves 37 dispatcher pods Pending, each requesting 350m CPU and 512Mi memory. - **`most-pods`** sees that every group can place all 37 pods, so it cannot choose and the tie is broken at random. - **`least-nodes`** prefers the 16-vCPU groups, which need fewer nodes. - **`least-waste`** compares the unused fraction of the new capacity and prefers whichever shape the 12,950m of CPU and about 18.5Gi of memory fill most tightly. - **`priority,least-waste`** sends the burst to the spot groups first, as the team intended. ## What an expander does not control - **When** a scale-up happens: that depends on pods being unschedulable. - **How fast** a node arrives: the cloud API, boot, kubelet registration and image pulls take minutes. If a node does not register within `--max-node-provision-time` (15 minutes by default), the group is backed off, starting at `--initial-node-group-backoff-duration` (5 minutes). - **Zone balance**: that is `--balance-similar-node-groups`, which is off by default.

  • A Pending pod needs capacity in one specific zone, but your Cluster Autoscaler node group spans three zones. What goes wrong, and what layout avoids it?
    The autoscaler simulates the group from a single template node, so it may decide the pod fits, and then the cloud provider creates the node in a different zone, where the pod still cannot run. Use one node group per zone, and enable `--balance-similar-node-groups` so the autoscaler spreads growth across those similar groups. That gives zone-spread workloads nodes in every zone.
  • Why can a pod stay Pending for several minutes even after Cluster Autoscaler has triggered a scale-up?
    Triggering is fast: a scan runs every 10 seconds by default. The node, though, has to be created by the cloud API, boot, register its kubelet, become Ready, start its DaemonSets and pull the pod's image, and that usually takes minutes. If it has not registered within `--max-node-provision-time` (15 minutes by default), the autoscaler backs the group off and tries other options.

saying these in an interview costs you the question

  • The default Cluster Autoscaler expander is still random
  • least-waste compares live CPU usage from metrics-server
  • With the priority expander, unlisted node groups get priority zero
  • Listing several expanders makes the autoscaler grow several groups at once
  • The expander decides when a scale-up is triggered