skip to content

You own node autoscaling for a 140-node Kubernetes cluster shared by 22 product teams. How do you set scale-down policy: its aggressiveness, who may block it, and what that costs?

level: principalimportance: should knowfreq 29%

answer

  1. cost versus churn, per pool
  2. four ways to veto a node
  3. budgets must allow one
  4. pins go to a tainted pool
  5. stranded node-hours per team

basics

~20 s

Set scale-down aggressiveness per pool, not globally. Allow vetoes only as reviewed exceptions: budgets that allow a disruption, and time-limited pins on a tainted pool. Report stranded node-hours per team so each veto's cost is visible.

solid answer

~40 s

I treat scale-down as a policy with three axes. Aggressiveness: Cluster Autoscaler's threshold, unneeded time and after-add delay, or Karpenter's `consolidationPolicy`, `consolidateAfter` and `budgets`, set per pool, so bursty latency-critical pools wait longer while batch pools stay tight. Vetoes: `safe-to-evict: "false"`, `karpenter.sh/do-not-disrupt`, zero-disruption PDBs, local storage and unique placement rules all pin nodes. I require budgets to allow at least one disruption, isolate pinned work on a tainted pool, give system add-ons budgets and bound hold time. Cost: I publish stranded node-hours and failed scale-downs per team. The price is quiet-hour spend, slower image rollouts and some fragmentation, and I make those costs explicit rather than chasing zero churn.

go deeper

for a junior

Recall that removing idle nodes saves money but evicts pods, and that some pods and budgets can stop a node from being removed.

for a middle

Explain which Cluster Autoscaler and Karpenter settings control how quickly and how many nodes are removed, and what each blocker is.

for a senior

Show you can find which team's pods pin nodes, and apply budgets, dedicated pools and bounded holds without breaking their workloads.

for a principal

Own the cost-versus-churn tradeoff per pool, govern vetoes through admission policy and showback, and state openly what the policy costs.

## The decision being made In a shared cluster, **node scale-down** is where cost and stability meet. Every node the autoscaler cannot remove is paid for, and every node it does remove evicts someone's pods. The platform owner does not pick a correct answer here. They set a **policy**: how aggressive removal is, who is allowed to veto it, and how the cost of those vetoes becomes visible. Take a 140-node cluster shared by 22 product teams, whose largest tenant is a webhook-delivery dispatcher running in a 1,180-pod namespace. ## Axis 1: aggressiveness | Lever | Cluster Autoscaler | Karpenter | |---|---|---| | What counts as underused | `--scale-down-utilization-threshold` (0.5) | `consolidationPolicy` | | How long before acting | `--scale-down-unneeded-time` (10m) | `consolidateAfter` (0s) | | Pause after growth | `--scale-down-delay-after-add` (10m) | no separate setting; `consolidateAfter` covers it | | Concurrency and timing | `--max-scale-down-parallelism` (10), `--max-drain-parallelism` (1), pod grace capped by `--max-graceful-termination-sec` (600) | `budgets` with `reasons` and `schedule` | Aggressive settings cut idle node-hours but increase **churn**: more evictions, more cold caches, more pods waiting for a node that was removed ten minutes earlier. A bursty tenant like the dispatcher argues for longer waits on its pool. A steady batch pool can be tightened. The usual answer is **different settings per pool**, not one global number. ## Axis 2: who may veto Four things can stop a node from going away. Each one is legitimate for some workload and harmful when used by default: - `cluster-autoscaler.kubernetes.io/safe-to-evict: "false"` or `karpenter.sh/do-not-disrupt`; - a **PodDisruptionBudget** that allows zero disruptions (`maxUnavailable: 0`, or `minAvailable` equal to the replica count); - **local storage** and pods with no controller; - **required placement rules** that only one node satisfies. A workable policy: 1. **Budgets must allow at least one disruption.** A policy engine rejects budgets that can never be satisfied, and the platform team reviews exceptions. 2. **Pins go on a dedicated pool.** Pods that must not be interrupted tolerate a taint on a separate node group or NodePool, so a pin strands one small pool instead of general capacity. 3. **Pins expire.** On Karpenter, `terminationGracePeriod` bounds how long a node can be held. On Cluster Autoscaler, the equivalent is an agreed maximum job length plus review. 4. **System add-ons get budgets**, so unbudgeted `kube-system` pods do not block general nodes. ## Axis 3: making the cost visible A veto is a spending decision that is usually made without the spender seeing the bill. Measure: - **stranded node-hours**: nodes the autoscaler reports as unremovable, multiplied by the hours they stay, attributed to the team whose pod pinned them. Eleven nodes pinned for a month is 11 × 24 × 30 = 7,920 node-hours; - **scale-down attempts that fail**, and their reasons; - **eviction rate per workload**, so aggressive settings show up as the churn they cause. Report these per team. The policy then holds up because teams can see what their exceptions cost. ## Axis 4: the scale-up side Scale-down policy only works if scale-up is predictable. When a HorizontalPodAutoscaler adds dispatcher replicas, they wait as Pending pods until a node exists, which takes minutes. Options: - keep **headroom** so new replicas land immediately (a pattern with its own design and cost); - use **per-zone node groups** with similar-group balancing, so zone-spread workloads get nodes in every zone; - set **node-count and core limits**, so a runaway HPA cannot grow the cluster without bound. ## A recommended starting point 1. General pools: default threshold and unneeded time, and on Karpenter a `consolidateAfter` of 10-15 minutes with a 10% budget. 2. Latency-critical pools: longer waits, and business-hours budgets of zero for `Underutilized` and `Drifted`. 3. A tainted "no-interrupt" pool for pinned work, with a maximum hold time. 4. Admission rules on budgets and pin annotations. 5. A monthly stranded-cost report per team, with the exception list attached. ## What you give up Longer waits cost money during quiet hours. Zero-disruption windows delay security image rollouts. Dedicated pools add fragmentation. A strict budget rule forces teams to run more replicas. State these costs openly: the aim is a policy whose price everyone can see, not zero churn or zero waste. Revisit the numbers each quarter. Teams add replicas, pools change shape, and a threshold that fitted last spring's traffic may now strand a dozen nodes every night or evict the dispatcher twice an hour.

  • A team insists their pods need `maxUnavailable: 0` in their PodDisruptionBudget. How do you respond?
    I ask what failure they are guarding against. Usually it is too few replicas or slow startup, and the fix is more replicas and a budget that allows one disruption. If the workload truly cannot lose a pod, it runs on the dedicated no-interrupt pool with a scheduled maintenance window, and the node-hours it strands are charged to that team.
  • How should an HPA-driven burst interact with your Cluster Autoscaler scale-down settings?
    The burst creates Pending replicas, and the autoscaler adds nodes for them. Each scale-up pauses scale-down for the after-add delay. When the burst ends, the HPA's own scale-down stabilisation and the node unneeded time run in sequence, so nodes linger for a while. That is acceptable when it is deliberate. Set the unneeded time from the burst pattern, and cap total nodes and cores so a misbehaving HPA cannot grow the cluster without bound.

saying these in an interview costs you the question

  • One global scale-down setting is right for every tenant
  • Pin annotations are free, because the pods are small
  • The goal is zero idle nodes at all times
  • PodDisruptionBudgets with zero allowed disruptions are the safe default
  • Scale-down policy can be designed without looking at scale-up latency