You own node autoscaling for a 140-node Kubernetes cluster shared by 22 product teams. How do you set scale-down policy: its aggressiveness, who may block it, and what that costs?
answer
- cost versus churn, per pool
- four ways to veto a node
- budgets must allow one
- pins go to a tainted pool
- stranded node-hours per team
basics
~20 sSet scale-down aggressiveness per pool, not globally. Allow vetoes only as reviewed exceptions: budgets that allow a disruption, and time-limited pins on a tainted pool. Report stranded node-hours per team so each veto's cost is visible.
solid answer
~40 sI treat scale-down as a policy with three axes. Aggressiveness: Cluster Autoscaler's threshold, unneeded time and after-add delay, or Karpenter's `consolidationPolicy`, `consolidateAfter` and `budgets`, set per pool, so bursty latency-critical pools wait longer while batch pools stay tight. Vetoes: `safe-to-evict: "false"`, `karpenter.sh/do-not-disrupt`, zero-disruption PDBs, local storage and unique placement rules all pin nodes. I require budgets to allow at least one disruption, isolate pinned work on a tainted pool, give system add-ons budgets and bound hold time. Cost: I publish stranded node-hours and failed scale-downs per team. The price is quiet-hour spend, slower image rollouts and some fragmentation, and I make those costs explicit rather than chasing zero churn.
go deeper
Recall that removing idle nodes saves money but evicts pods, and that some pods and budgets can stop a node from being removed.
Explain which Cluster Autoscaler and Karpenter settings control how quickly and how many nodes are removed, and what each blocker is.
Show you can find which team's pods pin nodes, and apply budgets, dedicated pools and bounded holds without breaking their workloads.
Own the cost-versus-churn tradeoff per pool, govern vetoes through admission policy and showback, and state openly what the policy costs.
## The decision being made In a shared cluster, **node scale-down** is where cost and stability meet. Every node the autoscaler cannot remove is paid for, and every node it does remove evicts someone's pods. The platform owner does not pick a correct answer here. They set a **policy**: how aggressive removal is, who is allowed to veto it, and how the cost of those vetoes becomes visible. Take a 140-node cluster shared by 22 product teams, whose largest tenant is a webhook-delivery dispatcher running in a 1,180-pod namespace. ## Axis 1: aggressiveness | Lever | Cluster Autoscaler | Karpenter | |---|---|---| | What counts as underused | `--scale-down-utilization-threshold` (0.5) | `consolidationPolicy` | | How long before acting | `--scale-down-unneeded-time` (10m) | `consolidateAfter` (0s) | | Pause after growth | `--scale-down-delay-after-add` (10m) | no separate setting; `consolidateAfter` covers it | | Concurrency and timing | `--max-scale-down-parallelism` (10), `--max-drain-parallelism` (1), pod grace capped by `--max-graceful-termination-sec` (600) | `budgets` with `reasons` and `schedule` | Aggressive settings cut idle node-hours but increase **churn**: more evictions, more cold caches, more pods waiting for a node that was removed ten minutes earlier. A bursty tenant like the dispatcher argues for longer waits on its pool. A steady batch pool can be tightened. The usual answer is **different settings per pool**, not one global number. ## Axis 2: who may veto Four things can stop a node from going away. Each one is legitimate for some workload and harmful when used by default: - `cluster-autoscaler.kubernetes.io/safe-to-evict: "false"` or `karpenter.sh/do-not-disrupt`; - a **PodDisruptionBudget** that allows zero disruptions (`maxUnavailable: 0`, or `minAvailable` equal to the replica count); - **local storage** and pods with no controller; - **required placement rules** that only one node satisfies. A workable policy: 1. **Budgets must allow at least one disruption.** A policy engine rejects budgets that can never be satisfied, and the platform team reviews exceptions. 2. **Pins go on a dedicated pool.** Pods that must not be interrupted tolerate a taint on a separate node group or NodePool, so a pin strands one small pool instead of general capacity. 3. **Pins expire.** On Karpenter, `terminationGracePeriod` bounds how long a node can be held. On Cluster Autoscaler, the equivalent is an agreed maximum job length plus review. 4. **System add-ons get budgets**, so unbudgeted `kube-system` pods do not block general nodes. ## Axis 3: making the cost visible A veto is a spending decision that is usually made without the spender seeing the bill. Measure: - **stranded node-hours**: nodes the autoscaler reports as unremovable, multiplied by the hours they stay, attributed to the team whose pod pinned them. Eleven nodes pinned for a month is 11 × 24 × 30 = 7,920 node-hours; - **scale-down attempts that fail**, and their reasons; - **eviction rate per workload**, so aggressive settings show up as the churn they cause. Report these per team. The policy then holds up because teams can see what their exceptions cost. ## Axis 4: the scale-up side Scale-down policy only works if scale-up is predictable. When a HorizontalPodAutoscaler adds dispatcher replicas, they wait as Pending pods until a node exists, which takes minutes. Options: - keep **headroom** so new replicas land immediately (a pattern with its own design and cost); - use **per-zone node groups** with similar-group balancing, so zone-spread workloads get nodes in every zone; - set **node-count and core limits**, so a runaway HPA cannot grow the cluster without bound. ## A recommended starting point 1. General pools: default threshold and unneeded time, and on Karpenter a `consolidateAfter` of 10-15 minutes with a 10% budget. 2. Latency-critical pools: longer waits, and business-hours budgets of zero for `Underutilized` and `Drifted`. 3. A tainted "no-interrupt" pool for pinned work, with a maximum hold time. 4. Admission rules on budgets and pin annotations. 5. A monthly stranded-cost report per team, with the exception list attached. ## What you give up Longer waits cost money during quiet hours. Zero-disruption windows delay security image rollouts. Dedicated pools add fragmentation. A strict budget rule forces teams to run more replicas. State these costs openly: the aim is a policy whose price everyone can see, not zero churn or zero waste. Revisit the numbers each quarter. Teams add replicas, pools change shape, and a threshold that fitted last spring's traffic may now strand a dozen nodes every night or evict the dispatcher twice an hour.
- A team insists their pods need `maxUnavailable: 0` in their PodDisruptionBudget. How do you respond?I ask what failure they are guarding against. Usually it is too few replicas or slow startup, and the fix is more replicas and a budget that allows one disruption. If the workload truly cannot lose a pod, it runs on the dedicated no-interrupt pool with a scheduled maintenance window, and the node-hours it strands are charged to that team.
- How should an HPA-driven burst interact with your Cluster Autoscaler scale-down settings?The burst creates Pending replicas, and the autoscaler adds nodes for them. Each scale-up pauses scale-down for the after-add delay. When the burst ends, the HPA's own scale-down stabilisation and the node unneeded time run in sequence, so nodes linger for a while. That is acceptable when it is deliberate. Set the unneeded time from the burst pattern, and cap total nodes and cores so a misbehaving HPA cannot grow the cluster without bound.
saying these in an interview costs you the question
- One global scale-down setting is right for every tenant
- Pin annotations are free, because the pods are small
- The goal is zero idle nodes at all times
- PodDisruptionBudgets with zero allowed disruptions are the safe default
- Scale-down policy can be designed without looking at scale-up latency