Karpenter keeps replacing nodes under a latency-sensitive Kubernetes workload. How do NodePool consolidation and drift work, and how do you limit the disruption they cause?
answer
- cost-driven versus spec-driven replacement
- template hash and the Drifted condition
- consolidateAfter defaults to zero
- ten percent unless budgets say otherwise
- a pod annotation plus a node deadline
basics
~20 sConsolidation deletes or replaces nodes whose pods fit elsewhere more cheaply, and drift replaces nodes that no longer match their NodePool or NodeClass. Limit both with consolidateAfter, disruption budgets scoped by reason and schedule, PDBs, and the do-not-disrupt annotation.
solid answer
~40 sKarpenter's NodePool has a `disruption` block. Consolidation, under the default `consolidationPolicy: WhenEmptyOrUnderutilized`, deletes a node whose pods fit on other nodes, or replaces nodes with a cheaper one; `WhenEmpty` restricts it to nodes with no workload pods. Drift marks a NodeClaim `Drifted` when the NodePool template, its requirements or the provider's NodeClass no longer match, and Karpenter replaces it, launching the new capacity before draining the old node. To tame it, I set `consolidateAfter` above the burst period, add `budgets` (10% by default) scoped by `reasons` and a cron `schedule` so business hours allow zero disruptions, and keep PDBs, which Karpenter honours. I use `karpenter.sh/do-not-disrupt` on truly uninterruptible pods, backed by `terminationGracePeriod` so nothing pins a node forever.
go deeper
Recall that Karpenter removes nodes for two reasons, cost (consolidation) and spec mismatch (drift), and that PDBs still apply.
Explain the two consolidation policies, what marks a NodeClaim Drifted, and why the replacement launches before the old node drains.
Show how you would read Karpenter's disruption events, then combine consolidateAfter, reason-scoped scheduled budgets, PDBs and do-not-disrupt without pinning nodes forever.
Weigh consolidation savings against tail-latency and rollout risk, and decide which pools may churn during business hours.
## How Karpenter meets Kubernetes **Karpenter** is a node provisioner that runs as a controller in the cluster. It watches Pending pods, batches them for a short window (one second idle, ten seconds at most, by default), and creates a **NodeClaim** for an instance shaped to those pods. The **NodePool** (`karpenter.sh/v1`) is the policy object: a `template` of what nodes may look like (requirements, taints, `nodeClassRef`, `expireAfter`) and a `disruption` block that says when Karpenter may take nodes away. Taking nodes away is **voluntary disruption**, and it comes from two mechanisms that interviewers ask about most: **consolidation** and **drift**. ## Consolidation Consolidation removes cost. For each candidate, Karpenter asks whether the node's pods would fit elsewhere, either on existing nodes (**delete** the node) or on a single cheaper node (**replace** it). It considers single nodes and groups of nodes. `spec.disruption.consolidationPolicy` controls which nodes qualify: | Value | Nodes Karpenter may consolidate | |---|---| | `WhenEmptyOrUnderutilized` (default) | Empty nodes, and nodes whose pods could be packed elsewhere or onto a cheaper node | | `WhenEmpty` | Only nodes Karpenter considers empty: no workload pods that would need rescheduling (DaemonSet pods do not count) | `spec.disruption.consolidateAfter` (default `0s`, and `Never` is allowed) is how long a node must stay a candidate before Karpenter acts on it. At `0s`, a service with bursty traffic can see nodes replaced minutes after a burst ends. ## Drift Drift keeps nodes matching their declared spec. A NodeClaim gets the **`Drifted`** status condition when: 1. the NodePool's template has changed (Karpenter compares a hash stored in the `karpenter.sh/nodepool-hash` annotation); 2. the node no longer satisfies the NodePool's `requirements`; 3. the provider's NodeClass or the provider itself reports a difference, such as a new machine image; 4. the node's instance type is no longer offered. Drifted nodes are replaced. Karpenter launches the replacement capacity and waits for it to be ready before it drains the old node, which it first taints with `karpenter.sh/disrupted:NoSchedule`. A routine image bump in the NodeClass therefore rolls every node in the pool. `template.spec.expireAfter` (default `720h`) replaces nodes on age alone, so a pool that seems stable still turns over roughly monthly. ## The brakes - **Disruption budgets.** `spec.disruption.budgets` caps how many of the pool's nodes may be disrupting at once. If the field is unset, a single budget of `nodes: "10%"` applies. A budget can be scoped by `reasons` (`Underutilized`, `Empty`, `Drifted`) and activated on a cron `schedule` with a `duration`. When several budgets are active, the most restrictive one wins, so `nodes: "0"` during business hours freezes those reasons. - **PodDisruptionBudgets.** Karpenter checks them before choosing a node and evicts through the Eviction API, so a PDB paces the drain. - **`karpenter.sh/do-not-disrupt`.** Set to `"true"` on a pod (or on the node), it blocks voluntary disruption of that node. Current Karpenter also accepts a duration value, which protects the pod only for that long after it starts. - **`template.spec.terminationGracePeriod`.** This is the escape hatch in the other direction. Once a node has been draining this long, Karpenter deletes the remaining pods, bypassing blocked PDBs and `do-not-disrupt`, so nodes cannot be pinned forever. ## Worked example: a webhook-delivery dispatcher The webhook-delivery dispatcher, in a 1,180-pod namespace, shows p99 delivery-latency spikes every weekday afternoon. Karpenter's events show nodes being disrupted with reason `Underutilized` right after the morning peak, and the PDB shows allowed disruptions going to zero during each wave. The team: 1. sets `consolidateAfter: 15m`, so post-burst dips do not trigger replacements; 2. adds a budget of `nodes: "0"` for `Underutilized` and `Drifted`, active 13:00-21:00 UTC on weekdays; 3. keeps a PDB of `maxUnavailable: 5%` on the dispatcher; 4. sets `terminationGracePeriod: 48h`, so the one pod with `do-not-disrupt` cannot hold a node indefinitely. ```yaml apiVersion: karpenter.sh/v1 kind: NodePool metadata: name: webhook-dispatch spec: template: spec: nodeClassRef: group: karpenter.kwok.sh kind: KWOKNodeClass name: default requirements: - key: karpenter.sh/capacity-type operator: In values: ["on-demand"] expireAfter: 720h terminationGracePeriod: 48h disruption: consolidationPolicy: WhenEmptyOrUnderutilized consolidateAfter: 15m budgets: - nodes: "10%" - nodes: "0" reasons: ["Underutilized", "Drifted"] schedule: "0 13 * * 1-5" duration: 8h ``` The `nodeClassRef` here points at Karpenter's in-repo test provider; a real cluster references its cloud provider's NodeClass kind. After the change, check that the afternoon windows show no `Underutilized` or `Drifted` disruptions, and that drift replacements, such as an image bump, now roll during the evening window instead. If drift must never wait that long for security fixes, give `Drifted` its own, smaller non-zero budget rather than zero.
- Is an empty Karpenter NodePool `budgets` list the same as having no limit on disruption?No. When the field is unset, the API defaults it to a single budget of `nodes: "10%"`, so at most a tenth of the pool's nodes may be disrupting at once. A budget of `nodes: "0"` blocks the reasons it names while it is active. When several budgets are active, the most restrictive one applies.
- How do you run Karpenter nodes for spot-tolerant and spot-intolerant workloads side by side?Use separate NodePools whose `requirements` constrain the well-known label `karpenter.sh/capacity-type` to `spot` or `on-demand`. Taint the spot pool so only tolerant workloads land there. Karpenter then provisions each instance shape per pool. How workloads handle interruption notices and reclaim is a separate concern from provisioning.
Consolidation is a warehouse manager merging half-empty shelves to save rent, and drift is the same manager swapping out shelves that no longer match the current blueprint. Budgets are the rule on how many aisles may be closed at once.
saying these in an interview costs you the question
- Karpenter ignores PodDisruptionBudgets because it deletes instances directly
- Drift only happens when someone edits the node by hand
- WhenEmpty stops drift replacements as well as consolidation
- Leaving budgets unset means Karpenter may disrupt every node at once
- do-not-disrupt is a Karpenter feature that also stops kubelet evictions