Why does kube-scheduler never move running pods when Kubernetes nodes become unbalanced, and what does the descheduler do about it?
answer
- binding is final for the scheduler
- evict, don't place
- controller recreates, scheduler re-places
- thresholds vs targetThresholds
- local storage skipped by default
basics
~20 skube-scheduler decides only for unbound pods; once bound, a pod stays put. The descheduler is a separate add-on that evicts selected pods through the Eviction API so their controllers recreate them and the scheduler places them again.
solid answer
~40 skube-scheduler only handles pods with no `nodeName`. Once it binds a pod, it never revisits that decision, so node additions, deletions and churn leave the cluster unbalanced. The **descheduler** (a kubernetes-sigs project run as a Job, CronJob or Deployment) fills that gap. It evaluates a `DeschedulerPolicy` and **evicts** pods that its plugins pick. Their controllers recreate them, and the normal scheduler places the replacements. It never chooses a destination itself. `LowNodeUtilization` moves pods off nodes above `targetThresholds` when there are nodes below `thresholds`. `HighNodeUtilization` empties under-used nodes and is documented to need `MostAllocated` scoring. Both measure node use as requests divided by Allocatable by default; metrics-based use is optional. Evictions go through the eviction subresource, so PodDisruptionBudgets are respected. By default it skips DaemonSet pods, pods with local storage, system-critical pods and bare pods.
code
bash · 2 lineskubectl -n kube-system get cronjob descheduler
kubectl -n kube-system logs job/descheduler-29018640 | grep -i evictgo deeper
Remember that the scheduler places a pod once and never moves it, and that an add-on called the descheduler evicts pods so they get placed again.
Explain the evict-then-reschedule loop, the thresholds and targetThresholds model on requests over Allocatable, and which pods are protected by default.
Run it safely: pair compaction with MostAllocated, rate-limit evictions, require PDBs, and keep single-replica stateful pods out of scope.
Decide whether ongoing eviction churn is worth the packing gain for your fleet, compared with fixing request shapes or node sizes at the source.
## Scheduling is a one-time decision **kube-scheduler** watches for pods whose `spec.nodeName` is empty, runs its filter and score phases, and binds each pod to a node. After binding, the pod belongs to that node's kubelet, and the scheduler does not look at it again. That keeps scheduling fast and predictable, but placements go stale: - a node added after a busy period stays nearly empty while old nodes stay full; - pods that moved during a node outage stay piled on the survivors when the node returns; - labels, taints or affinity rules change, and existing pods no longer satisfy them; - deletions leave holes that nothing fills (fragmentation). ## The descheduler's model The **descheduler** is a separate kubernetes-sigs component, not part of the control plane. It can run as a `Job`, a `CronJob` or a `Deployment` (with leader election when it has more than one replica). Each cycle it: 1. reads a `DeschedulerPolicy` (apiVersion `descheduler/v1alpha2`) made of profiles and plugins; 2. lets **balance** and **deschedule** plugins pick candidate pods; 3. filters those candidates through its default evictor, which applies protections; 4. **evicts** the survivors through the pods' eviction subresource. It never decides where a pod goes. The owning controller (ReplicaSet, StatefulSet, Job) creates a replacement, and kube-scheduler places it with the current scoring. If the scheduler would put it right back, the eviction achieved nothing. That is why the compaction plugin needs matching scheduler scoring. ## The utilisation plugins | Plugin | Goal | How it classifies nodes | |---|---|---| | `LowNodeUtilization` | spread load onto under-used nodes | below `thresholds` on every listed resource = under-used; above `targetThresholds` on any = over-used; evicts from over-used nodes, and does nothing if either group is empty | | `HighNodeUtilization` | compact pods onto fewer nodes | below `thresholds` = under-used; evicts from those nodes; documented to need `MostAllocated` scoring | - Percentages are **requested resources divided by Allocatable** by default, for `cpu`, `memory`, `pods` and optionally extended resources. That matches what the scheduler sees. - `LowNodeUtilization` can use real usage through `metricsUtilization` (a metrics source or a Prometheus query). Then its view differs from the scheduler's. - `thresholds` and `targetThresholds` must list the same resources, and a threshold cannot exceed its target. Other plugins handle placement drift, not capacity: `RemoveDuplicates`, `RemovePodsViolatingNodeAffinity`, `RemovePodsViolatingNodeTaints`, `RemovePodsViolatingInterPodAntiAffinity`, `RemovePodsViolatingTopologySpreadConstraint`, `PodLifeTime` and others. ```yaml apiVersion: "descheduler/v1alpha2" kind: "DeschedulerPolicy" profiles: - name: wiki-cluster-balance pluginConfig: - name: "LowNodeUtilization" args: thresholds: "cpu": 25 "memory": 30 "pods": 25 targetThresholds: "cpu": 70 "memory": 75 "pods": 70 plugins: balance: enabled: - "LowNodeUtilization" ``` ## What it will not evict by default - **DaemonSet pods**: they would come straight back on the same node. - **Pods with local storage** (`emptyDir` or `hostPath` volumes): moving them loses data. - **System-critical pods** (`system-cluster-critical`, `system-node-critical`). - **Bare pods** with no controller, including static and mirror pods, because nothing would recreate them. - Pods whose **PodDisruptionBudget** would be violated: the eviction subresource refuses them. Pods backed by a **PersistentVolumeClaim** *are* evictable by default unless the `PodsWithPVC` protection is turned on. An internal wiki that keeps its embedded database in an `emptyDir` is skipped. The same wiki on a PVC is a candidate, and a single-replica database pod being evicted means downtime. Protect it with a PDB or an explicit protection. ## When to use it The descheduler suits clusters where churn or node changes leave lasting imbalance or fragmentation, especially where nodes are fixed (bare metal) and consolidation is the only way to free a large block. It is best effort: each cycle is a set of evictions, so rate-limit it (for example `evictionLimits`), keep workloads PDB-protected, and make sure its policy agrees with the scheduler's scoring strategy.
- Why is HighNodeUtilization documented to require MostAllocated scoring in kube-scheduler?The descheduler only evicts. The scheduler decides where the replacements go. With the default LeastAllocated scoring, pods evicted from under-used nodes would be placed on the emptiest nodes again, often the ones just drained, so nothing compacts. MostAllocated sends them to fuller nodes, so the emptied nodes stay empty.
- What stops the descheduler from taking down every replica of a service at once?It evicts through the pods' eviction subresource, so a PodDisruptionBudget covering the service turns an eviction that would breach the budget into a refusal. Per-cycle limits such as `evictionLimits` cap how many pods it removes. Without a PDB, only those limits and the plugin's own logic stand in the way.
saying these in an interview costs you the question
- kube-scheduler periodically rebalances running pods across nodes
- The descheduler picks a new node and moves the pod there
- The descheduler always measures node usage from live metrics
- The descheduler deletes pods directly, ignoring PodDisruptionBudgets
- HighNodeUtilization works with any scheduler scoring strategy