A Kubernetes Deployment declares a topology spread constraint with `maxSkew: 1` across availability zones, yet after a rollout its pods sit 5/2/1 across three zones. Give the likely causes and how you would investigate.
answer
- Spread is placement-time only — nothing rebalances
- ScheduleAnyway can be outvoted by other scoring plugins
- Unlabelled nodes = no domain = excluded from counts
- Rollouts: old ReplicaSet pods counted unless matchLabelKeys pod-template-hash
- Scale-down and drain ignore topology spread entirely
basics
~20 sMost likely: the constraint is ScheduleAnyway so it only scores; or the domains were unequal at placement time (a zone was full or unavailable) and nothing rebalances afterwards; or nodes lack the zone label and are excluded; or the labelSelector is counting the wrong pod set. Check the constraint, the node labels, and the actual per-zone counts.
solid answer
~60 sSpread is enforced **only at placement time**, so a 5/2/1 distribution usually means it was legal — or unenforced — when those pods were bound. The candidates, in the order I check them: 1. **`whenUnsatisfiable: ScheduleAnyway`** — a preference other scoring plugins can outvote. Look at the actual constraint, not the intent. 2. **Unequal capacity at placement time.** A zone with no free nodes, or a taint the pod does not tolerate, means the scheduler places elsewhere; when that zone recovers nothing moves back. Spread never rebalances. 3. **Nodes missing `topology.kubernetes.io/zone`** — unlabelled nodes are not a domain and are excluded from the calculation entirely. 4. **`labelSelector` scope.** It counts every matching pod in the namespace. During a rollout the old ReplicaSet's pods count too, unless `matchLabelKeys: [pod-template-hash]` scopes the count per revision — this is the classic cause of skew appearing mid-rollout. 5. **`nodeAffinityPolicy` / `nodeTaintsPolicy`** changing which nodes count as eligible. Investigation: dump pod→node→zone counts, `kubectl describe node` the zone labels, and check Pending pods' events. Fix forward with `kubectl rollout restart` or a descheduler policy once the constraint itself is right.
code
bash · 7 lineskubectl get nodes -L topology.kubernetes.io/zone
for p in $(kubectl get pods -l app=web -o jsonpath='{.items[*].spec.nodeName}'); do
kubectl get node "$p" -o jsonpath='{.metadata.labels.topology\.kubernetes\.io/zone}{"\n"}'
done | sort | uniq -c
kubectl get deploy web -o jsonpath='{.spec.template.spec.topologySpreadConstraints}' | jq .go deeper
Recall that spread is applied when a pod is scheduled and never afterwards, and that a soft constraint may simply not be enforced.
List the concrete causes — soft constraint, missing node labels, unequal capacity, selector scope — and know the commands that check each.
Run the full investigation, explain matchLabelKeys and the node policies, and know that scale-down and drain create skew the scheduler will not repair; fix forward with rollout restart or the descheduler.
Treat balance as something that must be observed and maintained: alerting on per-zone counts, a descheduler with PDBs, standard constraint templates in the platform chart, and capacity per zone so the constraint is actually satisfiable.
## The governing fact: spreading is a placement-time decision `topologySpreadConstraints` is evaluated by the scheduler when it binds a pod to a node. It is never re-evaluated. There is no controller reconciling the distribution back towards balance. Every drift explanation follows from that: whatever the cluster looked like at bind time is what the constraint was judged against, and any later change — a zone recovering, nodes being added, other pods disappearing — leaves the existing placement untouched. ## Cause 1: the constraint is soft `whenUnsatisfiable: ScheduleAnyway` only contributes score. Resource-fit, image locality, affinity preferences and taint scoring are all summed alongside it, and on a cluster where one zone has much larger or emptier nodes the spread score can lose repeatedly. The result is exactly the kind of lopsided distribution described. Read the manifest as deployed (`kubectl get deploy -o yaml`) rather than trusting the chart's intent, and check whether a Helm value or an admission mutator changed it. ## Cause 2: domains were not equally able to accept pods Spread constraints do not create capacity. If zone C had no allocatable room, or its nodes carried a taint the pod does not tolerate, or its nodes failed the pod's own nodeSelector, the scheduler either skipped it (soft) or left pods Pending (hard). With `nodeTaintsPolicy: Ignore` — the default — nodes the pod cannot tolerate are still counted as part of the domain, which makes a domain look emptier than it usably is and can distort decisions. Setting `nodeTaintsPolicy: Honor` and `nodeAffinityPolicy: Honor` makes the computation reflect nodes the pod could actually use. ## Cause 3: nodes missing the topology label A node without `topology.kubernetes.io/zone` belongs to no domain. It is excluded from the skew computation, and pods that land on it are invisible to the counts. Mixed clusters — on-prem nodes added to a cloud cluster, or a node group whose bootstrap script omits the labels — produce distributions that look impossible until you check. `kubectl get nodes -L topology.kubernetes.io/zone` is a one-line diagnosis. ## Cause 4: the labelSelector counts the wrong population The selector is evaluated against all pods in the namespace, not just this ReplicaSet's. Two consequences: - **Other workloads sharing a label** are counted, so your balance is computed over a bigger, differently-distributed population. - **During a rolling update**, old and new ReplicaSet pods both match `app: web`. The scheduler counts them together, so the new pods are placed to balance the *combined* population; as the old pods are then deleted, the new set alone can be badly skewed. `matchLabelKeys: [pod-template-hash]` fixes this by having the scheduler add the incoming pod's `pod-template-hash` value to the selector, scoping the count to the current revision. ## Cause 5: scale-down and eviction chose badly The scheduler places pods; it does not choose which pods die. Scale-down picks victims by the ReplicaSet controller's own ordering, node drains remove whole nodes' worth of pods, and preemption evicts by priority. None of them consult topology spread. So a Deployment scaled 9→3 can end with all three survivors in one zone even though the original nine were perfectly balanced. ## Investigation procedure 1. **Measure.** Join pods to their nodes' zone labels and count — do not eyeball `kubectl get pods -o wide`. 2. **Read the effective constraint.** `kubectl get deploy web -o jsonpath='{.spec.template.spec.topologySpreadConstraints}'`. Confirm `whenUnsatisfiable`, `maxSkew`, `topologyKey`, and the selector. 3. **Check node labels.** `kubectl get nodes -L topology.kubernetes.io/zone` — any blanks are excluded domains. 4. **Check the selector's real population.** `kubectl get pods -l app=web -o wide | wc -l` versus the ReplicaSet's replica count; a mismatch means you are balancing more than you thought. 5. **Check Pending pods' events.** `didn't match pod topology spread constraints` proves a hard constraint is biting; the absence of Pending pods with a lopsided result points at a soft constraint or historic placement. 6. **Check for taints/capacity in the light zone.** `kubectl describe node` in the underweight zone. ## Fixing Correct the constraint first — tighten to `DoNotSchedule` if the workload genuinely needs it, add `matchLabelKeys`, fix node labels, set `nodeTaintsPolicy: Honor`. Then force re-placement: `kubectl rollout restart deployment/web` recreates pods so the corrected constraint is applied. For ongoing drift, run a descheduler with the `RemovePodsViolatingTopologySpreadConstraint` strategy so pods violating the constraint are evicted periodically and rescheduled — with a PodDisruptionBudget in place so the eviction is safe. Add an alert on observed per-zone counts; balance is not something you can assume once and forget, because nothing in the core control plane maintains it.
- You fix the constraint. How do you get the already-running pods rebalanced?Nothing in the core control plane moves running pods, so you must recreate them. `kubectl rollout restart deployment/web` replaces pods gradually under the corrected constraint and is usually enough. For continuous correction, run the descheduler with the RemovePodsViolatingTopologySpreadConstraint strategy, which evicts pods that violate the constraint so the scheduler places them again — pair it with a PodDisruptionBudget so those evictions cannot take the service below its availability floor.
- Why does `matchLabelKeys: [pod-template-hash]` matter during a rolling update?Without it, the labelSelector matches pods from both the old and the new ReplicaSet, so the scheduler balances the combined population. As the old pods are then deleted, the new pods alone can be badly skewed even though every individual decision was legal. `matchLabelKeys` tells the scheduler to add the incoming pod's own pod-template-hash value to the selector, so counts are scoped to the revision being rolled out and the new set is balanced on its own terms.
saying these in an interview costs you the question
- Assuming Kubernetes continuously reconciles pod distribution back to the declared skew
- Reading the intended manifest instead of the effective one and missing that the constraint is ScheduleAnyway
- Overlooking nodes without the topology label, which silently drop out of the calculation
- Forgetting the selector counts other Deployments and the old ReplicaSet during a rollout
- Blaming the scheduler for imbalance actually created by scale-down or node drain, neither of which consults spread constraints