skip to content

A Deployment scales from 3 to 6 replicas and the new pods stay Pending with the scheduler message `node(s) didn't match pod topology spread constraints`. Explain what that constraint does and how you would decide between relaxing it and adding capacity.

level: seniorimportance: should knowfreq 45%

answer

  1. skew = this domain count − least-populated eligible domain
  2. DoNotSchedule = filter → Pending; ScheduleAnyway = scoring only
  3. topologyKey: zone (hard) + hostname (soft) is the common pairing
  4. matchLabelKeys: pod-template-hash stops rollout deadlock
  5. podAntiAffinity hostname = one per node, blunt version

basics

~20 s

A topologySpreadConstraint limits how uneven pod placement may be across a topology key (zone, node): the count difference between domains must stay within maxSkew. With whenUnsatisfiable: DoNotSchedule, a pod that would break the skew stays Pending. Relax to ScheduleAnyway, raise maxSkew, or add capacity in the short domains.

solid answer

~60 s

`topologySpreadConstraints` express "spread these pods evenly across zones/nodes". Key fields: - **`topologyKey`** — the node label defining a domain (`topology.kubernetes.io/zone`, `kubernetes.io/hostname`). - **`maxSkew`** — the maximum allowed difference between the most- and least-populated matching domains. - **`labelSelector`** — which existing pods are counted. - **`whenUnsatisfiable`** — `DoNotSchedule` (hard filter → Pending) or `ScheduleAnyway` (scoring preference only). - **`minDomains`**, `nodeAffinityPolicy`/`nodeTaintsPolicy`, `matchLabelKeys` refine which domains and nodes count. The message means every candidate node sits in a domain that is already at the skew ceiling — typically three zones with two replicas each, `maxSkew: 1`, and one zone with no room, so placing a fourth anywhere would create a skew of 2. Deciding: if the constraint encodes a real availability requirement (survive a zone loss), add capacity in the short domain — that is exactly what the constraint is telling you. If it is a nice-to-have, switch to `ScheduleAnyway` so degraded capacity yields imbalance rather than an outage. `podAntiAffinity` with `kubernetes.io/hostname` is the older, stricter one-per-node variant and fails the same way.

code

yaml · 13 lines
yaml
spec:
  topologySpreadConstraints:
    - maxSkew: 1
      topologyKey: topology.kubernetes.io/zone
      whenUnsatisfiable: DoNotSchedule
      matchLabelKeys: ["pod-template-hash"]
      labelSelector:
        matchLabels: { app: api }
    - maxSkew: 1
      topologyKey: kubernetes.io/hostname
      whenUnsatisfiable: ScheduleAnyway
      labelSelector:
        matchLabels: { app: api }

go deeper

for a junior

Know that these constraints spread replicas across zones or nodes and that a hard version can leave pods Pending when the spread cannot be met.

for a middle

Define skew correctly, name maxSkew/topologyKey/whenUnsatisfiable/labelSelector, and know the two ways out: relax the constraint or add capacity where it is short.

for a senior

Diagnose from the actual distribution and the eligible domain set, know matchLabelKeys and the rollout deadlock, and argue the relax-versus-scale decision from the availability requirement.

for a principal

Own the policy: which tiers get hard zone spread, how the cluster autoscaler is configured to scale the short domain, cost of standing zone headroom, and how spread interacts with PDBs and disruption budgets during zone failure.

## The goal: not all eggs in one basket Running six replicas means nothing for availability if all six land in one availability zone or on one node. Kubernetes offers two mechanisms to prevent that. ### podAntiAffinity — the older, blunt tool ```yaml affinity: podAntiAffinity: requiredDuringSchedulingIgnoredDuringExecution: - labelSelector: matchLabels: { app: api } topologyKey: kubernetes.io/hostname ``` This says: no two `app=api` pods may share a node. It is binary — one per domain, no more. On a five-node cluster, replica six is permanently `Pending` with `node(s) didn't satisfy existing pods anti-affinity rules`. It is also expensive to evaluate at scale, because the scheduler must inspect the pods on every candidate node. ### topologySpreadConstraints — the modern, tunable tool ```yaml topologySpreadConstraints: - maxSkew: 1 topologyKey: topology.kubernetes.io/zone whenUnsatisfiable: DoNotSchedule labelSelector: matchLabels: { app: api } ``` Instead of "at most one", it bounds the **skew**: for the set of pods matching the selector, `skew = count(this domain) − count(least-populated eligible domain)`. Placing a pod is allowed only if the resulting skew stays within `maxSkew`. Worked example: three zones, `maxSkew: 1`, and existing counts `a=2, b=2, c=2`. Adding one pod makes that domain 3 against a minimum of 2 — skew 1, allowed. But if zone `c` has no schedulable capacity, placing pods only in `a` and `b` quickly gives `a=3, b=3, c=2`… then the next pod would make `a=4` against `c=2`, skew 2, forbidden. Every node fails the predicate and the pod stays `Pending` with `node(s) didn't match pod topology spread constraints`. The crucial insight: **the constraint is working as designed**. It is refusing to concentrate risk. `Pending` is the constraint reporting that you do not have the capacity to remain as available as you asked to be. ## The knobs that change the outcome - **`whenUnsatisfiable: ScheduleAnyway`** — demotes the rule from a filter to a scoring preference. Pods always schedule; balance is best-effort. This is the right choice when imbalance is worse-but-acceptable and downtime is not. - **`maxSkew`** — raising it to 2 or 3 tolerates more imbalance while still preventing total concentration. - **`minDomains`** — requires at least N eligible domains to exist; with fewer, the constraint fails. Useful to force multi-zone presence, but it can also be the hidden reason for `Pending` when a zone's nodes are all cordoned. - **`nodeAffinityPolicy` / `nodeTaintsPolicy`** (`Honor` / `Ignore`) — whether nodes excluded by the pod's own affinity or by taints still count as domains. `Honor` (the default for nodeAffinityPolicy) prevents a zone your pod could never use from being counted as "empty" and distorting the skew. - **`matchLabelKeys`** — commonly `[pod-template-hash]`, so a rolling update counts only pods from the same revision. Without it, old and new ReplicaSet pods are counted together and a rollout can deadlock: the old pods occupy the domains, so the new ones cannot satisfy the skew, so the old ones are never removed. ## Diagnosis 1. Read the full event — it names the constraint that failed and often the skew. 2. Count reality: `kubectl get pods -l app=api -o wide` grouped by node/zone tells you the current distribution instantly. 3. Check which domains are actually usable: cordoned nodes, taints, and zones with no free allocatable capacity all shrink the eligible set. 4. Look for a rollout in progress — a spread deadlock during an update is usually the `matchLabelKeys` problem or a `maxUnavailable: 0` strategy combined with a hard spread. ## The decision Ask what the constraint exists for. - If it encodes an availability requirement — "we must survive losing a zone" — then `Pending` is the honest signal that the cluster cannot host six replicas at that guarantee. **Add capacity in the under-populated domain** (or let the cluster autoscaler do it; modern autoscalers understand spread constraints and will scale the right zone's node group). Relaxing the constraint here silently converts an availability guarantee into an outage waiting for a zone failure. - If it is a best-effort balance for a stateless service where partial concentration is tolerable, use `ScheduleAnyway`. Running six imbalanced replicas beats running three. A good production pattern is layered: a hard constraint at `maxSkew: 1` across **zones**, and a soft (`ScheduleAnyway`) constraint across **hostnames** — guaranteed zone diversity, best-effort node diversity — and `matchLabelKeys: [pod-template-hash]` so rollouts do not deadlock. ## Interview framing Define skew precisely, show that `DoNotSchedule` makes it a filter, name the deciding question (is this an availability guarantee or a preference?), and mention the rollout deadlock and `matchLabelKeys`. That last detail is what separates someone who has operated this from someone who has read about it.

  • How can a topology spread constraint deadlock a rolling update?
    By default the constraint counts every pod matching its labelSelector, including the old ReplicaSet's pods. During a rollout the old pods already fill the domains, so a new pod cannot be placed without breaking maxSkew, and the old pods are not removed until the new ones are Ready. Setting `matchLabelKeys: [pod-template-hash]` scopes counting to the pod's own revision and breaks the cycle; allowing some `maxUnavailable` also helps.
  • When is `whenUnsatisfiable: ScheduleAnyway` the wrong choice?
    When the spread encodes a real availability contract — for instance a quorum-based system that must keep replicas in separate zones or failure domains to survive losing one. With ScheduleAnyway the scheduler will happily concentrate replicas when capacity is tight, and you only discover the concentration during the failure the constraint was meant to survive.

saying these in an interview costs you the question

  • Describing maxSkew as a maximum pods-per-domain rather than a maximum difference between domains
  • Assuming ScheduleAnyway still guarantees some spread
  • Not knowing that cordoned/tainted nodes and node affinity change which domains count
  • Blaming the scheduler rather than reading the constraint as a capacity signal
  • Treating podAntiAffinity per hostname and topology spread as interchangeable — anti-affinity is one-per-domain and cannot be tuned

context