skip to content

Topology Spread Constraints

topologySpreadConstraints keep replicas evenly distributed across zones or nodes within a tunable maxSkew, and decide whether to hard-fail or place anyway when the skew cannot be met. Asked in HA rounds as the cheaper replacement for pod anti-affinity.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

5

What do Kubernetes `topologySpreadConstraints` in a pod spec do, and what do the `maxSkew`, `topologyKey`, `labelSelector` and `whenUnsatisfiable` fields each mean?

level: middleimportance: must knowfreq 52%

answer

  1. topologyKey = node label defining the domain
  2. skew = count in this domain − global minimum
  3. maxSkew ≥ 1; 1 = as even as arithmetic allows
  4. labelSelector counts all matching pods in the namespace
  5. Nodes lacking the topology label are excluded entirely

basics

~20 s

They tell the scheduler to spread matching pods evenly across failure domains. topologyKey is the node label defining a domain (zone, hostname); labelSelector picks which pods are counted; maxSkew is the largest allowed difference in count between domains; whenUnsatisfiable says whether to block scheduling or place anyway.

solid answer

~50 s

`topologySpreadConstraints` is a list on `pod.spec` that keeps a group of pods evenly distributed across failure domains, so losing one zone or node does not take out most of the replicas. The fields: - **`topologyKey`** — the node label that defines a domain. Common values are `topology.kubernetes.io/zone` and `kubernetes.io/hostname`. Every node sharing a value is one domain. - **`labelSelector`** — which existing pods are counted in the domain totals. Normally the workload's own labels; it counts pods cluster-wide (within the namespace), not just this ReplicaSet. - **`maxSkew`** — the maximum allowed difference between the number of matching pods in the fullest and the emptiest eligible domain. `maxSkew: 1` means as even as arithmetic allows. - **`whenUnsatisfiable`** — `DoNotSchedule` (hard: pod stays Pending if placement would exceed the skew) or `ScheduleAnyway` (soft: the scheduler scores in favour of balance but still places the pod). You can list several constraints; all of them must be satisfied simultaneously, e.g. spread across zones *and* across nodes.

code

yaml · 19 lines
yaml
spec:
  template:
    metadata:
      labels:
        app: web
    spec:
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: topology.kubernetes.io/zone
          whenUnsatisfiable: DoNotSchedule
          labelSelector:
            matchLabels:
              app: web
        - maxSkew: 1
          topologyKey: kubernetes.io/hostname
          whenUnsatisfiable: ScheduleAnyway
          labelSelector:
            matchLabels:
              app: web

go deeper

for a junior

Name the four fields and say the feature spreads replicas across zones or nodes so one failure does not take everything down.

for a middle

Compute skew correctly on an example, know that unlabelled nodes are excluded, and that multiple constraints are ANDed.

for a senior

Add the operational nuances: selector scope across ReplicaSets during rollouts, matchLabelKeys, minDomains, and the fact that spreading is decided only at placement time.

for a principal

Discuss it as an availability policy — what skew the SLO actually needs, hard versus soft cluster-wide defaults, interaction with capacity and autoscaling, and the cost of Pending pods versus imbalance.

## The problem A Deployment with six replicas is not highly available if the scheduler bin-packs all six onto two nodes in one availability zone. The default scheduler already spreads a little — its `PodTopologySpread` plugin contributes score, and by default the cluster ships soft constraints on zone and hostname — but nothing guarantees a distribution. `topologySpreadConstraints` is the explicit, declarative way to state "these pods must be spread across these domains, this evenly". ## Domains A *topology domain* is the set of nodes that share the same value for a given node label. Choose the label with `topologyKey`: - `topology.kubernetes.io/zone` — one domain per availability zone; protects against a zone outage. - `kubernetes.io/hostname` — one domain per node; protects against a single-node failure and stops bin-packing. - Any custom label — `rack`, `topology.kubernetes.io/region`, `nodepool`. Nodes that do not carry the label at all are **not** a domain; they are excluded from the calculation entirely, and pods on them are invisible to it. That is a common source of surprise on clusters with mixed or unlabelled nodes. ## The skew calculation For a given constraint the scheduler counts, in each eligible domain, the pods matching `labelSelector`. Then: `skew = pods in the domain being considered − minimum pod count across eligible domains` A candidate placement is feasible only if the resulting skew stays ≤ `maxSkew`. In effect the scheduler must place the next pod in one of the currently-emptiest domains once the spread reaches the limit. Worked example: 3 zones, `maxSkew: 1`, and current counts A=2, B=1, C=1. Minimum is 1. Placing in A gives 3 − 1 = 2 > 1, infeasible. Placing in B gives 2 − 1 = 1, fine. So the 5th replica goes to B or C. Six replicas across three zones land 2/2/2; seven land 3/2/2, which is the best arithmetic allows. `maxSkew` must be at least 1. Larger values loosen the requirement: `maxSkew: 2` over three zones permits 4/2/2, which is often the pragmatic choice when zones have unequal capacity. ## labelSelector counts more than you think The selector is evaluated against **all pods in the namespace**, not just the ones created by this controller. If two Deployments share the label `app: web`, a constraint selecting `app: web` balances their combined population. That is sometimes exactly what you want (all frontends spread together) and sometimes a bug (a canary Deployment perturbing the stable one's placement). Select precisely; where you want per-revision balance, `matchLabelKeys: [pod-template-hash]` makes the scheduler add the current revision's hash to the selector automatically so old and new ReplicaSets are counted separately during a rollout. ## whenUnsatisfiable `DoNotSchedule` makes the constraint a hard filter: if no domain can take the pod without breaking `maxSkew`, the pod stays `Pending` with an event naming the constraint. `ScheduleAnyway` makes it a scoring preference: the scheduler prefers balanced placements but never refuses to place the pod. See the dedicated comparison for how to choose. ## Related knobs - **`minDomains`** (with `DoNotSchedule`) — asserts a minimum number of eligible domains; if fewer exist, the global minimum is treated as zero, which forces the scheduler to keep domains free rather than piling into the few that exist. Useful with cluster autoscaling so new zones get provisioned. - **`nodeAffinityPolicy` / `nodeTaintsPolicy`** — whether nodes excluded by the pod's own nodeSelector/affinity or by untolerated taints are counted when computing domains. Defaults are `Honor` for affinity and `Ignore` for taints. ## Cluster-wide defaults If a pod declares no constraints, the scheduler applies default constraints from its configuration. The shipped defaults are soft (`ScheduleAnyway`) with `maxSkew: 3` on both zone and hostname — enough to nudge spreading, nowhere near a guarantee. Administrators can replace them in the `KubeSchedulerConfiguration` for the `PodTopologySpread` plugin. ## Verifying `kubectl get pods -o wide` shows nodes; joining pods to node zone labels shows the actual zone distribution. When a pod is Pending, `kubectl describe pod` prints something like `didn't match pod topology spread constraints`, which tells you the constraint, not the resource, is the blocker.

  • With 3 zones, maxSkew 1 and 7 replicas, what distribution do you expect, and does the scheduler rebalance if a zone later comes back?
    You get 3/2/2, since the skew between fullest and emptiest is 1, which is the best arithmetic allows for 7 across 3. The scheduler makes decisions only at placement time, so if a zone was unavailable when the pods were created they will have landed 4/3/0 or similar and will stay that way after the zone returns. Rebalancing requires deleting pods — a rollout restart, or a descheduler that evicts pods violating the constraint.
  • Your constraint uses topologyKey `topology.kubernetes.io/zone` but a third of the nodes were added without zone labels. What happens?
    Those nodes form no domain and are excluded from the skew computation entirely. With `DoNotSchedule` this can make placements fail on a cluster that appears to have plenty of capacity, because the eligible domains are full while the unlabelled nodes are idle. With `ScheduleAnyway` pods can still land on them, but the resulting distribution is not what the constraint describes. The fix is to ensure the cloud provider or node bootstrap applies the well-known topology labels consistently.

Seating guests across tables: topologyKey is what counts as a table, labelSelector is which guests you are balancing, and maxSkew is how many more guests the fullest table may have than the emptiest.

saying these in an interview costs you the question

  • Thinking topologyKey names a pod label rather than a node label
  • Believing maxSkew is a per-domain maximum pod count instead of a difference between domains
  • Forgetting labelSelector counts every matching pod in the namespace, including other Deployments and old ReplicaSets
  • Assuming the scheduler rebalances existing pods after a domain returns — placement decisions are one-shot
  • Setting maxSkew: 0, which is invalid; the minimum is 1

context

open as a page

In a Kubernetes topology spread constraint, what is the difference between `whenUnsatisfiable: DoNotSchedule` and `whenUnsatisfiable: ScheduleAnyway`, and how do you decide which to use?

level: middleimportance: must knowfreq 46%

basics

~20 s

DoNotSchedule is a hard filter: if no domain can take the pod without exceeding maxSkew, the pod stays Pending. ScheduleAnyway is a scoring preference: the scheduler favours the emptiest domain but always places the pod. Choose hard when imbalance is worse than unavailability, soft otherwise.

open as a page

A Kubernetes Deployment declares a topology spread constraint with `maxSkew: 1` across availability zones, yet after a rollout its pods sit 5/2/1 across three zones. Give the likely causes and how you would investigate.

level: seniorimportance: should knowfreq 36%

basics

~20 s

Most likely: the constraint is ScheduleAnyway so it only scores; or the domains were unequal at placement time (a zone was full or unavailable) and nothing rebalances afterwards; or nodes lack the zone label and are excluded; or the labelSelector is counting the wrong pod set. Check the constraint, the node labels, and the actual per-zone counts.

open as a page

Kubernetes offers both `topologySpreadConstraints` and pod anti-affinity for keeping replicas apart. What are the practical differences, and when would you still reach for anti-affinity?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Anti-affinity is binary per domain — required anti-affinity allows at most one matching pod per domain, which caps replicas at the number of domains. Spread constraints express a degree of evenness via maxSkew, so many pods per domain are fine. Use anti-affinity when you truly need at most one, or to express repulsion from different pods.

open as a page

You are setting a default multi-zone spreading policy for every workload in a shared Kubernetes cluster. What would you make the default, what would you require teams to opt into, and what does the policy cost in capacity and incident behaviour?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Default to soft zone and hostname spreading for everything (ScheduleAnyway, small maxSkew) via the scheduler's default constraints, and require an explicit opt-in for hard DoNotSchedule constraints, reserved for quorum and compliance workloads. Hard defaults strand pods during zone outages and scale-ups; soft defaults cost some imbalance but never block capacity.

open as a page