What do Kubernetes `topologySpreadConstraints` in a pod spec do, and what do the `maxSkew`, `topologyKey`, `labelSelector` and `whenUnsatisfiable` fields each mean?
answer
- topologyKey = node label defining the domain
- skew = count in this domain − global minimum
- maxSkew ≥ 1; 1 = as even as arithmetic allows
- labelSelector counts all matching pods in the namespace
- Nodes lacking the topology label are excluded entirely
basics
~20 sThey tell the scheduler to spread matching pods evenly across failure domains. topologyKey is the node label defining a domain (zone, hostname); labelSelector picks which pods are counted; maxSkew is the largest allowed difference in count between domains; whenUnsatisfiable says whether to block scheduling or place anyway.
solid answer
~50 s`topologySpreadConstraints` is a list on `pod.spec` that keeps a group of pods evenly distributed across failure domains, so losing one zone or node does not take out most of the replicas. The fields: - **`topologyKey`** — the node label that defines a domain. Common values are `topology.kubernetes.io/zone` and `kubernetes.io/hostname`. Every node sharing a value is one domain. - **`labelSelector`** — which existing pods are counted in the domain totals. Normally the workload's own labels; it counts pods cluster-wide (within the namespace), not just this ReplicaSet. - **`maxSkew`** — the maximum allowed difference between the number of matching pods in the fullest and the emptiest eligible domain. `maxSkew: 1` means as even as arithmetic allows. - **`whenUnsatisfiable`** — `DoNotSchedule` (hard: pod stays Pending if placement would exceed the skew) or `ScheduleAnyway` (soft: the scheduler scores in favour of balance but still places the pod). You can list several constraints; all of them must be satisfied simultaneously, e.g. spread across zones *and* across nodes.
code
yaml · 19 linesspec:
template:
metadata:
labels:
app: web
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: web
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: webgo deeper
Name the four fields and say the feature spreads replicas across zones or nodes so one failure does not take everything down.
Compute skew correctly on an example, know that unlabelled nodes are excluded, and that multiple constraints are ANDed.
Add the operational nuances: selector scope across ReplicaSets during rollouts, matchLabelKeys, minDomains, and the fact that spreading is decided only at placement time.
Discuss it as an availability policy — what skew the SLO actually needs, hard versus soft cluster-wide defaults, interaction with capacity and autoscaling, and the cost of Pending pods versus imbalance.
## The problem A Deployment with six replicas is not highly available if the scheduler bin-packs all six onto two nodes in one availability zone. The default scheduler already spreads a little — its `PodTopologySpread` plugin contributes score, and by default the cluster ships soft constraints on zone and hostname — but nothing guarantees a distribution. `topologySpreadConstraints` is the explicit, declarative way to state "these pods must be spread across these domains, this evenly". ## Domains A *topology domain* is the set of nodes that share the same value for a given node label. Choose the label with `topologyKey`: - `topology.kubernetes.io/zone` — one domain per availability zone; protects against a zone outage. - `kubernetes.io/hostname` — one domain per node; protects against a single-node failure and stops bin-packing. - Any custom label — `rack`, `topology.kubernetes.io/region`, `nodepool`. Nodes that do not carry the label at all are **not** a domain; they are excluded from the calculation entirely, and pods on them are invisible to it. That is a common source of surprise on clusters with mixed or unlabelled nodes. ## The skew calculation For a given constraint the scheduler counts, in each eligible domain, the pods matching `labelSelector`. Then: `skew = pods in the domain being considered − minimum pod count across eligible domains` A candidate placement is feasible only if the resulting skew stays ≤ `maxSkew`. In effect the scheduler must place the next pod in one of the currently-emptiest domains once the spread reaches the limit. Worked example: 3 zones, `maxSkew: 1`, and current counts A=2, B=1, C=1. Minimum is 1. Placing in A gives 3 − 1 = 2 > 1, infeasible. Placing in B gives 2 − 1 = 1, fine. So the 5th replica goes to B or C. Six replicas across three zones land 2/2/2; seven land 3/2/2, which is the best arithmetic allows. `maxSkew` must be at least 1. Larger values loosen the requirement: `maxSkew: 2` over three zones permits 4/2/2, which is often the pragmatic choice when zones have unequal capacity. ## labelSelector counts more than you think The selector is evaluated against **all pods in the namespace**, not just the ones created by this controller. If two Deployments share the label `app: web`, a constraint selecting `app: web` balances their combined population. That is sometimes exactly what you want (all frontends spread together) and sometimes a bug (a canary Deployment perturbing the stable one's placement). Select precisely; where you want per-revision balance, `matchLabelKeys: [pod-template-hash]` makes the scheduler add the current revision's hash to the selector automatically so old and new ReplicaSets are counted separately during a rollout. ## whenUnsatisfiable `DoNotSchedule` makes the constraint a hard filter: if no domain can take the pod without breaking `maxSkew`, the pod stays `Pending` with an event naming the constraint. `ScheduleAnyway` makes it a scoring preference: the scheduler prefers balanced placements but never refuses to place the pod. See the dedicated comparison for how to choose. ## Related knobs - **`minDomains`** (with `DoNotSchedule`) — asserts a minimum number of eligible domains; if fewer exist, the global minimum is treated as zero, which forces the scheduler to keep domains free rather than piling into the few that exist. Useful with cluster autoscaling so new zones get provisioned. - **`nodeAffinityPolicy` / `nodeTaintsPolicy`** — whether nodes excluded by the pod's own nodeSelector/affinity or by untolerated taints are counted when computing domains. Defaults are `Honor` for affinity and `Ignore` for taints. ## Cluster-wide defaults If a pod declares no constraints, the scheduler applies default constraints from its configuration. The shipped defaults are soft (`ScheduleAnyway`) with `maxSkew: 3` on both zone and hostname — enough to nudge spreading, nowhere near a guarantee. Administrators can replace them in the `KubeSchedulerConfiguration` for the `PodTopologySpread` plugin. ## Verifying `kubectl get pods -o wide` shows nodes; joining pods to node zone labels shows the actual zone distribution. When a pod is Pending, `kubectl describe pod` prints something like `didn't match pod topology spread constraints`, which tells you the constraint, not the resource, is the blocker.
- With 3 zones, maxSkew 1 and 7 replicas, what distribution do you expect, and does the scheduler rebalance if a zone later comes back?You get 3/2/2, since the skew between fullest and emptiest is 1, which is the best arithmetic allows for 7 across 3. The scheduler makes decisions only at placement time, so if a zone was unavailable when the pods were created they will have landed 4/3/0 or similar and will stay that way after the zone returns. Rebalancing requires deleting pods — a rollout restart, or a descheduler that evicts pods violating the constraint.
- Your constraint uses topologyKey `topology.kubernetes.io/zone` but a third of the nodes were added without zone labels. What happens?Those nodes form no domain and are excluded from the skew computation entirely. With `DoNotSchedule` this can make placements fail on a cluster that appears to have plenty of capacity, because the eligible domains are full while the unlabelled nodes are idle. With `ScheduleAnyway` pods can still land on them, but the resulting distribution is not what the constraint describes. The fix is to ensure the cloud provider or node bootstrap applies the well-known topology labels consistently.
Seating guests across tables: topologyKey is what counts as a table, labelSelector is which guests you are balancing, and maxSkew is how many more guests the fullest table may have than the emptiest.
saying these in an interview costs you the question
- Thinking topologyKey names a pod label rather than a node label
- Believing maxSkew is a per-domain maximum pod count instead of a difference between domains
- Forgetting labelSelector counts every matching pod in the namespace, including other Deployments and old ReplicaSets
- Assuming the scheduler rebalances existing pods after a domain returns — placement decisions are one-shot
- Setting maxSkew: 0, which is invalid; the minimum is 1