Kubernetes offers both `topologySpreadConstraints` and pod anti-affinity for keeping replicas apart. What are the practical differences, and when would you still reach for anti-affinity?
answer
- Anti-affinity = binary per domain; spread = quantitative maxSkew
- Required hostname anti-affinity caps replicas at node count
- Anti-affinity can repel *other* workloads; spread only balances one selected group
- No spread equivalent of pod affinity (co-location)
- Both are IgnoredDuringExecution — neither rebalances
basics
~20 sAnti-affinity is binary per domain — required anti-affinity allows at most one matching pod per domain, which caps replicas at the number of domains. Spread constraints express a degree of evenness via maxSkew, so many pods per domain are fine. Use anti-affinity when you truly need at most one, or to express repulsion from different pods.
solid answer
~60 s**Pod anti-affinity** answers "don't put me with pods like that". In its `required` form it is effectively a per-domain limit of one matching pod, so a Deployment with more replicas than domains leaves the surplus Pending. The `preferred` form only adds score, and gives no control over *how* uneven things may get. **Topology spread constraints** answer "distribute this group evenly". `maxSkew` expresses degree, so 30 replicas across 3 zones is perfectly expressible (10/10/10, or 11/10/9 with maxSkew 1). This is why spread constraints are the default choice for replica distribution today. I still reach for anti-affinity when: - The requirement really is **at most one per domain** — one control-plane component per node. - I need **repulsion between different workloads** — keep the cache away from the batch crunchers — which spread constraints, being about balancing one selected group, do not express. - I need **pod affinity** (co-location), which has no spread-constraint equivalent. Operationally, required anti-affinity across `kubernetes.io/hostname` is also the more expensive predicate on large clusters, and it is the classic cause of "pods Pending after scaling past node count".
code
yaml · 7 linesaffinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- topologyKey: kubernetes.io/hostname
labelSelector:
matchLabels:
app: etcdgo deeper
Know that both keep replicas apart, and that required anti-affinity means at most one per node while spread constraints allow several as long as counts stay close.
Explain the replica-count ceiling of required anti-affinity and that maxSkew expresses degree, and give the cross-workload case where anti-affinity is still needed.
Add scheduling cost, IgnoredDuringExecution semantics, namespace scope differences, and how the two compose for quorum workloads; diagnose from the Pending event text.
Set the guidance: spread constraints as the default for replica distribution, anti-affinity reserved for relational and cardinal requirements, and a limit on how many hard placement rules any workload may stack before availability suffers.
## Two different questions They look similar because both keep pods apart, but they encode different statements. **Pod anti-affinity** is a *relational* rule evaluated per candidate node: "do not schedule me into a topology domain that already contains a pod matching this selector". It is binary — the domain either contains such a pod or it does not. There is no notion of how many. **Topology spread** is a *distributional* rule about a population: "across these domains, the counts of pods matching this selector must not differ by more than maxSkew". It is quantitative. That difference drives almost everything else. ## The replica-count ceiling `requiredDuringSchedulingIgnoredDuringExecution` anti-affinity with `topologyKey: kubernetes.io/hostname`, selecting the pod's own labels, is the classic "one replica per node" recipe. It works beautifully until replicas exceed nodes: replica N+1 finds every node already occupied by a sibling and stays `Pending` forever. On an autoscaled cluster this is worse than it sounds, because the autoscaler adds nodes to satisfy it, so scaling the Deployment silently scales the cluster — an expensive coupling nobody intended. With spread constraints the same intent is `maxSkew: 1` on hostname, which allows 2, 3 or 10 per node as long as nodes stay within one of each other. The workload scales past node count naturally. ## The softness problem Anti-affinity's `preferred` form takes a `weight` and contributes score, but it cannot say "spread reasonably evenly" — it only says "nodes with siblings are less attractive". Real clusters routinely end up with a soft anti-affinity rule and a 6/1/1 zone distribution, which looks compliant and provides little protection. `ScheduleAnyway` spread constraints score proportionally to skew, which is a better-behaved preference, and `maxSkew` lets you state exactly how uneven is tolerable. ## What anti-affinity still does better 1. **Hard cardinality of one.** If the requirement literally is "never two of these on the same node" — because they contend for a device, a host port, a local disk, or a licence — required anti-affinity states it directly and enforces it. Expressing it as spread requires `maxSkew: 1` plus the knowledge that the total equals the node count, which is fragile. 2. **Cross-workload repulsion.** "Keep the latency-sensitive API away from the batch crunchers" selects *other* pods' labels. Spread constraints balance a selected group across domains; they do not express "stay away from that other group". Only anti-affinity does. 3. **Co-location.** Pod *affinity* — put the cache in the same zone as its consumer — has no spread-constraint counterpart at all. ## Cost and behaviour details - **Scheduling cost.** Inter-pod affinity/anti-affinity requires the scheduler to examine pods across the cluster per candidate node; the documentation itself warns against using it in clusters of several hundred nodes with non-hostname topology keys. `PodTopologySpread` is cheaper, working from precomputed domain counts. - **Both are `IgnoredDuringExecution`.** Neither moves a pod after placement. Rules are evaluated when the pod is scheduled and never re-checked, so a distribution that was legal at creation persists after it stops being ideal. - **Namespace scope.** Anti-affinity can select across namespaces via `namespaces` / `namespaceSelector`. Spread constraints count only pods in the pod's own namespace. - **They compose.** A common production combination is required anti-affinity on hostname for a small quorum set (never two members on a node) plus a spread constraint on zone with `maxSkew: 1` and `DoNotSchedule` (members balanced across zones). All rules must hold simultaneously, so the more you stack, the more likely something is Pending — which is a reason to keep the set small and deliberate. ## How to answer the "which should I use" question Default to topology spread constraints for distributing replicas of one workload — that is what they were built for and they scale with replica count. Reach for anti-affinity when the semantics are genuinely relational (repel *those* pods) or genuinely cardinal (at most one here). Treat required hostname anti-affinity on an autoscaled Deployment as a smell and check whether the author meant "spread out" rather than "exactly one per node". ## Debugging both A Pending pod's events distinguish them clearly: `node(s) didn't match pod anti-affinity rules` versus `node(s) didn't match pod topology spread constraints`. That single line usually resolves the argument about which rule is over-constraining the workload.
- A team reports that scaling their Deployment from 8 to 12 replicas leaves 4 pods Pending, and the cluster has 8 nodes with spare CPU. What do you suspect?Required pod anti-affinity on `kubernetes.io/hostname` selecting the workload's own labels — it permits at most one replica per node, so replicas beyond the node count can never be placed no matter how much CPU is free. The Pending pods' events will say the nodes did not match pod anti-affinity rules. If the intent was distribution rather than strict one-per-node, replacing it with a hostname spread constraint at maxSkew 1 fixes it.
- Can topology spread constraints keep one workload away from a different workload?No. A spread constraint balances the population of pods matched by its own labelSelector across domains; it has no way to say 'avoid domains containing those other pods'. If you point its selector at another workload's labels you are asking to balance that workload's counts, which is not repulsion. Cross-workload separation requires pod anti-affinity, which is explicitly relational.
Anti-affinity is a rule that no two people from the same team may sit at one table; spread constraints are a rule that no table may have more than one extra guest compared to the emptiest.
saying these in an interview costs you the question
- Claiming anti-affinity and spread constraints are interchangeable
- Using required hostname anti-affinity on an autoscaled Deployment and being surprised by Pending pods past node count
- Thinking preferred anti-affinity gives a bounded imbalance — it has no maxSkew equivalent
- Believing either mechanism rebalances running pods; both are IgnoredDuringExecution
- Ignoring the scheduling cost of inter-pod anti-affinity on large clusters with non-hostname topology keys