In a Kubernetes topology spread constraint, what is the difference between `whenUnsatisfiable: DoNotSchedule` and `whenUnsatisfiable: ScheduleAnyway`, and how do you decide which to use?
answer
- DoNotSchedule = filter phase; ScheduleAnyway = score phase
- Hard trades pod availability for distribution; soft trades distribution for capacity
- Quorum systems hard; stateless services soft
- Hard hostname spread caps replicas at node count
- Default constraints are ScheduleAnyway, maxSkew 3
basics
~20 sDoNotSchedule is a hard filter: if no domain can take the pod without exceeding maxSkew, the pod stays Pending. ScheduleAnyway is a scoring preference: the scheduler favours the emptiest domain but always places the pod. Choose hard when imbalance is worse than unavailability, soft otherwise.
solid answer
~60 sThe field decides whether the constraint is a **filter** or a **score**. - **`DoNotSchedule`** runs in the scheduler's filter phase. Nodes whose domains would break `maxSkew` are removed from the feasible set; if none survive, the pod stays `Pending` with an event naming the spread constraint. You get the distribution you asked for, or no pod at all. - **`ScheduleAnyway`** runs in the score phase. Balanced domains score higher, so the scheduler prefers them, but a pod is always placed somewhere. How I choose: use `DoNotSchedule` when running the replicas unevenly is genuinely worse than running fewer — a quorum-based system where two of three members in one zone means a zone outage loses quorum. Use `ScheduleAnyway` for ordinary stateless services where capacity beats perfect balance, and for the `kubernetes.io/hostname` constraint where hard spreading breaks scale-up on small clusters. Common pattern: hard across zones, soft across hostnames. And remember the risk — a hard constraint plus a lost zone plus an autoscaler that cannot add capacity in the remaining zones means replicas sit Pending during exactly the incident you were protecting against.
code
yaml · 13 linestopologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: payments
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: paymentsgo deeper
State plainly that DoNotSchedule can leave a pod Pending while ScheduleAnyway always places it, just preferring balance.
Connect them to the scheduler's filter and score phases and give one workload example for each choice.
Frame it as pod availability versus service distribution, name the zone-outage plus scale-up failure mode, and give mitigations such as looser maxSkew or hard-zone/soft-hostname.
Set it as policy: which workload classes get hard constraints, how it interacts with capacity planning and autoscaling per zone, and the cost of Pending replicas during an incident versus running unbalanced.
## Two phases, two behaviours The kube-scheduler processes a pending pod in two phases: **filter**, which eliminates nodes that cannot host the pod, and **score**, which ranks the survivors. `whenUnsatisfiable` chooses which phase the `PodTopologySpread` plugin uses for a given constraint. `DoNotSchedule` participates in filtering. For each candidate node the plugin simulates placing the pod in that node's domain and computes the resulting skew against the least-populated eligible domain; if it exceeds `maxSkew`, the node is filtered out. When every node is filtered out the pod is `Pending`, and `kubectl describe pod` reports something like `3 node(s) didn't match pod topology spread constraints`. `ScheduleAnyway` participates only in scoring. Every node stays feasible; nodes in emptier domains simply receive a higher score. The score is combined with all the other scoring plugins — resource balance, image locality, affinity preferences, taint tolerance — and can be outweighed by them. So a soft constraint gives you a tendency, not a guarantee, and on a lopsided cluster the tendency can lose. ## The real tradeoff: availability of the pod versus availability of the service This is the framing that gets you full marks. A hard constraint protects the *distribution* of the service by sacrificing individual pods: rather than run a badly-distributed replica, it runs nothing. A soft constraint protects *capacity*: it always runs the replica, accepting that the distribution may be poor. Which is right depends on whether imbalance actually hurts: - **Quorum systems** (etcd, ZooKeeper, Kafka controllers, anything with a raft group) — imbalance is genuinely dangerous. Two of three members in one zone means that zone's loss costs you quorum, which is worse than running with two members deliberately. `DoNotSchedule`. - **Compliance or contractual multi-zone requirements** — the distribution is the requirement. `DoNotSchedule`. - **Ordinary stateless HTTP services** — an extra replica in the wrong zone still serves traffic. Losing a zone costs you a share of capacity either way; having *more* total capacity is usually better. `ScheduleAnyway`. - **Hostname-level spreading** — hard hostname spreading caps your replica count at roughly the node count, so a scale-up beyond that leaves pods Pending forever. Almost always `ScheduleAnyway`, unless you truly require one replica per node (in which case a DaemonSet may be the better tool). ## The failure mode of a hard constraint The scenario to name explicitly: three zones, `maxSkew: 1`, `DoNotSchedule`, 9 replicas at 3/3/3. Zone C is lost. The 3 pods there are rescheduled — but placing any of them in A or B would push those to 4 against a minimum of 3 in the now-empty-but-still-eligible... actually the more common bite is scale-out: an HPA doubles the Deployment during the incident, and with only two eligible zones the constraint permits at most an even split, so pods that would exceed it sit `Pending`. You have optimised for balance during exactly the moment you needed capacity. Mitigations worth mentioning: - Loosen `maxSkew` (2 or 3) instead of going fully soft — you still bound imbalance but leave slack. - Use `minDomains` deliberately so the scheduler keeps expecting the missing domain and the cluster autoscaler is prompted to provision there. - Pair a hard zone constraint with a soft hostname constraint, so the strictness is only where it buys real failure isolation. - Ensure node groups exist in every zone so the autoscaler can satisfy the constraint rather than being stuck. ## Interaction with defaults If a pod declares no constraints at all, the scheduler's default constraints apply. The shipped defaults are `ScheduleAnyway` with `maxSkew: 3` on both zone and hostname — deliberately soft, because a hard cluster-wide default would strand workloads. Declaring your own constraints replaces the defaults for that pod entirely, so if you write a single hostname constraint you also lose the default zone nudge; write both if you want both. ## Neither option rebalances Worth stating for either value: the scheduler evaluates spread only when placing a pod. Nothing moves existing pods when a domain becomes available again. Restoring balance after an incident means recreating pods — a `kubectl rollout restart`, or a descheduler policy that evicts pods violating topology spread. ## Checking which one bit you If pods are Pending, the describe output names the constraint — that is `DoNotSchedule` at work. If pods are all running but clustered, you have a soft constraint being outvoted by other scoring plugins, or a selector/label problem, and no event will tell you: you have to inspect the actual distribution.
- You use DoNotSchedule with maxSkew 1 across three zones and one zone goes down. What happens when the HPA scales the Deployment up during the incident?Only two zones remain eligible, so the constraint permits at most an even split between them; any replica that would push one zone more than one pod ahead of the other stays Pending. You get less capacity precisely when you need more. Mitigations are a looser maxSkew, making the hostname constraint soft while keeping zone hard, or ensuring the autoscaler has node groups able to grow in the surviving zones.
- Your pods declare only a hostname spread constraint and you notice zone balance got worse. Why?Declaring any topologySpreadConstraints replaces the scheduler's default constraints for that pod, and the shipped defaults include a soft zone constraint with maxSkew 3. By writing only a hostname constraint you dropped the zone nudge entirely. If you want both behaviours you must declare both constraints explicitly in the pod spec.
A hard constraint is a fire-code occupancy limit — the door simply does not open. A soft constraint is a host suggesting the emptier table; if you insist, you still get seated.
saying these in an interview costs you the question
- Describing ScheduleAnyway as a weaker filter rather than a scoring preference that never blocks placement
- Using DoNotSchedule with maxSkew 1 on kubernetes.io/hostname for a Deployment that can exceed the node count
- Assuming a soft constraint guarantees balance — it can be outweighed by other scoring plugins
- Not realising that declaring any constraint replaces the cluster's default constraints for that pod
- Believing either setting rebalances already-running pods when a zone recovers