skip to content

You are setting a default multi-zone spreading policy for every workload in a shared Kubernetes cluster. What would you make the default, what would you require teams to opt into, and what does the policy cost in capacity and incident behaviour?

level: principalimportance: nice to knowfreq 26%

answer

  1. Three levers: scheduler defaultConstraints, platform templates, admission policy
  2. Soft default; hard is an opt-in with obligations
  3. Hard + zone outage + HPA scale-up = Pending during the incident
  4. Spreading fights bin-packing → per-zone headroom, weaker scale-in
  5. Declaring any constraint replaces defaults; unlabelled nodes vanish from domains

basics

~20 s

Default to soft zone and hostname spreading for everything (ScheduleAnyway, small maxSkew) via the scheduler's default constraints, and require an explicit opt-in for hard DoNotSchedule constraints, reserved for quorum and compliance workloads. Hard defaults strand pods during zone outages and scale-ups; soft defaults cost some imbalance but never block capacity.

solid answer

~60 s

**Default: soft.** I set the scheduler's `PodTopologySpread` default constraints to `ScheduleAnyway` on both `topology.kubernetes.io/zone` and `kubernetes.io/hostname` with a small `maxSkew` (1–2). Every workload gets a meaningful nudge with no risk of Pending pods, and teams that declare nothing still land reasonably spread. **Opt-in: hard.** `DoNotSchedule` is reserved for workloads where imbalance genuinely defeats the availability model — quorum stores, compliance-driven multi-zone requirements — and comes with obligations: node groups in every zone that the autoscaler can grow, a PodDisruptionBudget, and an owner who accepts Pending pods as the failure mode. **Costs to state plainly.** Balanced placement fights bin-packing, so you carry more headroom and cannot consolidate onto fewer nodes as tightly; hard constraints turn a zone outage into reduced capacity precisely when an HPA wants more. Balance also decays — scale-down, drains and preemption ignore spread, and nothing rebalances — so the policy needs a descheduler and an alert on observed per-zone counts, not just a manifest standard. And note that any pod declaring its own constraints replaces the defaults entirely — so the platform chart should emit both constraints, not one.

code

yaml · 15 lines
yaml
apiVersion: kubescheduler.config.k8s.io/v1
kind: KubeSchedulerConfiguration
profiles:
  - schedulerName: default-scheduler
    pluginConfig:
      - name: PodTopologySpread
        args:
          defaultingType: List
          defaultConstraints:
            - maxSkew: 1
              topologyKey: topology.kubernetes.io/zone
              whenUnsatisfiable: ScheduleAnyway
            - maxSkew: 2
              topologyKey: kubernetes.io/hostname
              whenUnsatisfiable: ScheduleAnyway

go deeper

for a junior

Recognise that a cluster can set default spreading behaviour and that soft defaults are safer than hard ones.

for a middle

Name the levers (scheduler defaultConstraints, templates), pick soft as the default with a reason, and note that declaring constraints replaces the defaults.

for a senior

Argue the incident behaviour concretely — zone outage plus HPA scale-up under a hard constraint — and add descheduler plus skew alerting because balance decays.

for a principal

Own the whole policy: default posture, an exception process with obligations attached, quantified capacity cost against zone-loss risk, node-labelling guarantees, and game days that verify the policy does what the document says.

## What a cluster-wide policy actually is Three levers exist and a real policy uses all three: 1. **Scheduler default constraints.** The `PodTopologySpread` plugin in `KubeSchedulerConfiguration` carries `defaultConstraints`, applied to any pod that declares none of its own. Kubernetes ships soft constraints with `maxSkew: 3` on zone and hostname. This is the floor for workloads nobody has thought about. 2. **Platform templates.** The shared Helm library or generator that most teams use to emit Deployments; this is where an explicit, well-formed pair of constraints belongs. 3. **Admission policy.** Kyverno/Gatekeeper/`ValidatingAdmissionPolicy` to mutate in defaults for workloads above a replica threshold, or to reject hard constraints from namespaces that have not opted in. ## Why the default must be soft A hard cluster-wide default is a foot-gun with a delay fuse. It behaves perfectly in steady state and fails on the days that matter: - **Zone outage.** Eligible domains shrink from three to two. With `maxSkew: 1` and `DoNotSchedule`, additional replicas can only be added in near-lockstep across the survivors, so an HPA scaling up during the incident produces Pending pods. You have converted a partial-capacity event into a smaller-capacity event. - **Uneven zone capacity.** Cloud instance availability is not uniform; a shortage in one zone with a hard constraint stalls the whole workload rather than degrading its distribution. - **Hostname spreading.** Hard `maxSkew: 1` on hostname caps a workload near the node count and couples Deployment scaling to cluster scaling. Soft defaults have the opposite profile: they never block, and they degrade to "less balanced than ideal" — a condition you can observe and repair, rather than an outage. ## What earns the hard opt-in Make hardness a request with a justification, not a default: - **Quorum systems.** Three etcd/ZooKeeper/Kafka-controller members must be in three zones; two in one zone means a zone loss costs quorum. Here Pending is genuinely preferable to a bad layout. - **Contractual or regulatory multi-zone presence.** - **Workloads whose failover is manual** and whose distribution therefore cannot be repaired quickly. Attach obligations to the grant: node groups present and scalable in every relevant zone (otherwise the constraint is unsatisfiable by construction), `minDomains` set so the autoscaler is prompted to restore a missing domain rather than the scheduler concluding two domains are fine, a PodDisruptionBudget so voluntary disruption cannot break the layout, and an on-call owner who understands that the failure mode is a Pending pod. ## The capacity conversation Spreading is in direct tension with bin-packing. A perfectly balanced workload occupies more distinct nodes, which: - Raises the floor of nodes you must keep running, reducing the autoscaler's ability to consolidate and scale in. - Requires per-zone headroom, since a zone must be able to absorb its share of a spike locally — you cannot borrow another zone's slack for a hard-constrained workload. - Interacts with reserved-capacity purchases, which are usually zonal; a policy that forces even spread must be matched by even reservations or you pay on-demand rates in the short zone. Quantify it for the cluster rather than asserting it: what fraction of capacity is idle purely to satisfy balance, and what would a zone loss actually cost if the policy were softer. ## Balance is not self-maintaining The most under-appreciated point. The scheduler enforces spread at placement only. Scale-down chooses victims without consulting topology; node drains remove a node's whole cohort; preemption evicts by priority. So a cluster that was balanced on Monday is measurably skewed by Friday. A credible policy therefore includes: - **A descheduler** running `RemovePodsViolatingTopologySpreadConstraint`, rate-limited, always with PodDisruptionBudgets in force so the corrective eviction cannot itself cause an outage. - **Observability**: export per-workload, per-zone pod counts and alert on sustained skew beyond policy, because no Kubernetes event tells you a soft constraint lost. - **Game days**: cordon a zone and confirm what the policy actually does to Pending counts and capacity. ## Two implementation traps to name - **Declaring one constraint drops the defaults.** A pod that specifies only a hostname constraint loses the default zone constraint entirely. Platform templates must always emit the full set. - **Unlabelled nodes silently vanish** from the topology calculation, so node bootstrap must guarantee `topology.kubernetes.io/zone` and `kubernetes.io/hostname` on every node, in every node group, including any on-prem or specialty additions. ## The recommendation Soft-by-default at the scheduler, well-formed constraint pairs in the platform template, hard constraints by exception with obligations attached, a descheduler plus skew alerting to keep the promise true over time, and a periodic drill to confirm the policy behaves the way the document claims.

  • How do you stop the declared policy from decaying over the following months?
    Because scale-down, node drains and preemption all ignore topology spread and nothing rebalances running pods, balance decays even in a well-configured cluster. I run a descheduler with the RemovePodsViolatingTopologySpreadConstraint strategy, rate-limited and with PodDisruptionBudgets enforced so the corrective evictions are safe, and I export per-workload per-zone pod counts with an alert on sustained skew. A soft constraint that loses produces no event, so measurement is the only way to know.
  • A team asks for DoNotSchedule with maxSkew 1 on hostname for a 40-replica service on a 30-node cluster. What do you say?
    I decline that specific shape. Hard hostname spreading at maxSkew 1 caps the workload near the node count, so 10 replicas would be permanently Pending and Deployment scaling would be coupled to cluster scaling. If the real intent is 'do not let many replicas pile onto one node', a soft hostname constraint or a larger maxSkew delivers it without the ceiling. If the intent is genuinely one per node, a DaemonSet is the correct object.

A soft default is a recommended speed limit that keeps traffic sensible; a hard default is a barrier across the road — perfect until a lane closes, when it stops everyone rather than slowing them.

saying these in an interview costs you the question

  • Making DoNotSchedule the cluster-wide default and treating Pending pods during a zone outage as acceptable
  • Assuming the declared policy stays true over time when scale-down, drain and preemption all ignore spread
  • Emitting a single constraint in a template and silently dropping the scheduler's default zone constraint
  • Ignoring that spreading fights bin-packing, so per-zone headroom and weaker scale-in are part of the bill
  • Setting hard zone constraints without node groups able to scale in every zone, making the constraint unsatisfiable by construction

context