skip to content

How would you use Kubernetes podAntiAffinity to keep the replicas of a Deployment off the same machine, and what exactly does the topologyKey field control?

level: middleimportance: must knowfreq 58%

answer

  1. labelSelector + namespaces + topologyKey
  2. topologyKey names a NODE label; equal values = one domain
  3. hostname = one per node; zone = one per zone
  4. required caps replicas at domain count -> Pending
  5. unlabelled node = invisible domain

basics

~20 s

Add podAntiAffinity whose labelSelector matches the Deployment's own Pod labels, with topologyKey kubernetes.io/hostname. topologyKey names a node label; nodes sharing a value form one domain, and the scheduler refuses to place a second matching Pod into a domain that already has one.

solid answer

~50 s

Inter-Pod anti-affinity is a rule about *other Pods already placed*, not about node hardware. You write it under `spec.affinity.podAntiAffinity` with three parts: a `labelSelector` picking which Pods you are repelled from, an optional `namespaces`/`namespaceSelector` (default: the Pod's own namespace), and a mandatory `topologyKey`. `topologyKey` is the name of a **node label**. All nodes sharing the same value for that label form one topology domain. `kubernetes.io/hostname` makes each node its own domain — one replica per node. `topology.kubernetes.io/zone` makes each availability zone a domain — one replica per zone. For a Deployment you point the labelSelector at the Deployment's own Pod labels, e.g. `app=checkout`, so replicas repel each other. Hard (`requiredDuringSchedulingIgnoredDuringExecution`) means a matching domain is filtered out entirely; if every domain is taken, the extra replicas sit `Pending`. Soft (`preferredDuringSchedulingIgnoredDuringExecution`, with a weight) only penalises the score, so replicas spread when they can and double up rather than stay Pending.

code

yaml · 31 lines
yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: checkout
spec:
  replicas: 3
  selector:
    matchLabels:
      app: checkout
  template:
    metadata:
      labels:
        app: checkout
    spec:
      affinity:
        podAntiAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
          - labelSelector:
              matchLabels:
                app: checkout
            topologyKey: topology.kubernetes.io/zone
          preferredDuringSchedulingIgnoredDuringExecution:
          - weight: 100
            podAffinityTerm:
              labelSelector:
                matchLabels:
                  app: checkout
              topologyKey: kubernetes.io/hostname
      containers:
      - name: app
        image: registry.example.com/checkout:2.1.0

go deeper

for a junior

Be able to state that anti-affinity keeps matching Pods apart and that topologyKey chooses the granularity, with kubernetes.io/hostname meaning one per node.

for a middle

Explain the three fields of the term, the namespace default, hard versus soft, and the replicas-capped-by-domains failure mode.

for a senior

Discuss combining zone-level hard with hostname-level soft, rolling-update deadlock with maxSurge, unlabelled nodes as silent holes, and reading FailedScheduling counts.

for a principal

Argue about where the guarantee belongs — placement rules per Deployment versus platform-enforced policy — and when a hard availability guarantee is worth capacity that sits idle.

## What inter-Pod affinity is Node affinity answers "what kind of machine do I need?". Inter-Pod affinity and anti-affinity answer "who else is already there?". `podAffinity` attracts a Pod towards nodes where matching Pods run; `podAntiAffinity` repels it away. Both live under `spec.affinity` and both have the same required/preferred split as node affinity. ## The three fields of a term 1. **labelSelector** — which existing Pods count. Standard `matchLabels`/`matchExpressions`. For spreading a Deployment's replicas you select the Deployment's own Pod template labels. 2. **namespaces** / **namespaceSelector** — which namespaces those Pods are looked for in. Omitted, it means the scheduling Pod's own namespace. This trips people up: a rule intended to be cluster-wide silently applies to one namespace. 3. **topologyKey** — mandatory and non-empty. It names a node label, and nodes carrying the same value for that label form one *topology domain*. ## How topologyKey actually evaluates For a candidate node N, the scheduler reads N's value for the topologyKey label — say `topology.kubernetes.io/zone=eu-west-1a`. It then looks at every Pod matching the labelSelector across the cluster, finds their nodes, and reads those nodes' values for the same label. For hard anti-affinity, if any matching Pod sits on a node with value `eu-west-1a`, node N is filtered out. So the unit of exclusion is the *domain*, not the node. Common topologyKeys: - `kubernetes.io/hostname` — one Pod per node. - `topology.kubernetes.io/zone` — one Pod per availability zone. - `topology.kubernetes.io/region` — one Pod per region. - Any custom label you apply consistently, e.g. `rack`, `failure-domain`. A node missing the topologyKey label is not part of any domain and is skipped for anti-affinity purposes — an unlabelled node pool is a silent hole in your spreading guarantee. ## Hard versus soft `requiredDuringSchedulingIgnoredDuringExecution` is a filter: the domain is either allowed or not. Its value is that the guarantee is real. Its cost is that the number of schedulable replicas is capped by the number of domains. Three replicas with hard hostname anti-affinity on a two-node cluster means one Pod stays `Pending` indefinitely, and a rolling update can deadlock because the new Pod cannot be placed until the old one on that node goes away — with the default `maxSurge` strategy that is a real hang. `preferredDuringSchedulingIgnoredDuringExecution` wraps each term in `{weight: 1-100, podAffinityTerm: {...}}`. Matching domains lose score, so the scheduler spreads if it can and packs if it must. For most stateless services this is the right default; for quorum-based systems (etcd, ZooKeeper, a database with synchronous replicas) where two replicas on one machine defeats the point, hard is justified. A frequent production pattern is both at once: hard anti-affinity at zone level for correctness, soft at hostname level for extra resilience. ## podAffinity, the mirror image The same structure attracts instead of repels: put a cache Pod in the same zone as the service that reads it by selecting `app=api` with `topologyKey: topology.kubernetes.io/zone` to cut cross-zone traffic charges and latency. Hostname-level podAffinity is rarely wise — it packs Pods onto one machine and turns that machine into a correlated failure domain. ## Practical rules - Anti-affinity and affinity are evaluated only at scheduling. Adding nodes later does not rebalance existing Pods; you must delete Pods or run a descheduler. - The `LimitPodHardAntiAffinityTopology` admission plugin, where enabled, forces `topologyKey: kubernetes.io/hostname` for required anti-affinity, precisely because arbitrary topologyKeys on hard rules are expensive. - Kubernetes offers a purpose-built alternative for even spreading; anti-affinity is the on/off tool, not the balancing tool. ## Diagnosis A blocked Pod's `FailedScheduling` event reads like "0/6 nodes are available: 3 node(s) didn't match pod anti-affinity rules, 3 node(s) had untolerated taint". That message tells you how many nodes each predicate eliminated, which is usually enough to see whether the anti-affinity or something else is the binding constraint.

  • You set required podAntiAffinity with topologyKey kubernetes.io/hostname on a Deployment with 5 replicas, and the cluster has 3 nodes. What happens?
    Three replicas schedule, one per node, and the remaining two stay Pending with a FailedScheduling event saying no node matched the pod anti-affinity rules. The Deployment reports 3/5 ready indefinitely. Cluster Autoscaler may add nodes if it can simulate that a new node satisfies the rule, but if the node group is at its maximum the Pods stay stuck.
  • Which namespaces does a podAntiAffinity labelSelector search by default, and why does that matter?
    By default only the scheduling Pod's own namespace. If you intend to repel from Pods in other namespaces you must set the namespaces list or a namespaceSelector explicitly. Teams that copy a rule into a second namespace often believe they have a cluster-wide guarantee when each namespace is in fact spreading independently.

topologyKey chooses the granularity of "the same place". With hostname, two replicas in the same building but different rooms are fine. With zone, the whole building counts as one place, so the second replica must go to a different building.

saying these in an interview costs you the question

  • Saying topologyKey is a Pod label — it is a node label read from the candidate node.
  • Omitting or emptying topologyKey; it is mandatory and the API rejects an empty value for these terms.
  • Assuming anti-affinity rebalances running Pods when nodes are added — it is scheduling-time only.
  • Believing the labelSelector matches all namespaces by default; it defaults to the Pod's own namespace.
  • Using required hostname anti-affinity with more replicas than nodes and then blaming the scheduler for Pending Pods.

context