What does node affinity do for a Kubernetes Pod, and how does it differ from the simpler nodeSelector field in the Pod spec?
answer
- nodeSelector = exact match, AND-ed, mandatory
- affinity = matchExpressions: In / NotIn / Exists / Gt / Lt
- required = filter phase; preferred = score phase, weight 1-100
- nodeSelectorTerms OR-ed, matchExpressions AND-ed
- IgnoredDuringExecution = never evicts
basics
~20 sBoth steer a Pod onto nodes carrying particular labels. nodeSelector is exact key=value matching and every entry must match. Node affinity adds an expression language (In, NotIn, Exists, Gt, Lt), OR-ed alternatives, and a soft preferred form the scheduler may ignore.
solid answer
~50 sNode affinity is a Pod-side placement rule written against **node labels**. `spec.nodeSelector` is the primitive form: a flat map of exact key=value pairs, all of which must match, with no alternatives and no fallback. `spec.affinity.nodeAffinity` has two fields: - `requiredDuringSchedulingIgnoredDuringExecution` — a hard filter. It holds `nodeSelectorTerms`; terms are OR-ed, and the `matchExpressions` inside one term are AND-ed. Operators are `In`, `NotIn`, `Exists`, `DoesNotExist`, `Gt`, `Lt`, so you get sets and negation. - `preferredDuringSchedulingIgnoredDuringExecution` — soft. A list of `{weight: 1-100, preference}`; matching nodes gain that weight during scoring, and the Pod still schedules if nothing matches. `IgnoredDuringExecution` in both names means the rule is enforced only at placement time — relabelling a node afterwards never evicts the running Pod. Use nodeSelector for one trivial exact match; use node affinity when you need a set of acceptable values, a negation, or a preference with a fallback such as "prefer the spot pool, accept on-demand".
code
yaml · 29 linesapiVersion: v1
kind: Pod
metadata:
name: worker
spec:
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/arch
operator: In
values: ["amd64", "arm64"]
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
preference:
matchExpressions:
- key: pool
operator: In
values: ["spot"]
- weight: 1
preference:
matchExpressions:
- key: pool
operator: In
values: ["on-demand"]
containers:
- name: app
image: registry.example.com/app:1.4.0go deeper
Know that both match node labels, that nodeSelector is exact-match only, and that affinity has a hard required form and a soft preferred form.
Add the operator list, the OR/AND semantics of terms versus expressions, weights 1-100 in the scoring phase, and what IgnoredDuringExecution really means.
Talk about diagnosing FailedScheduling events, the prefer-spot-fallback-on-demand pattern, and making sure node-group templates carry the labels so autoscaling can react.
Frame it as a labelling contract: which labels are platform-owned versus team-owned, whether hard rules should be allowed at all, and how placement policy is enforced through admission rather than left to each Deployment.
## The problem The Kubernetes scheduler assigns every new Pod to a node. Left alone it picks any node with enough free CPU and memory. That is often not good enough: a Pod needs a GPU, must run on an ARM machine, or should land on a cheap spot-instance pool. Node labels plus node affinity are how you express that from the Pod side. ## Node labels Every Node object carries labels — key/value strings. The kubelet and cloud provider populate well-known ones: `kubernetes.io/os`, `kubernetes.io/arch`, `node.kubernetes.io/instance-type`, `topology.kubernetes.io/zone`, `topology.kubernetes.io/region`, `kubernetes.io/hostname`. Operators add their own, e.g. `pool=spot` or `hardware=gpu`. Placement rules are always written against labels, never node names. ## nodeSelector — the simple form `spec.nodeSelector` is a flat map. `{disktype: ssd}` means "only nodes whose `disktype` label equals `ssd`". Multiple entries are AND-ed. That is the entire feature: exact string equality, always mandatory, no OR, no negation, no preference. If no node matches, the Pod stays `Pending` forever. ## nodeAffinity — the expressive form Under `spec.affinity.nodeAffinity` there are two fields. **requiredDuringSchedulingIgnoredDuringExecution** is a hard rule evaluated in the scheduler's filter phase. It contains `nodeSelectorTerms`, a list. The list is OR-ed; the `matchExpressions` inside a single term are AND-ed. So you can express "(zone in [eu-west-1a, eu-west-1b]) OR (pool = fallback)". **preferredDuringSchedulingIgnoredDuringExecution** is a soft rule evaluated in the scoring phase. It is a list of `{weight, preference}` where weight is 1–100 and `preference` holds `matchExpressions`. Every node matching a preference gets that weight added to its score; the highest-scoring feasible node wins. Nothing matching is fine — the Pod still schedules. Each expression is `{key, operator, values}`. Operators: `In`, `NotIn`, `Exists`, `DoesNotExist`, `Gt`, `Lt`. `NotIn` and `DoesNotExist` give you node-level anti-affinity. `Gt`/`Lt` compare a single integer-valued label. There is also `matchFields` for a small set of non-label fields such as `metadata.name`. ## Reading the long field names Split each name in two. The first half — `requiredDuringScheduling` or `preferredDuringScheduling` — says when the rule applies: at placement. The second half — `IgnoredDuringExecution` — says what happens afterwards: nothing. If someone relabels the node so the Pod no longer matches, the running Pod stays put. A `RequiredDuringExecution` variant that would evict was proposed and never shipped; taints with the `NoExecute` effect are the mechanism that actually evicts running Pods, and that is a separate feature working from the node side. ## Choosing between them nodeSelector is fine and very readable when the rule is one exact match you never want a fallback for. Reach for node affinity when you need a set of acceptable values, a negation, an OR of alternatives, or — the most valuable case — a soft preference. A common production shape is: prefer `pool=spot` with weight 100, prefer `pool=on-demand` with weight 1, require neither. Workloads then drift onto cheap capacity but never go Pending when spot capacity disappears. The two can be combined on one Pod; both must be satisfied. Required node affinity and nodeSelector together are AND-ed. ## Failure mode and diagnosis A hard rule no node satisfies leaves the Pod `Pending`. `kubectl describe pod` shows a `FailedScheduling` event reading roughly "0/12 nodes are available: 12 node(s) didn't match Pod's node affinity/selector". The usual causes are a typo in the label key, a label applied to the wrong nodes, or a value written as an integer in YAML where the label is a string. ## Cost Node affinity is cheap: one pass over candidate nodes comparing label maps, both in filtering and scoring. It costs nothing like inter-pod affinity, which must inspect the Pods already placed across whole topology domains. ## Interaction with cluster autoscaling Cluster Autoscaler simulates node-group templates, and those templates carry labels, so a Pending Pod requiring `pool=gpu` can trigger a scale-up of that pool — provided the node group's template genuinely advertises the label. Labels bolted on by a boot script after the node registers are invisible to that simulation, and the Pod stays Pending.
- If someone removes the label a running Pod's required node affinity depends on, what happens to that Pod?Nothing — it keeps running. All current affinity rules end in IgnoredDuringExecution, so they are only evaluated when the scheduler places the Pod. The rule would only bite again if the Pod were recreated, at which point it might not schedule at all. To actively evict Pods from a node you need a taint with the NoExecute effect, which is a different mechanism.
- Inside nodeAffinity, how do multiple nodeSelectorTerms combine, and how do multiple matchExpressions inside one term combine?The list of nodeSelectorTerms is OR-ed: a node is feasible if it satisfies any one term. Within a single term, all matchExpressions must hold, so they are AND-ed. That asymmetry is the usual source of bugs — putting two expressions in separate terms accidentally loosens the rule instead of tightening it.
nodeSelector is a job ad saying "must live in Berlin". Node affinity is "must live in Germany or Austria (hard), Berlin preferred (weight 100), Munich acceptable (weight 10)" — and once you are hired, moving house does not get you fired.
saying these in an interview costs you the question
- Saying node affinity evicts Pods when node labels change — every current form is IgnoredDuringExecution.
- Claiming preferred affinity is a guarantee; it only adds weight during scoring and can be ignored entirely.
- Confusing node affinity with taints and tolerations — affinity attracts a Pod to nodes, tolerations only permit a Pod onto tainted nodes and never attract anything.
- Believing matchExpressions across separate nodeSelectorTerms are AND-ed (they are OR-ed).
- Thinking nodeSelector is deprecated — it is not, it is simply the limited form.