skip to content

A pod stays Pending with the scheduler event `node(s) had untolerated taint {dedicated=gpu: NoSchedule}` and, on the remaining nodes, `node(s) didn't match Pod's node affinity/selector`. Explain both mechanisms and how you would resolve each.

level: middleimportance: must knowfreq 58%

answer

  1. taint = node repels; toleration = permission, not attraction
  2. effects: NoSchedule / PreferNoSchedule / NoExecute
  3. nodeSelector exact; nodeAffinity required vs preferred
  4. IgnoredDuringExecution = label changes don't evict
  5. dedicated pool = taint + label + affinity

basics

~20 s

Taints are set on nodes to repel pods; a pod needs a matching toleration to be allowed there. nodeSelector/nodeAffinity are set on the pod to require node labels. Taints repel, selectors attract — you usually need both: add the toleration and make sure node labels match.

solid answer

~50 s

They are opposite-direction controls. **Taints (node-side, repel):** `kubectl taint node n1 dedicated=gpu:NoSchedule` keeps every pod off unless the pod carries a matching **toleration**. Effects: `NoSchedule` (no new pods), `PreferNoSchedule` (soft), `NoExecute` (also evicts running pods). **nodeSelector / nodeAffinity (pod-side, attract):** the pod demands node labels. `nodeSelector` is exact-match; `nodeAffinity` adds operators (`In`, `NotIn`, `Exists`, `Gt`) and two strengths — `requiredDuringSchedulingIgnoredDuringExecution` (hard filter) and `preferredDuringScheduling…` (weighted preference only). Fixes: - Untolerated taint → add the toleration to the pod (key, value, effect must match; `operator: Exists` matches any value), or remove the taint with `kubectl taint node n1 dedicated-`. - Affinity mismatch → check the actual labels (`kubectl get nodes --show-labels`); either relabel nodes or correct/loosen the selector. Typos and stale keys like `beta.kubernetes.io/os` are the usual culprits. A toleration only *permits*; it does not attract. Dedicating hardware needs taint **plus** a label and matching affinity, or other pods will land there too.

code

yaml · 21 lines
yaml
spec:
  tolerations:
    - key: dedicated
      operator: Equal
      value: gpu
      effect: NoSchedule
  affinity:
    nodeAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:
        nodeSelectorTerms:
          - matchExpressions:
              - key: hardware
                operator: In
                values: ["gpu"]
      preferredDuringSchedulingIgnoredDuringExecution:
        - weight: 50
          preference:
            matchExpressions:
              - key: topology.kubernetes.io/zone
                operator: In
                values: ["eu-west-1a"]

go deeper

for a junior

State the direction of each mechanism: taints on nodes repel, tolerations permit, nodeSelector on pods requires node labels. Know how to inspect both.

for a middle

Name the three taint effects, the required-versus-preferred affinity strengths and their operators, and explain why dedicating hardware needs a taint plus a label plus affinity.

for a senior

Diagnose from the event text, distinguish an intended taint from an automatic condition/cordon taint, and know when to downgrade a hard affinity to a preference to avoid Pending in degraded conditions.

for a principal

Design the node-pool topology: which workloads get dedicated pools, who owns taints and labels, how tolerations are governed so they are not copy-pasted everywhere, and how the scheme survives node-pool replacement.

## Two directions of control Kubernetes gives you a repelling mechanism owned by the node and an attracting mechanism owned by the pod. Confusing them is the single most common source of `Pending`. ### Taints and tolerations — the node repels A **taint** is a property of a node: a key, an optional value, and an effect. ``` kubectl taint node gpu-1 dedicated=gpu:NoSchedule ``` Effects: - **`NoSchedule`** — the scheduler will not place a pod here unless it tolerates the taint. - **`PreferNoSchedule`** — a soft version; the scheduler avoids the node if it can. - **`NoExecute`** — as `NoSchedule`, and additionally evicts pods already running that do not tolerate it. A **toleration** in the pod spec is permission to ignore a taint: ```yaml tolerations: - key: dedicated operator: Equal value: gpu effect: NoSchedule ``` Matching rules: `operator: Equal` requires key, value, and effect to match; `operator: Exists` matches any value for that key; omitting `effect` tolerates all effects of that key; an empty `key` with `operator: Exists` tolerates *everything* (used by cluster-critical DaemonSets, and dangerous elsewhere). Kubernetes itself uses taints for control-plane nodes (`node-role.kubernetes.io/control-plane:NoSchedule`) and for node conditions (`node.kubernetes.io/memory-pressure:NoSchedule`, `not-ready:NoExecute`). ### nodeSelector and nodeAffinity — the pod attracts `nodeSelector` is the simple form: a map of label keys to values, all of which the node must have. ```yaml nodeSelector: disktype: ssd ``` `nodeAffinity` is the expressive form: - **`requiredDuringSchedulingIgnoredDuringExecution`** — a hard filter. A pod is only placed on a matching node. `nodeSelectorTerms` are OR-ed; `matchExpressions` within a term are AND-ed. Operators: `In`, `NotIn`, `Exists`, `DoesNotExist`, `Gt`, `Lt`. - **`preferredDuringSchedulingIgnoredDuringExecution`** — a scoring preference with a `weight` (1–100). It never causes `Pending`; it only nudges placement. `IgnoredDuringExecution` in both names means: once the pod is running, later label changes on the node do not evict it. Only `NoExecute` taints move running pods. ## Why you usually need both Suppose you buy GPU nodes for machine-learning jobs. - **Taint them** so ordinary web pods do not get scheduled onto expensive hardware. - **Label them** (`hardware=gpu`) and give the ML pods a **nodeAffinity/nodeSelector** for that label, so the ML pods actually go there. With only the taint, a tolerating pod *may* land there but may equally land anywhere else. With only the label and affinity, other pods still consume the GPU nodes. The pairing — taint to exclude, label plus affinity to attract — is the standard dedicated-node-pool idiom, and interviewers ask for exactly this reasoning. ## Diagnosing the two messages **`node(s) had untolerated taint {dedicated=gpu: NoSchedule}`** ``` kubectl get nodes -o json | jq '.items[] | {name:.metadata.name, taints:.spec.taints}' kubectl describe node gpu-1 | grep -i taints ``` Then decide: is the taint correct (add a toleration to the pod) or accidental — for example a leftover maintenance taint or a node-condition taint like `disk-pressure` that reveals a sick node? Remove an obsolete one with a trailing dash: `kubectl taint node gpu-1 dedicated-`. **`node(s) didn't match Pod's node affinity/selector`** ``` kubectl get nodes --show-labels kubectl get nodes -l disktype=ssd kubectl get pod api-1 -o jsonpath='{.spec.nodeSelector}{.spec.affinity.nodeAffinity}' ``` Compare literally. The usual causes are a typo, a value case mismatch, a label that only exists in one environment, a node pool that was replaced without re-applying labels, or a deprecated key (`beta.kubernetes.io/instance-type` versus `node.kubernetes.io/instance-type`). Either relabel the node (`kubectl label node n1 disktype=ssd`) or fix the pod spec — and if the requirement is a nice-to-have rather than a necessity, downgrade it from `required…` to `preferred…` so it can never cause `Pending`. ## A note on cordoned nodes `kubectl cordon` sets `spec.unschedulable: true` and also adds `node.kubernetes.io/unschedulable:NoSchedule`. The scheduler message is `node(s) were unschedulable`. It is a related but distinct reason — uncordon with `kubectl uncordon`. ## Interview framing State the direction of each mechanism in one sentence ("taints repel from the node side, selectors attract from the pod side"), give the three taint effects and the two affinity strengths, explain that a toleration is permission and not attraction, and describe the dedicated-node-pool pattern that needs both. Then show the two `kubectl` commands you would actually run.

  • If you add the toleration but the pod still lands on ordinary nodes, what is missing?
    A toleration is only permission to ignore a taint — it exerts no pull. Without a nodeSelector or required nodeAffinity matching a label unique to the intended nodes, the scheduler is free to place the pod anywhere it fits. Add the label to those nodes and the matching affinity to the pod.
  • What does `IgnoredDuringExecution` mean in `requiredDuringSchedulingIgnoredDuringExecution`?
    The rule is enforced only at scheduling time. If someone later removes or changes the node label the pod required, the already-running pod is not evicted — Kubernetes does not currently implement a `RequiredDuringExecution` variant for node affinity. The only mechanism that removes running pods based on node properties is a `NoExecute` taint.

saying these in an interview costs you the question

  • Believing a toleration attracts a pod to the tainted node
  • Thinking nodeSelector is set on the node rather than the pod
  • Confusing NoSchedule with NoExecute when explaining what happens to running pods
  • Assuming `preferredDuringScheduling…` can leave a pod Pending — it cannot
  • Forgetting that node-condition and cordon taints are added automatically and can be the real reason

context