In Kubernetes, what is a node taint and what is a pod toleration, and what do the NoSchedule, PreferNoSchedule and NoExecute taint effects each do?
answer
- Taint = node repels; toleration = pod's permission slip
- NoSchedule blocks new · PreferNoSchedule soft · NoExecute also evicts
- Toleration is permission, not attraction
- key=value:effect · operator Equal vs Exists
- Dedication = taint + label/nodeSelector, both sides
basics
~20 sA taint marks a node so it repels pods; a matching toleration on a pod lets that pod ignore the taint. NoSchedule blocks new pods, PreferNoSchedule is a soft preference to schedule elsewhere, and NoExecute also evicts already-running pods that do not tolerate it.
solid answer
~50 sTaints are set on the **node** (`node.spec.taints`) as `key=value:effect` and repel pods. Tolerations are set on the **pod** (`spec.tolerations`) and grant that pod an exemption from a matching taint. The three effects: - **NoSchedule** — a hard scheduling filter. The scheduler will not place a pod on the node unless the pod tolerates the taint. Pods already running there are untouched. - **PreferNoSchedule** — soft. The scheduler penalizes the node during scoring but will still use it if nothing better exists. - **NoExecute** — hard filter *and* eviction: pods already running on the node that do not tolerate the taint are evicted by the taint manager. The key mental model is that a toleration is *permission, not attraction*. It only stops the node being filtered out; it never pulls the pod there. To actually pin a workload to those nodes you pair the toleration with `nodeSelector` or node affinity.
code
bash · 3 lineskubectl taint nodes worker-3 dedicated=payments:NoSchedule
kubectl describe node worker-3 | grep -A3 Taints
kubectl taint nodes worker-3 dedicated=payments:NoSchedule-go deeper
Be able to state which object carries which field (taint on node, toleration on pod), name the three effects, and show the kubectl taint command.
Add the matching semantics — Equal vs Exists, empty key/effect wildcards, all taints must be tolerated — and explain why toleration alone does not pin a pod to a node.
Frame taints as the only node-side exclusion mechanism, describe the dedication recipe (taint + label + nodeSelector), and mention condition taints, drain/cordon and how you read the scheduler's Pending reason.
Discuss taints as a capacity-partitioning policy: who owns the taint keys, whether pools are tainted at node-pool creation so no window exists, the fragmentation and cost of exclusive pools, and how blanket tolerations erode the guarantee cluster-wide.
## Why taints exist By default the Kubernetes scheduler considers every schedulable, Ready node a candidate for every pending pod. Pod-side mechanisms such as `nodeSelector` and node affinity let a *pod* say "I want to run on nodes like that" — but they are purely attraction, and they do nothing to keep *other* pods off those nodes. Taints are the mirror image: a property on the **node** that repels pods unless the pod explicitly opts in. If you need exclusivity — a GPU pool, a licensed-software pool, control-plane nodes, a tenant's dedicated hardware — taints are the only mechanism that provides it. ## Anatomy of a taint A taint is a triple: `key=value:effect`. The value is optional (`key:effect` is legal). It lives in `node.spec.taints` and is normally applied with the CLI: `kubectl taint nodes worker-3 dedicated=payments:NoSchedule` A trailing minus removes it: `kubectl taint nodes worker-3 dedicated=payments:NoSchedule-`. In managed clusters you usually set taints on the node *pool* definition so replacement nodes are born tainted rather than being tainted after they join — otherwise there is a window where arbitrary pods land on a fresh node. ## The three effects in detail **NoSchedule** acts during the scheduler's filter phase: any node carrying an un-tolerated `NoSchedule` taint is removed from the feasible set. It is retroactive to nothing — pods already bound to the node keep running happily. This is the effect you use for node dedication. **PreferNoSchedule** acts during the score phase. The `TaintToleration` scoring plugin lowers the node's score, so the scheduler avoids the node when alternatives exist but will still use it under pressure. It is a hint, not a guarantee, and it is the wrong tool when you need real isolation. **NoExecute** does both jobs. It filters at schedule time, and the node lifecycle / taint manager component of `kube-controller-manager` continuously evicts running pods on that node whose tolerations do not match. This is what makes taints usable as a live drain / quarantine mechanism: taint a misbehaving node `NoExecute` and its workloads move away. ## Anatomy of a toleration A toleration is an entry in `pod.spec.tolerations` with `key`, `operator`, `value`, `effect` and `tolerationSeconds`. The matching rules: - `operator: Equal` (the default) matches when key **and** value are equal. - `operator: Exists` matches any value for that key; you must leave `value` empty. - An **empty `key` with `operator: Exists`** matches *every* taint — a blanket exemption, used sparingly (DaemonSets, some cluster agents). - An **empty `effect`** matches all effects for that key. - `tolerationSeconds` is only meaningful together with `NoExecute`; it turns the exemption into a timed one. A node may carry several taints. The pod must tolerate **all** the `NoSchedule` and `NoExecute` ones to be feasible; un-tolerated `PreferNoSchedule` taints only cost score. ## Toleration is permission, not attraction This is the single most common misunderstanding. Adding `tolerations: [{key: dedicated, operator: Exists}]` to a Deployment does not send it to the dedicated nodes — it merely stops those nodes being excluded. The pod remains equally eligible for every untainted node in the cluster and will usually land on one of them. The complete dedication recipe is two-sided: 1. **Taint** the nodes so nothing else can come in. 2. **Label** the nodes and add a `nodeSelector` or `requiredDuringSchedulingIgnoredDuringExecution` node affinity so your workload is pulled in. ## Where you meet taints in a real cluster Control-plane nodes are tainted `node-role.kubernetes.io/control-plane:NoSchedule` by kubeadm so user workloads stay off them. GPU pools carry something like `nvidia.com/gpu=true:NoSchedule`. Kubernetes itself adds condition taints such as `node.kubernetes.io/not-ready` and `node.kubernetes.io/unreachable` with `NoExecute`. `kubectl drain` works by cordoning and evicting, and `kubectl cordon` sets `spec.unschedulable`, which the scheduler honours much like a taint. DaemonSet pods are given broad tolerations automatically so node agents keep running on nodes that are repelling everything else. ## Debugging When a pod is stuck `Pending`, `kubectl describe pod` prints the scheduler's reason verbatim, e.g. `0/6 nodes are available: 3 node(s) had untolerated taint {dedicated: payments}, 3 Insufficient cpu`. Check the node side with `kubectl describe node worker-3 | grep -A3 Taints` or `kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints`.
- If you add a NoSchedule taint to a node that already has ten pods running on it, what happens to those pods?Nothing — they keep running. NoSchedule is evaluated only when the scheduler is choosing a node for a pending pod, so it is not retroactive. If you want the existing pods off the node you either use the NoExecute effect, or drain the node with `kubectl drain`, which cordons it and evicts the pods through the Eviction API so PodDisruptionBudgets are respected.
- What does a toleration with an empty key, operator Exists and no effect match?Everything. An empty key with `operator: Exists` matches any taint key, and an empty `effect` matches all three effects, so the pod is exempt from every taint in the cluster including the built-in not-ready and unreachable ones. It is appropriate for a handful of node-level agents but is a red flag on an application workload, because it lets the pod be scheduled onto control-plane or quarantined nodes.
A taint is a velvet rope on a club door; a toleration is being on the guest list. Being on the list lets you in — it does not mean you will go to that club rather than the bar next door.
saying these in an interview costs you the question
- Saying a toleration forces or prefers the pod onto the tainted node — it only removes the node from the exclusion list
- Claiming NoSchedule evicts pods that are already running (only NoExecute does)
- Confusing the two sides: putting the taint on the pod and the toleration on the node
- Treating PreferNoSchedule as isolation — it is only a scoring penalty and pods can still land there
- Believing a pod must tolerate only one of several taints on a node; every NoSchedule/NoExecute taint must be tolerated