Kubernetes automatically adds taints such as `node.kubernetes.io/not-ready`, `node.kubernetes.io/unreachable` and `node.kubernetes.io/disk-pressure` to nodes. Explain what adds them, what they do to running pods, and how you would use them when operating a cluster.
answer
- node.kubernetes.io/ prefix is reserved for condition taints
- not-ready + unreachable = NoExecute; pressure taints = NoSchedule
- Failover = node-monitor grace + tolerationSeconds (≈300s default)
- Pressure taints never evict — kubelet eviction manager does that
- DaemonSets auto-tolerate built-ins, not your custom taints
basics
~20 sThe node lifecycle controller taints nodes from their reported conditions. not-ready and unreachable are NoExecute, so pods are evicted once their injected 300-second toleration expires; pressure taints such as disk-pressure and memory-pressure are NoSchedule and only keep new pods away. DaemonSets tolerate them so node agents keep running.
solid answer
~50 sThe node lifecycle controller in `kube-controller-manager` mirrors node conditions into taints under the `node.kubernetes.io/` prefix: - **NoExecute**: `not-ready` (kubelet reports NotReady) and `unreachable` (no kubelet heartbeat). Pods are evicted when their toleration expires — by default 300 seconds, injected by the DefaultTolerationSeconds admission plugin. - **NoSchedule**: `memory-pressure`, `disk-pressure`, `pid-pressure`, `network-unavailable`, `unschedulable` (set by `kubectl cordon`). These stop new pods arriving; they never evict. Eviction under resource pressure is the kubelet's own eviction manager, a different mechanism. Operationally this means: node failure timing is *node-monitor grace period + tolerationSeconds*, not instant. `kubectl cordon` works through the `unschedulable` taint; `kubectl drain` adds eviction on top. DaemonSet pods get automatic tolerations for these taints so logging, CNI and CSI agents keep running on a degraded node. When debugging "why did my pods move / why is nothing scheduling here", `kubectl describe node` and its Taints block is the first place to look.
code
bash · 4 lineskubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints
kubectl describe node worker-4 | sed -n '/Conditions/,/Addresses/p'
kubectl cordon worker-4
kubectl drain worker-4 --ignore-daemonsets --delete-emptydir-datago deeper
Recognise that Kubernetes taints unhealthy nodes itself and that this is why pods move away from a broken node; name not-ready and unreachable.
Split the list into NoExecute (not-ready, unreachable) and NoSchedule (pressure, unschedulable) and explain the 300-second delay before eviction.
Walk the full failover timeline, distinguish taint eviction from kubelet resource eviction, explain cordon vs drain vs NoExecute taint, and know DaemonSets auto-tolerate the built-ins.
Treat it as a failure-detection policy question: what failover latency the business needs, per-workload toleration overrides versus global flags, eviction-storm risk during control-plane partitions, and using custom taints from a health agent to automate evacuation of bad hardware.
## Conditions become taints Every node reports conditions: `Ready`, `MemoryPressure`, `DiskPressure`, `PIDPressure`, `NetworkUnavailable`. The node lifecycle controller inside `kube-controller-manager` watches these and maintains a corresponding set of taints on `node.spec.taints`, all under the reserved `node.kubernetes.io/` prefix. This means the scheduler needs to understand only one concept — taints — instead of special-casing each condition, and it means you can reason about node health with the same vocabulary you use for dedication. The main ones: | Taint key | Effect | Meaning | |---|---|---| | `node.kubernetes.io/not-ready` | NoExecute | Node's Ready condition is False | | `node.kubernetes.io/unreachable` | NoExecute | Ready condition is Unknown — no kubelet heartbeat | | `node.kubernetes.io/memory-pressure` | NoSchedule | Kubelet reports memory pressure | | `node.kubernetes.io/disk-pressure` | NoSchedule | Kubelet reports disk pressure | | `node.kubernetes.io/pid-pressure` | NoSchedule | Process-ID exhaustion | | `node.kubernetes.io/network-unavailable` | NoSchedule | Network not correctly configured (set by cloud/CNI) | | `node.kubernetes.io/unschedulable` | NoSchedule | Node cordoned (`spec.unschedulable`) | Cloud provider integrations add `node.cloudprovider.kubernetes.io/uninitialized` with NoSchedule while the cloud controller manager is still filling in labels and addresses, so pods are not placed on a half-registered node. ## What actually happens when a node dies 1. The kubelet stops posting node status leases. 2. After the node-monitor grace period, the node controller flips Ready to Unknown and adds `node.kubernetes.io/unreachable:NoExecute`. 3. The scheduler immediately stops placing new pods there (the taint filters it out). 4. The taint manager begins evicting pods whose toleration for that key has expired. Because the DefaultTolerationSeconds admission plugin injected `tolerationSeconds: 300` into pods that did not declare their own, that is about five more minutes. 5. Owning controllers create replacements, which land elsewhere. So end-to-end failover is grace period plus toleration seconds — commonly six-ish minutes out of the box. Teams that need faster failover override the toleration per workload rather than shortening the cluster-wide default, because short global values cause eviction storms when the control plane or network hiccups and many nodes look unreachable at once. ## Pressure taints do not evict A frequent confusion: `disk-pressure` and `memory-pressure` are `NoSchedule`, so they only stop *new* pods arriving. The pods already there are dealt with by a completely separate mechanism — the **kubelet eviction manager**, which watches its own soft/hard eviction thresholds and terminates pods locally in QoS order (BestEffort first, then Burstable over requests, Guaranteed last). Do not claim the disk-pressure taint evicts anything. ## Cordon, drain and quarantine `kubectl cordon` sets `spec.unschedulable`, which surfaces as the `unschedulable:NoSchedule` taint — new pods stay away, existing ones keep running. `kubectl drain` cordons and then evicts through the Eviction API, which *does* respect PodDisruptionBudgets and is the safe, voluntary path for maintenance. Manually tainting a node `NoExecute` is the blunt instrument: it deletes pods without consulting budgets, appropriate for a node you believe is actively harming traffic. ## Why DaemonSets keep running A degraded node still needs its CNI agent, CSI node plugin, log shipper and node exporter — arguably more than a healthy one. The DaemonSet controller therefore adds tolerations to its pods for `not-ready`, `unreachable`, `disk-pressure`, `memory-pressure`, `pid-pressure` and `unschedulable`. Note this covers only the built-in taints: your own custom taints must be tolerated explicitly in the DaemonSet spec. ## Using them yourself The prefix `node.kubernetes.io/` is reserved for the system — do not invent keys under it. But the pattern generalises: a node-problem-detector style agent can watch for kernel deadlocks, failing NICs or degraded disks and apply your own taints (`hardware=degraded:NoExecute`), which gives you automatic workload evacuation from bad hardware without a human in the loop. Pair it with alerting so a node quietly evacuating itself is visible. ## Debugging checklist - `kubectl get nodes` — Ready / SchedulingDisabled column. - `kubectl describe node <n>` — Conditions and Taints blocks together tell the whole story. - A pending pod's events name the exact untolerated taint, e.g. `1 node(s) had untolerated taint {node.kubernetes.io/disk-pressure: }`. - If pods vanished from a node, correlate the node's condition transition timestamps with the pod deletion timestamps; the gap should equal the toleration seconds.
- A node reports DiskPressure. Why do its existing pods still get terminated even though the disk-pressure taint is only NoSchedule?Two different mechanisms are at work. The `node.kubernetes.io/disk-pressure:NoSchedule` taint only prevents new pods being scheduled there. The terminations come from the kubelet's own eviction manager, which enforces its configured soft and hard eviction thresholds locally and kills pods in QoS order — BestEffort first, then Burstable exceeding requests, Guaranteed last. Neither path consults PodDisruptionBudgets.
- What is the difference between kubectl cordon, kubectl drain, and manually tainting a node NoExecute?Cordon sets `spec.unschedulable`, which appears as an `unschedulable:NoSchedule` taint — no new pods, existing ones untouched. Drain cordons and then evicts pods through the Eviction API, so PodDisruptionBudgets are honoured and the departure is graceful; it is the correct maintenance path. Tainting NoExecute makes the taint manager delete pods directly, ignoring budgets — the blunt option for a node you want emptied immediately regardless of availability guarantees.
saying these in an interview costs you the question
- Saying the disk-pressure or memory-pressure taint evicts running pods
- Thinking pods are removed the instant a node goes NotReady, with no grace period
- Assuming DaemonSets tolerate custom taints automatically the way they tolerate built-in ones
- Creating custom taints under the reserved node.kubernetes.io/ prefix
- Confusing cordon (no new pods) with drain (also evicts, respecting PodDisruptionBudgets)