A Kubernetes node goes `NotReady`, and after roughly five minutes its pods are deleted and recreated elsewhere. Explain the mechanism that does this, and how it differs from a kubelet evicting pods because the node is short on memory.
answer
- lease stops → Ready=Unknown → NotReady
- taints not-ready / unreachable, effect NoExecute
- DefaultTolerationSeconds admission → tolerationSeconds 300
- taint manager deletes pod object; containers may live on
- control plane + health-driven vs kubelet + resource-driven
basics
~20 sThe node controller adds a NoExecute taint (node.kubernetes.io/not-ready or unreachable). Pods carry a default toleration with tolerationSeconds 300, so after 5 minutes the taint manager deletes them and controllers recreate them elsewhere. Kubelet node-pressure eviction is different: local, resource-driven, ranked by QoS.
solid answer
~60 sWhen a node stops heartbeating, the **node controller** in kube-controller-manager marks it `Ready=Unknown`/`False` and applies a `NoExecute` taint — `node.kubernetes.io/unreachable` or `node.kubernetes.io/not-ready`. `NoExecute` means *existing* pods that do not tolerate the taint are evicted. The API server's DefaultTolerationSeconds admission plugin injects a toleration for both taints with `tolerationSeconds: 300` into every pod that lacks one, so the taint manager waits five minutes and then deletes the pods. Their controllers create replacements, which the scheduler places on healthy nodes. Differences from kubelet node-pressure eviction: | | Taint-based (NoExecute) | Node-pressure | |---|---|---| | Actor | control plane (node controller + taint manager) | kubelet on the node | | Trigger | node unhealthy/unreachable, or an operator-applied taint | memory/disk/PID signal breached | | Victims | all pods without a matching toleration | ranked by usage-vs-request, priority, QoS | | Timing | `tolerationSeconds` (default 300) | immediate (hard) or after soft grace | Neither honours PodDisruptionBudgets. If the node is merely unreachable, its pods may still be running — which is why StatefulSet pods are not force-recreated without fencing.
code
yaml · 10 linesspec:
tolerations:
- key: node.kubernetes.io/not-ready
operator: Exists
effect: NoExecute
tolerationSeconds: 30
- key: node.kubernetes.io/unreachable
operator: Exists
effect: NoExecute
tolerationSeconds: 30go deeper
Recall that an unhealthy node gets tainted and its pods are moved after a few minutes, and that a controller recreates them elsewhere.
Name the not-ready/unreachable NoExecute taints, the injected 300-second toleration, and the distinction between NoSchedule and NoExecute.
Explain the partition ambiguity, why pods are deleted but containers may survive, why StatefulSets need fencing, and how this differs in actor and victim-selection from kubelet node-pressure eviction.
Reason about failover policy fleet-wide: tolerationSeconds tuning versus flap sensitivity, capacity headroom to absorb a node's worth of rescheduling, and fencing/storage-detach requirements for stateful tiers.
## The heartbeat Each kubelet reports liveness to the API server through a **Lease** object in `kube-node-lease`, renewed every few seconds. The **node controller** inside kube-controller-manager watches those leases. If a node's lease stops being renewed for `node-monitor-grace-period` (40s in typical upstream configuration), the controller sets the node's `Ready` condition to `Unknown` — this is what `kubectl get nodes` shows as `NotReady`. The node may be genuinely dead (hardware failure, kernel panic), or merely **partitioned** — still running every container, just unable to talk to the control plane. From the control plane's viewpoint these are indistinguishable, and that ambiguity drives everything that follows. ## Taint-based eviction On marking the node unhealthy, the node controller applies a taint: - `node.kubernetes.io/not-ready:NoExecute` when `Ready=False`, - `node.kubernetes.io/unreachable:NoExecute` when `Ready=Unknown`. A **taint** repels pods; a **toleration** in a pod spec lets it stay. Taint effects: `NoSchedule` (do not place new pods), `PreferNoSchedule` (avoid if possible), and `NoExecute` (also remove pods **already running** that do not tolerate it). The component that acts on `NoExecute` is the **taint manager**. For each running pod on the tainted node it checks the pod's tolerations: - no matching toleration → delete immediately; - matching toleration with no `tolerationSeconds` → stay forever; - matching toleration with `tolerationSeconds: N` → delete after N seconds. ## Where the five minutes comes from The API server runs the **DefaultTolerationSeconds** admission plugin. Every pod created without explicit tolerations for these two taints gets: ```yaml tolerations: - key: node.kubernetes.io/not-ready operator: Exists effect: NoExecute tolerationSeconds: 300 - key: node.kubernetes.io/unreachable operator: Exists effect: NoExecute tolerationSeconds: 300 ``` Hence the familiar ~5-minute delay before pods on a dead node are recreated (roughly 40s to notice plus 300s of toleration). The value is tunable per pod, or cluster-wide via the API server's `--default-not-ready-toleration-seconds` / `--default-unreachable-toleration-seconds`. Lowering it recovers faster but reacts to transient network blips; raising it tolerates flaky networks at the cost of longer outages. DaemonSet pods are given tolerations without `tolerationSeconds` for these taints, so they stay put — they belong to the node, and there is nowhere to move them. ## Deleted, not stopped The taint manager deletes the **pod object**. If the node is unreachable, the kubelet never receives the delete, so the containers may keep running — the pod becomes an orphan still holding whatever the workload holds. Kubernetes calls the deleted-but-possibly-alive state exactly why StatefulSet replicas are not automatically recreated for unreachable nodes: recreating a pod with the same identity while the original might still be writing risks split-brain. Forcing it (`kubectl delete pod --force --grace-period=0`) tells the API server to drop the record without confirmation and is safe only once you have genuinely fenced the node (powered it off, detached its volumes). ## Contrast with node-pressure eviction Node-pressure eviction is a **local, resource-driven** action by the kubelet: a signal like `memory.available` breaches a threshold, the kubelet ranks pods by usage-versus-request and priority, and kills enough of them to recover. The node stays `Ready`; only some pods die; the evicted pods leave `Evicted` tombstones behind. Taint-based eviction is a **remote, health-driven** action by the control plane: the node's health, not its resource usage, is the trigger; *all* non-tolerating pods go, regardless of QoS; the pods are deleted outright rather than left as `Evicted` records. Both ignore PodDisruptionBudgets — PDBs constrain only API-initiated eviction through the `/eviction` subresource, which is what `kubectl drain` and cluster autoscalers use. ## Operator-applied NoExecute taints The same machinery serves deliberate operations: `kubectl taint node n1 maintenance=true:NoExecute` drains non-tolerating pods off a node. Pressure conditions, by contrast, generate `NoSchedule` taints (`node.kubernetes.io/memory-pressure`, `disk-pressure`) — they stop new placements while the kubelet handles removals locally. ## Interview framing Name the actors (node controller → taint → taint manager), explain the 300-second default toleration as the source of the five minutes, note that pods are *deleted* while containers may still run on a partitioned node, and contrast trigger/actor/victim-selection with kubelet node-pressure eviction.
- Why are StatefulSet pods not automatically recreated when their node is unreachable?An unreachable node may still be running the containers — the control plane cannot tell a dead machine from a network partition. Recreating a StatefulSet replica with the same identity and volume while the original might still be writing risks split-brain data corruption, so Kubernetes waits for the pod to be confirmed gone. Recovery requires fencing the node or force-deleting the pod once you know it is truly dead.
- What is the risk of lowering `tolerationSeconds` to something like 10 seconds cluster-wide?Short network blips or a briefly overloaded kubelet would be treated as node failure, mass-deleting pods and rescheduling them elsewhere. That causes churn, cold caches, and possibly a rescheduling storm that overloads healthy nodes. Lower it selectively for workloads that genuinely need fast failover and can restart cheaply.
saying these in an interview costs you the question
- Attributing NotReady pod removal to the kubelet — the kubelet is unreachable or dead; the control plane acts
- Believing PodDisruptionBudgets delay taint-based eviction
- Assuming pod deletion means the containers definitely stopped on a partitioned node
- Confusing NoSchedule with NoExecute — only NoExecute removes running pods
- Not knowing the 5-minute default comes from an injected tolerationSeconds value