skip to content

In a Kubernetes pod spec, what does the `tolerationSeconds` field on a toleration do, and how does it interact with a taint whose effect is NoExecute?

level: middleimportance: must knowfreq 58%

answer

  1. tolerationSeconds only means anything with NoExecute
  2. Clock starts when the taint appears, not at pod start
  3. Omitted = forever · 0 = evict now
  4. DefaultTolerationSeconds admission injects not-ready/unreachable at 300s
  5. Taint-manager eviction ignores PodDisruptionBudgets

basics

~20 s

tolerationSeconds makes a toleration temporary: with a NoExecute taint the pod is allowed to keep running for that many seconds after the taint appears, then it is evicted. Omitted means tolerate forever; zero or negative means evict immediately. It is ignored for other effects.

solid answer

~50 s

`tolerationSeconds` only has meaning for the **NoExecute** effect. NoExecute both blocks scheduling and makes the taint manager in `kube-controller-manager` evict running pods that do not tolerate the taint. - **No `tolerationSeconds`** — the toleration is unbounded; the pod stays as long as it likes. - **`tolerationSeconds: N`** — the pod is bound to the node for N seconds after the matching taint is added, then the taint manager deletes it. If the taint is removed before the timer fires, nothing happens and the timer is cancelled. - **`tolerationSeconds: 0`** — evict as soon as the taint appears. The field is why node failure has a delay by default: the API server's admission plugin injects `node.kubernetes.io/not-ready` and `node.kubernetes.io/unreachable` tolerations with `tolerationSeconds: 300` into pods that lack them, so a pod survives a five-minute blip rather than being torn down on the first missed heartbeat. Eviction here is a straight deletion by the controller — it does **not** consult PodDisruptionBudgets.

code

yaml · 10 lines
yaml
spec:
  tolerations:
    - key: node.kubernetes.io/not-ready
      operator: Exists
      effect: NoExecute
      tolerationSeconds: 30
    - key: node.kubernetes.io/unreachable
      operator: Exists
      effect: NoExecute
      tolerationSeconds: 30

go deeper

for a junior

Know that tolerationSeconds is a countdown that applies to NoExecute only, and that omitting it means the pod may stay indefinitely.

for a middle

Explain the timer semantics including cancellation when the taint is removed, and connect the 300-second default to the injected not-ready/unreachable tolerations.

for a senior

Add the operational picture: node-monitor grace period plus toleration seconds equals real failover time, taint eviction bypasses PDBs, and StatefulSet pods can stick in Terminating on an unreachable node.

for a principal

Discuss failover-time budgets as a cluster policy — per-workload overrides versus changing the API server defaults, the eviction-storm risk during partial control-plane outages, and how spot-instance termination handling plugs into the same mechanism.

## The NoExecute effect The `NoSchedule` and `PreferNoSchedule` taint effects are pure scheduling-time constructs. `NoExecute` is different: besides filtering the node out for pending pods, it activates a controller loop — commonly called the *taint manager*, part of the node lifecycle controller inside `kube-controller-manager` — that watches nodes and pods and removes pods that are running on a node carrying an un-tolerated `NoExecute` taint. That is what makes taints usable as a live quarantine mechanism instead of only a placement rule. ## What tolerationSeconds adds Without `tolerationSeconds`, a toleration is a permanent exemption: the pod can sit on the tainted node indefinitely. `tolerationSeconds` converts it into a *grace period* — "I can live with this taint, but only for so long". The semantics: - The clock starts when the matching taint is **added** to the node, not when the pod started. - When the timer expires the taint manager deletes the pod. The workload controller behind it (Deployment, StatefulSet, Job) then creates a replacement, which the scheduler places on some other node — the tainted node has been filtered out for it. - If the taint is removed before expiry, the eviction is cancelled. A node that flaps Ready → NotReady → Ready inside the window costs you nothing. - `0` (or a negative value) means evict at once. - The field is **only** consulted for `NoExecute`. Setting it alongside `NoSchedule` is legal YAML but has no effect, which is a common source of confusion in review. ## Why every pod already has one The node lifecycle controller automatically taints nodes on condition changes. The two that matter most are `node.kubernetes.io/not-ready` (the node reports NotReady) and `node.kubernetes.io/unreachable` (the node controller has stopped hearing from the kubelet), both with effect `NoExecute`. If pods had no tolerations for these, a single missed heartbeat would immediately delete every pod on the node. To avoid that, the `DefaultTolerationSeconds` admission plugin in the API server injects tolerations for both keys with `tolerationSeconds: 300` into any pod that does not already declare them. That 300 seconds is where the familiar "pods take about five minutes to move after a node dies" behaviour comes from — and it is *in addition to* the node-monitor grace period the controller waits before deciding the node is unreachable at all, so real-world failover is often closer to six minutes. The default is cluster-wide and tunable through the API server flags `--default-not-ready-toleration-seconds` and `--default-unreachable-toleration-seconds`. A pod can override it individually by declaring its own toleration for those keys with a smaller value, which is the standard trick for latency-sensitive stateless services that want faster failover. Lowering it globally is risky: short values make the cluster reschedule aggressively during transient control-plane or network hiccups, causing eviction storms exactly when the cluster is least healthy. ## Eviction here bypasses PodDisruptionBudgets This catches people out. There are two very different things called "eviction": - **API-initiated eviction** — what `kubectl drain` and descheduling tools use. It goes through the Eviction subresource, respects PodDisruptionBudgets, and can be refused. - **Taint-manager eviction** — a plain pod deletion by the controller. It is not budget-aware. A node going NotReady is an involuntary disruption; the pods are already gone in practice, so blocking their removal would only leave the API object stranded. Do not rely on a PodDisruptionBudget to protect a workload from taint-based eviction. ## StatefulSets and the stuck-terminating case With an unreachable node the kubelet cannot confirm the pod is gone, so the pod object may sit in `Terminating` indefinitely. For a Deployment this is cosmetic, but a StatefulSet will not create the replacement `web-0` while the old `web-0` might still be running and holding a volume — the at-most-one guarantee demands it. Recovery is either deleting the Node object (which lets the API server force-delete its pods) or an explicit force delete. This is why NoExecute timing discussions almost always turn into a StatefulSet discussion in interviews. ## Practical uses Operators taint a node `maintenance=true:NoExecute` to push workloads off before a kernel patch; a hardware-health agent taints a node with failing disks; a cost controller taints spot/preemptible nodes on a termination notice with a short `tolerationSeconds` so pods have a bounded window to drain. In all of these, `tolerationSeconds` on the pods that *do* tolerate the taint is how you express "you may stay a bit longer, but not forever".

  • A node goes unreachable and its pods are deleted about five minutes later. Where does that five minutes come from, and how would you make one particular Deployment fail over faster?
    It comes from the `tolerationSeconds: 300` toleration for `node.kubernetes.io/unreachable` that the DefaultTolerationSeconds admission plugin injects into pods that do not declare their own. On top of it the node controller first waits its node-monitor grace period before marking the node unreachable at all. To speed up one Deployment, declare explicit tolerations for the `not-ready` and `unreachable` keys with a smaller `tolerationSeconds`, such as 30, in that pod template — this keeps the cluster-wide default intact for everything else.
  • Will a PodDisruptionBudget stop taint-based eviction?
    No. Taint-manager eviction is a direct pod deletion by the controller, not an API-initiated eviction through the Eviction subresource, so PodDisruptionBudgets are not consulted. Budgets protect against voluntary disruptions such as `kubectl drain` or a descheduler. Node failure is an involuntary disruption and honouring a budget there would only keep dead pod objects alive.
  • Why might a StatefulSet pod not be recreated after its node becomes unreachable?
    The pod object stays in Terminating because the kubelet is unreachable and cannot confirm the container is gone. A StatefulSet guarantees at most one pod per ordinal, so it refuses to create the replacement while the old one might still be running and attached to its volume. You resolve it by deleting the Node object, which lets the API server clean up its pods, or by force-deleting the pod once you are certain the machine is really down.

A plain toleration is a permanent residence permit; tolerationSeconds is a visa with an expiry date that starts ticking the day the eviction notice goes up on the building.

saying these in an interview costs you the question

  • Thinking tolerationSeconds works with NoSchedule or gates initial scheduling
  • Believing the timer starts when the pod is created rather than when the taint is applied
  • Claiming PodDisruptionBudgets protect pods from taint-based eviction
  • Assuming the ~5 minute failover delay is hardcoded in the scheduler rather than an injected default toleration
  • Setting the cluster-wide defaults to a few seconds without considering eviction storms during transient network issues

context