A Kubernetes DaemonSet that ships node logs runs on all ordinary worker nodes, but no pod ever appears on the control-plane nodes or on a pool that was marked with a custom taint. Why, and what do you change?
answer
- taint repels, toleration opts in, affinity attracts
- control-plane NoSchedule taint blocks agents
- `- operator: Exists` = tolerate every taint
- controller auto-adds not-ready/unreachable/pressure tolerations
- silent gap: DESIRED < node count
basics
~20 sThose nodes carry taints, and your pod template has no matching toleration, so pods are repelled. Add tolerations for each taint — for example the control-plane NoSchedule taint — or a blanket toleration with operator Exists to accept every taint on any node.
solid answer
~50 sTaints repel pods unless the pod tolerates them. Control-plane nodes carry `node-role.kubernetes.io/control-plane:NoSchedule`, and custom pools are usually tainted so only opted-in workloads land there. Your agent tolerates neither, so it is excluded. The fix is a toleration per taint you want to accept: ```yaml tolerations: - key: node-role.kubernetes.io/control-plane operator: Exists effect: NoSchedule ``` For a genuinely fleet-wide agent, the usual choice is `- operator: Exists` with no key and no effect, which tolerates **every** taint, including future ones. That is right for logging and security agents and wrong for anything you would not want on a dedicated node. Also know what the controller adds for you: DaemonSet pods automatically get tolerations for `not-ready`, `unreachable`, `memory-pressure`, `disk-pressure`, `pid-pressure`, `unschedulable` and, for host-network pods, `network-unavailable`. That is deliberate — agents must stay on exactly the nodes that are unhealthy.
code
yaml · 8 linesspec:
template:
spec:
tolerations:
- operator: Exists # any key, any value, any effect
containers:
- name: agent
image: registry.example.com/log-agent:1.9.2go deeper
Know that a taint blocks pods and a toleration is the opt-in, and that control-plane nodes are tainted by default.
Write the correct toleration shapes (Exists vs Equal, per-effect vs blanket) and know that scope and tolerations are two separate conditions.
Treat coverage gaps as a monitored invariant, use blanket tolerations only for genuine fleet-wide agents, and diagnose Pending pods from describe output rather than guesswork.
Own the taint taxonomy for the fleet — which pools are isolated and why, which agents are permitted to bypass isolation, and how coverage is proven for audit and incident response.
## Taints and tolerations in one paragraph A **taint** on a node says "repel pods that do not explicitly accept me". It is a key/value plus an effect: `NoSchedule` (do not place new pods here), `PreferNoSchedule` (avoid if possible), `NoExecute` (do not place, and evict pods already here that do not tolerate it). A **toleration** in a pod spec is the matching opt-in. Tolerating a taint does not attract a pod to that node — it only removes the barrier. Attraction is the job of `nodeSelector`/affinity. Candidates who blur these two produce agents that either miss nodes or land everywhere. ## Why the agent is missing Control-plane nodes are tainted by kubeadm and by every managed distribution with `node-role.kubernetes.io/control-plane:NoSchedule` (older clusters used `node-role.kubernetes.io/master`, and some clusters still carry both). Dedicated pools — GPU, spot, tenant-isolated, Windows — are typically tainted so only opted-in workloads run there. Your log agent's template carries no tolerations, so the scheduler rejects those nodes and the DaemonSet quietly covers a subset of the fleet. This failure mode is dangerous precisely because it is silent: `kubectl get ds` shows DESIRED equal to CURRENT equal to READY, all green, just with a smaller DESIRED than the node count. Blind spots in log or security coverage are found during an incident, not by an alert. Comparing `kubectl get nodes --no-headers | wc -l` against the DaemonSet's DESIRED is a cheap audit worth automating. ## Writing the tolerations Three shapes matter: 1. **Specific taint, specific effect** — `key: node-role.kubernetes.io/control-plane`, `operator: Exists`, `effect: NoSchedule`. Precise, self-documenting, and the right choice when you want the agent on control-plane nodes only. 2. **Key with a value** — `operator: Equal` plus `value:`. Use for pool taints such as `dedicated=tenant-a:NoSchedule` where different values mean different pools. 3. **Tolerate everything** — a single entry `- operator: Exists` with no key and no effect. This matches all taints, present and future, including `NoExecute`. It is the standard configuration for cluster-wide logging, metrics, CNI and security agents, because coverage must not depend on someone remembering to update the DaemonSet when a new tainted pool is created. It is also why you should never copy that snippet into an application workload: it will happily land on nodes that were isolated for a reason. For `NoExecute` taints you can also set `tolerationSeconds`, which defines how long a running pod may stay after the taint appears. Omitting it means "stay indefinitely", which is what agents want. ## What Kubernetes adds automatically The DaemonSet controller injects tolerations into every DaemonSet pod, and the list encodes an important design intent: agents must survive exactly the conditions that evict ordinary pods. - `node.kubernetes.io/not-ready:NoExecute` and `node.kubernetes.io/unreachable:NoExecute` — regular pods are evicted after ~5 minutes when a node goes bad; the agent stays, so you still get telemetry from a failing node. - `node.kubernetes.io/memory-pressure`, `disk-pressure`, `pid-pressure` (`NoSchedule`) — the agent is not blocked from a node under pressure, which is when you most need it. - `node.kubernetes.io/unschedulable:NoSchedule` — cordoning a node does not remove its agents. - `node.kubernetes.io/network-unavailable:NoExecute` — added only for host-network pods, so CNI installers can run before networking is ready. These are added regardless of what you write, and they explain the bootstrap ordering of a cluster: a CNI DaemonSet must be able to run on a node that is `NotReady` precisely because the node is `NotReady` until CNI is installed. ## Diagnosing in practice When an expected node has no agent, check taints first: `kubectl describe node <n> | grep -A5 Taints`, then compare with the template's tolerations. If the pod exists but is `Pending`, `kubectl describe pod` states the reason verbatim ("node(s) had untolerated taint {dedicated: gpu}"). If the pod does not exist at all, the node fell out of the *scope* (selector/affinity) rather than being repelled by a taint — a different fix. Keeping those two causes distinct is the whole diagnostic skill here.
- Which tolerations does Kubernetes add to DaemonSet pods automatically, and why those?The controller injects tolerations for node.kubernetes.io/not-ready and unreachable with the NoExecute effect, for memory-pressure, disk-pressure, pid-pressure and unschedulable with NoSchedule, and network-unavailable for host-network pods. The intent is that node agents must keep running on precisely the nodes that are unhealthy, cordoned or still bootstrapping — otherwise you would lose logs and metrics exactly when a node is failing, and a CNI installer could never run on a NotReady node.
- If a DaemonSet pod tolerates every taint, will it now run on every node in the cluster?Not necessarily. Tolerations only remove a barrier; they do not attract. The node must still be inside the DaemonSet's scope — matching any nodeSelector or required node affinity — and must have enough allocatable capacity for the pod's requests, otherwise the pod stays Pending. Tolerating everything also means the agent will land on isolated pools, which is intended for logging or security agents and wrong for ordinary workloads.
A taint is a bouncer at a door; a toleration is being on the guest list. Being on the list does not make you walk in — that's what the invitation (node affinity) is for.
saying these in an interview costs you the question
- Saying a toleration attracts pods to a node, confusing it with node affinity.
- Adding a toleration but expecting it to override a nodeSelector that excludes the node.
- Not knowing control-plane nodes are tainted by default and assuming a cluster-wide DaemonSet covers them.
- Copying `- operator: Exists` into ordinary application workloads, defeating dedicated-node isolation.
- Believing DaemonSet pods are evicted like other pods when a node goes NotReady — the controller's NoExecute tolerations keep them.