You need a Kubernetes DaemonSet's pods to run only on the subset of nodes that carry GPUs rather than on every node in the cluster. How do you scope it, and what happens if a node's labels change afterwards?
answer
- scope lives in the pod template
- nodeSelector simple, nodeAffinity expressive
- label removed → controller deletes the pod
- DESIRED 0 = selector matches nothing
- don't set nodeName; arch/os labels on mixed fleets
basics
~20 sPut a nodeSelector, or a requiredDuringSchedulingIgnoredDuringExecution node affinity, in the DaemonSet's pod template. The controller then creates pods only on matching nodes. If a node stops matching, the DaemonSet controller deletes its pod; if a node starts matching, it gets one.
solid answer
~50 sScope lives in the **pod template**, not on the DaemonSet itself. The simple form is `spec.template.spec.nodeSelector: { gpu: "true" }`. The expressive form is `nodeAffinity.requiredDuringSchedulingIgnoredDuringExecution`, which supports `In`, `NotIn`, `Exists` and multiple terms — useful for "nvidia or amd" or "any pool except spot". The key difference from an ordinary pod is what happens on label change. For a bare pod, `IgnoredDuringExecution` means a running pod is never disturbed. For a DaemonSet, the *controller* re-evaluates continuously: label a node and it gains a pod; remove the label and the controller **deletes** the pod on that node. So node labels behave like live membership, not a one-time placement decision. In production the label usually comes from the machine pool (`cloud.google.com/gke-nodepool`, `eks.amazonaws.com/nodegroup`) or from node-feature discovery, so scoping is declarative rather than hand-labelled. Watch out for typos: a selector that matches nothing yields DESIRED 0 and a healthy-looking but useless DaemonSet.
code
yaml · 14 linesspec:
template:
spec:
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: accelerator
operator: In
values: ["nvidia", "amd"]
- key: kubernetes.io/os
operator: In
values: ["linux"]go deeper
Know that nodeSelector in the pod template limits which nodes get a pod, and that adding the label to a node makes an agent appear there.
Explain nodeSelector vs required node affinity, and the fact that the controller deletes pods on nodes that stop matching.
Use scoping deliberately: pool-based labels rather than manual ones, label-driven canaries for agent upgrades, and a diagnosis loop that separates selector, taint and capacity causes.
Decide the fleet's labelling contract — which labels are machine-pool-owned versus operator-applied, and how agent coverage is guaranteed and audited as pools are added.
## Where scoping is expressed A DaemonSet's "eligible node set" is computed from its **pod template**. There is no field on the DaemonSet spec itself for node targeting; you use the same primitives any pod uses: - **`nodeSelector`** — a flat map of label key/value pairs that must all match. Simple, readable, covers most cases. - **`affinity.nodeAffinity.requiredDuringSchedulingIgnoredDuringExecution`** — a list of node selector terms (OR'ed) each holding match expressions (AND'ed), with operators `In`, `NotIn`, `Exists`, `DoesNotExist`, `Gt`, `Lt`. This is how you express "GPU nodes of either vendor" or "all nodes except the spot pool". What you should *not* do is set `nodeName` in the template. The controller injects its own required affinity on `metadata.name` to pin each generated pod to its node, and a hand-written `nodeName` fights that mechanism. `preferredDuringSchedulingIgnoredDuringExecution` (soft affinity) is nearly meaningless here: the target node is already fixed for each generated pod, so there is nothing to prefer between. ## The controller re-evaluates, continuously This is the part interviewers probe. For a normal pod, `IgnoredDuringExecution` promises that once scheduled, later label changes never move it. That promise still holds for the *pod*, but the DaemonSet **controller** sits above it and reconciles membership on every sync: - Node gains the matching label → controller creates a pod for it. - Node loses the matching label → controller **deletes** the pod on that node. So removing a label is a live, destructive operation on running agents. `kubectl label node worker-3 gpu-` will terminate the GPU agent there within seconds. Treat node labels that DaemonSets key on as production configuration, and prefer labels applied by the machine pool or by node-feature discovery over ad-hoc human labelling. The same logic explains a common confusion: a DaemonSet's `DESIRED` column is not "number of nodes" but "number of nodes matching the scope, after tolerations are considered". A selector with a typo produces `DESIRED 0` — no pods, no errors, no events. Always verify with `kubectl get ds` and compare against `kubectl get nodes -l <your selector>`. ## Scoping patterns that come up 1. **Hardware-specific agents.** GPU device plugins and monitoring agents (DCGM exporter) belong only on GPU nodes; deploying them cluster-wide wastes memory on every CPU node and produces crash loops where the driver is absent. 2. **Pool-specific behaviour.** Different log pipelines for the `system` and `tenant` pools, expressed by selecting on the cloud provider's node-pool label. 3. **Progressive rollout by label.** Ship a risky agent version scoped to `agent-canary: "true"`, label three nodes, watch, then widen. This is the closest thing DaemonSets have to a canary, since the update strategy has no traffic-based gating. 4. **OS or architecture splits.** `kubernetes.io/os: linux` is worth adding explicitly on mixed Windows clusters; `kubernetes.io/arch` matters on mixed arm64/amd64 fleets where the image is not multi-arch. Forgetting these produces pods that pull an incompatible image and crash-loop on a subset of nodes. ## Interaction with taints Scoping narrows; tolerations widen. A node that matches your selector but carries a taint your pod does not tolerate will still not run the agent — the pod is created and then sits `Pending` with a scheduling error, or is never created if the controller filters it out. When an expected node has no agent, check both directions: does the node match the selector, and does the pod tolerate every taint on that node? ## Verifying The fast loop is: `kubectl get nodes -l gpu=true` (how many nodes *should* match), `kubectl get ds -n <ns>` (what the controller thinks DESIRED/CURRENT/READY are), then `kubectl describe pod` on any pod stuck `Pending` to see the scheduler's own explanation. Those three commands separate "my selector is wrong" from "my toleration is missing" from "the node is genuinely full".
- A DaemonSet pod is already running on a node and you remove the label its nodeSelector matches. Does the pod keep running, given that node affinity is IgnoredDuringExecution?No. IgnoredDuringExecution applies to the scheduler's treatment of an already-bound pod, but the DaemonSet controller independently recomputes which nodes should carry a pod on every sync. Once the node stops matching, the controller deletes that pod. Node labels used by DaemonSets are therefore live membership switches, not one-time placement hints.
- Why is preferred (soft) node affinity of little use inside a DaemonSet template?Because the controller already generates one pod per target node and pins it there with a required affinity on the node's name. There is no set of candidate nodes left for a soft preference to rank, so a preferred term has essentially nothing to influence. Narrowing must be done with required affinity or nodeSelector.
saying these in an interview costs you the question
- Believing IgnoredDuringExecution protects a running DaemonSet pod when the node label is removed — the controller deletes it.
- Setting nodeName in the pod template instead of using a selector or affinity.
- Reading DESIRED 0 as a bug in Kubernetes rather than a selector that matches no node.
- Assuming a matching label is sufficient, and forgetting that a taint on the node still blocks the pod.
- Omitting kubernetes.io/os or kubernetes.io/arch on mixed clusters and then blaming the image for crash loops.