skip to content

In Kubernetes, what does a DaemonSet guarantee that a Deployment with a fixed replica count does not, and what kind of workload is it the right controller for?

level: juniorimportance: must knowfreq 72%

answer

  1. one pod per node, no replicas field
  2. count = f(node set), autoscaler-aware
  3. agents: logs, metrics, CNI, CSI node
  4. drain needs --ignore-daemonsets
  5. default scheduler binds since 1.12

basics

~20 s

A DaemonSet keeps one pod on every node matching its scope. It has no replica count: a new node automatically gets a pod, a deleted node loses its pod. Use it for per-node agents such as log shippers, metrics collectors, CNI and CSI plugins.

solid answer

~50 s

A **Deployment** says "keep N replicas running somewhere"; a **DaemonSet** says "keep exactly one pod per eligible node". There is no `replicas` field — the desired count is derived from the node set. When the cluster autoscaler adds a node, the DaemonSet controller creates a pod there within seconds; when the node object is deleted, its pod goes with it. Eligibility is every schedulable node by default, narrowed with `nodeSelector` or node affinity and widened onto tainted nodes with tolerations. The natural users are infrastructure agents that observe or serve the node itself: log shippers, metrics agents like node-exporter, CNI network plugins, CSI node drivers, security agents. They usually mount host paths or use `hostNetwork`. Anything that should scale with *traffic* rather than with *cluster size* belongs in a Deployment — a DaemonSet hard-couples pod count to node count, which is almost never what an application wants.

code

yaml · 25 lines
yaml
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: node-log-agent
  namespace: observability
spec:
  selector:
    matchLabels:
      app: node-log-agent
  template:
    metadata:
      labels:
        app: node-log-agent
    spec:
      containers:
        - name: agent
          image: registry.example.com/log-agent:1.9.2
          volumeMounts:
            - name: varlog
              mountPath: /var/log
              readOnly: true
      volumes:
        - name: varlog
          hostPath:
            path: /var/log

go deeper

for a junior

Recall the invariant (one pod per matching node, no replica count) and name two agent use cases such as log shipping and node metrics.

for a middle

Add the mechanics: the controller reacts to node add/remove events, the default scheduler binds the pod, and scope is narrowed by nodeSelector/affinity and widened by tolerations.

for a senior

Bring operational consequences: drain with --ignore-daemonsets, agents that must survive cordoned or unhealthy nodes, hostPath/hostNetwork privilege review, and per-node resource overhead.

for a principal

Frame it as a platform contract: which agents every node is required to carry, who owns them, the fixed capacity tax they impose per node, and where the CSI/CNI controller-vs-node split belongs.

## The core semantic Kubernetes workload controllers differ mainly in how they answer "how many pods should exist, and where?". A **Deployment** answers with a number you pick (`replicas: 5`), placed wherever the scheduler finds room. A **DaemonSet** answers with a function of the cluster: one pod per node that matches its scope. You never set a count, and `kubectl scale` on a DaemonSet fails because there is no `spec.replicas` field to scale. The object itself looks familiar: `spec.selector` plus `spec.template` (a pod template), exactly like a Deployment. What differs is the controller loop. The DaemonSet controller lists nodes, computes the set that should carry a pod, and reconciles: create where missing, delete where the node no longer qualifies. ## What the controller actually does Since Kubernetes 1.12 the DaemonSet controller does **not** place pods itself by writing `nodeName`. It creates a normal pod with a node affinity term pinning `metadata.name` to the target node, and the **default scheduler** binds it. That matters in practice: DaemonSet pods respect priority and preemption, and they show `Pending` with normal scheduler events when a node has no room for them, instead of silently forcing their way on. Node churn is handled automatically: - **Node joins** (autoscaler scale-up, new machine pool) → a pod is created for it immediately, so agents come up alongside the first application pods. - **Node is cordoned** (`kubectl cordon`) → existing DaemonSet pods keep running; the controller tolerates the `node.kubernetes.io/unschedulable` taint by default so agents are not dropped from nodes an operator has marked off-limits. - **Node is drained** → `kubectl drain` refuses to proceed until you pass `--ignore-daemonsets`, precisely because evicting an agent is pointless: the controller would immediately recreate it. - **Node object is deleted** → its pods are garbage-collected with it. ## What it is for The test is: *does one copy of this need to exist per machine, because it consumes something machine-scoped?* Canonical cases: - **Log collection** — the agent reads `/var/log/containers` on the host, so it must be on every node. - **Node metrics** — node-exporter reads `/proc` and `/sys` of that host. - **Networking** — CNI plugins install binaries and program routes/iptables per node. - **Storage** — CSI *node* plugins mount volumes on the node (the CSI *controller* half is a Deployment; that split is a good detail to volunteer). - **Security/compliance agents** — kernel-level or file-integrity monitoring per host. These pods usually need elevated access: `hostPath` mounts, `hostNetwork: true`, `hostPID`, or specific capabilities. That is expected for infrastructure agents but makes DaemonSets a privileged surface worth reviewing carefully. ## What it is not for A frequent misuse is "I want a copy of my API near every node, so I'll use a DaemonSet". That makes your capacity a function of cluster size: scale the cluster for a batch job and you have just multiplied your database connections. Application replicas should be driven by load (a Deployment, optionally with an autoscaler). Node-locality for latency is a service-routing concern (topology-aware routing), not a reason to pin one replica per node. The other misuse is caching or state: a DaemonSet pod dies with its node and is not rescheduled elsewhere, so it offers no durability guarantees. Stable identity and per-replica storage are StatefulSet territory. ## Interview framing State the invariant in one line — "one pod per matching node, count derived from the node set" — then name two or three real agents, then mention the three levers that define "matching": `nodeSelector`/affinity to narrow, tolerations to reach tainted nodes, and the update strategy for rollouts. That shows you have operated one rather than just read the definition.

  • What happens to DaemonSet pods when you run kubectl drain on a node?
    The drain refuses to start unless you pass `--ignore-daemonsets`, and with that flag the DaemonSet pods are simply left running while other pods are evicted. The reason is that evicting them is futile: the DaemonSet controller would recreate a pod on that node immediately, since the node still exists and still matches. Agents therefore keep shipping logs and metrics during the drain, which is usually what you want.
  • How does a DaemonSet pod get onto a node — does it go through kube-scheduler?
    Yes, since Kubernetes 1.12. The controller creates a pod carrying a required node affinity on `metadata.name` for the target node, and the default scheduler binds it. Consequently DaemonSet pods obey priority, preemption and resource checks, and will sit in `Pending` with normal scheduling events if the node has no allocatable room. Before 1.12 the controller set `nodeName` directly and bypassed the scheduler.

A Deployment is like hiring N staff for a company and letting HQ decide which office they sit in. A DaemonSet is like the rule "every office must have exactly one fire warden" — open a new office and a warden appears; close it and the post disappears.

saying these in an interview costs you the question

  • Saying you set replicas on a DaemonSet to control how many nodes it covers — there is no replicas field.
  • Claiming a DaemonSet pod is rescheduled to another node if its node dies; it is deleted with the node, and coverage is restored only when a matching node exists.
  • Using a DaemonSet to scale a stateless application 'so there is one near every node', coupling application capacity to cluster size.
  • Assuming DaemonSet pods bypass the scheduler and therefore always fit on a node — they can be Pending like any other pod.
  • Being surprised that kubectl drain aborts, instead of knowing --ignore-daemonsets is expected.

context