How does a Kubernetes DaemonSet roll a new pod template out across a large fleet of nodes, and which settings control the speed and blast radius of that rollout?
answer
- RollingUpdate default, OnDelete manual
- maxUnavailable = blast-radius budget, default 1
- maxSurge 1.22+, needs maxUnavailable 0, no hostPort
- readiness + minReadySeconds gate progress
- rollout undo works; no auto-rollback
basics
~20 sWith updateStrategy RollingUpdate (the default), the controller replaces pods node by node, keeping at most maxUnavailable (default 1) unavailable at a time; maxSurge can instead start the new pod before deleting the old. OnDelete updates a node only when you delete its pod manually.
solid answer
~60 sA DaemonSet has two strategies. **RollingUpdate** (default) walks the nodes, deleting the old pod and letting the controller create the new one, never exceeding `maxUnavailable` unavailable pods at once — default `1`, which on a 2,000-node cluster means a very long rollout, so percentages such as `10%` are common. **OnDelete** changes nothing until you delete a pod yourself; the node picks up the new template only then. That is the escape hatch for agents whose restart is disruptive and must be coordinated with node maintenance. Since Kubernetes 1.22 `maxSurge` is also available: the new pod starts *alongside* the old one on the same node before the old is removed, which avoids an observability gap — but it requires `maxUnavailable: 0` and fails for agents using `hostPort` or otherwise unable to run two copies on a node. Progress is gated by readiness plus `minReadySeconds`. `kubectl rollout status/history/undo` work on DaemonSets, but there is no automatic rollback: a crash-looping agent stalls the rollout at `maxUnavailable` nodes, which is exactly the containment you want.
code
yaml · 22 linesapiVersion: apps/v1
kind: DaemonSet
metadata:
name: node-log-agent
spec:
minReadySeconds: 30
updateStrategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
selector:
matchLabels: { app: node-log-agent }
template:
metadata:
labels: { app: node-log-agent }
annotations:
checksum/config: "8f14e45fceea167a"
spec:
containers:
- name: agent
image: registry.example.com/log-agent:1.9.2go deeper
Know that changing the pod template triggers a node-by-node replacement and that maxUnavailable limits how many pods are down at once.
Explain both strategies, the defaults, how readiness and minReadySeconds gate progress, and that rollout status/history/undo apply to DaemonSets.
Size the blast-radius budget for the specific agent, use surge or OnDelete deliberately, force config rollouts with a template hash, and drive rollouts from a pipeline with timeouts.
Define the fleet upgrade policy: canary tiers, digest pinning for privileged agents, per-agent availability budgets, and the recovery time implied when undo is itself a slow rolling update.
## Two strategies `spec.updateStrategy.type` is either `RollingUpdate` (the default since 1.6) or `OnDelete`. **RollingUpdate** reconciles automatically. After you change the pod template — a new image tag, a changed env var, a new config hash annotation — the controller identifies pods whose template hash is stale and replaces them a few at a time. It does not batch by node pool or zone; ordering is essentially arbitrary, which is worth saying out loud because people often assume zone-aware sequencing. **OnDelete** makes the controller passive: it records the new template but leaves existing pods alone. A node gets the new version only when its pod is deleted, whether by you, by a node drain, or by node replacement. Teams use it for agents whose restart causes real disruption (a CNI plugin that briefly reprograms datapath rules, a storage driver mid-attach) so the upgrade rides along with planned node maintenance or an image-based node roll. ## The knobs - **`maxUnavailable`** (default `1`, absolute number or percentage) — the ceiling on how many DaemonSet pods may be unavailable simultaneously. `1` is safe and slow; on a 1,000-node cluster with a 30-second agent start, a serial rollout takes hours. `10%` is a common compromise; percentages round **down** but never to zero when non-zero was requested. - **`maxSurge`** (Kubernetes 1.22+, default `0`) — when non-zero, the controller creates the new pod on a node while the old one still runs, then removes the old one once the new is ready. This removes the observability blind spot during agent restarts. Constraints: `maxSurge` and `maxUnavailable` cannot both be non-zero, and two copies must actually be able to coexist on one node — an agent binding a `hostPort` or an exclusive host resource cannot surge. - **`minReadySeconds`** — a pod must stay ready this long before it counts as available and the rollout advances. It converts a fast crash-after-ready into a stalled rollout rather than a fleet-wide outage. Availability is judged by readiness, so an agent with no readiness probe is considered available the moment its container is running. For a rollout knob to mean anything, the agent needs a probe that reflects real function. ## Rollout containment The useful mental model: `maxUnavailable` is a blast-radius budget. A bad image cannot break more nodes than that budget at any instant, because the controller refuses to proceed while the budget is spent. A broken agent therefore shows up as a rollout stuck at N updated nodes, with the rest untouched — a much better failure than an all-at-once restart. DaemonSets keep revision history in `ControllerRevision` objects (bounded by `revisionHistoryLimit`, default 10), so `kubectl rollout history daemonset/x` and `kubectl rollout undo daemonset/x --to-revision=3` both work. Undo is itself a rolling update under the same constraints, so recovery from a bad agent takes as long as the rollout did — a good argument for a larger `maxUnavailable` on agents you trust and label-based canaries for those you do not. There is no automatic rollback for DaemonSets. Nothing watches error rates and reverts; `progressDeadlineSeconds` does not exist here as it does for Deployments. Detection is on your monitoring. ## Practical rollout discipline 1. **Canary by label.** Scope a copy of the DaemonSet to `agent-canary: "true"`, label a handful of nodes across zones and hardware types, bake, then roll the fleet-wide one. 2. **Pin digests.** Agents often run privileged with host mounts; `image: agent:latest` plus `imagePullPolicy: Always` means a node reboot silently upgrades one node. Use a tag plus digest so the version is a property of the manifest. 3. **Roll config too.** A DaemonSet reading a ConfigMap does not restart when the map changes. Put a hash of the config in a pod-template annotation so a config change becomes a template change and inherits the same controlled rollout. 4. **Size the budget honestly.** Ask what happens to the platform while `maxUnavailable` nodes have no agent: 10% of nodes briefly not shipping logs is usually fine; 10% of nodes with no CNI is not. 5. **Watch, do not tail.** `kubectl rollout status ds/<name>` blocks until done and is the right thing to put in a pipeline, with a timeout.
- When would you choose OnDelete over RollingUpdate for a DaemonSet?When restarting the agent is itself disruptive and must be coordinated with node maintenance — CNI plugins that reprogram the datapath, CSI node drivers that may be mid-attach, or kernel-adjacent security agents. With OnDelete the new template is stored but applied per node only when that node's pod is deleted, so the upgrade rides along with a planned drain or a node-image roll instead of sweeping the fleet on its own schedule.
- Your DaemonSet uses maxSurge: 1 and the rollout fails immediately with pods stuck Pending. What is the likely cause?Surge requires two copies of the agent to coexist briefly on the same node. If the pod binds a hostPort, claims an exclusive host resource such as a device, or the node simply lacks allocatable capacity for a second copy, the surged pod can never schedule. The fix is to fall back to maxUnavailable-based rolling, or to remove the exclusive host binding.
saying these in an interview costs you the question
- Thinking a DaemonSet rolls out to all nodes at once, or that maxUnavailable does not exist for DaemonSets.
- Setting both maxSurge and maxUnavailable to non-zero values, which the API rejects.
- Expecting an automatic rollback when the new agent crash-loops — the rollout just stalls.
- Leaving maxUnavailable at 1 on a multi-thousand-node cluster and being surprised the rollout takes many hours.
- Assuming a ConfigMap change restarts DaemonSet pods; without a template change nothing rolls.