skip to content

Pods and Workloads

Pods are what Kubernetes schedules, and a controller decides how many run and for how long: stateless replicas, stable identities, per-node agents, batch runs. The wrong controller only hurts during a deploy or an outage, which is why interviewers probe it.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

In Kubernetes, what does a DaemonSet guarantee that a Deployment with a fixed replica count does not, and what kind of workload is it the right controller for?

level: juniorimportance: must knowfreq 72%

answer

  1. one pod per node, no replicas field
  2. count = f(node set), autoscaler-aware
  3. agents: logs, metrics, CNI, CSI node
  4. drain needs --ignore-daemonsets
  5. default scheduler binds since 1.12

basics

~20 s

A DaemonSet keeps one pod on every node matching its scope. It has no replica count: a new node automatically gets a pod, a deleted node loses its pod. Use it for per-node agents such as log shippers, metrics collectors, CNI and CSI plugins.

solid answer

~50 s

A **Deployment** says "keep N replicas running somewhere"; a **DaemonSet** says "keep exactly one pod per eligible node". There is no `replicas` field — the desired count is derived from the node set. When the cluster autoscaler adds a node, the DaemonSet controller creates a pod there within seconds; when the node object is deleted, its pod goes with it. Eligibility is every schedulable node by default, narrowed with `nodeSelector` or node affinity and widened onto tainted nodes with tolerations. The natural users are infrastructure agents that observe or serve the node itself: log shippers, metrics agents like node-exporter, CNI network plugins, CSI node drivers, security agents. They usually mount host paths or use `hostNetwork`. Anything that should scale with *traffic* rather than with *cluster size* belongs in a Deployment — a DaemonSet hard-couples pod count to node count, which is almost never what an application wants.

code

yaml · 25 lines
yaml
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: node-log-agent
  namespace: observability
spec:
  selector:
    matchLabels:
      app: node-log-agent
  template:
    metadata:
      labels:
        app: node-log-agent
    spec:
      containers:
        - name: agent
          image: registry.example.com/log-agent:1.9.2
          volumeMounts:
            - name: varlog
              mountPath: /var/log
              readOnly: true
      volumes:
        - name: varlog
          hostPath:
            path: /var/log

go deeper

for a junior

Recall the invariant (one pod per matching node, no replica count) and name two agent use cases such as log shipping and node metrics.

for a middle

Add the mechanics: the controller reacts to node add/remove events, the default scheduler binds the pod, and scope is narrowed by nodeSelector/affinity and widened by tolerations.

for a senior

Bring operational consequences: drain with --ignore-daemonsets, agents that must survive cordoned or unhealthy nodes, hostPath/hostNetwork privilege review, and per-node resource overhead.

for a principal

Frame it as a platform contract: which agents every node is required to carry, who owns them, the fixed capacity tax they impose per node, and where the CSI/CNI controller-vs-node split belongs.

## The core semantic Kubernetes workload controllers differ mainly in how they answer "how many pods should exist, and where?". A **Deployment** answers with a number you pick (`replicas: 5`), placed wherever the scheduler finds room. A **DaemonSet** answers with a function of the cluster: one pod per node that matches its scope. You never set a count, and `kubectl scale` on a DaemonSet fails because there is no `spec.replicas` field to scale. The object itself looks familiar: `spec.selector` plus `spec.template` (a pod template), exactly like a Deployment. What differs is the controller loop. The DaemonSet controller lists nodes, computes the set that should carry a pod, and reconciles: create where missing, delete where the node no longer qualifies. ## What the controller actually does Since Kubernetes 1.12 the DaemonSet controller does **not** place pods itself by writing `nodeName`. It creates a normal pod with a node affinity term pinning `metadata.name` to the target node, and the **default scheduler** binds it. That matters in practice: DaemonSet pods respect priority and preemption, and they show `Pending` with normal scheduler events when a node has no room for them, instead of silently forcing their way on. Node churn is handled automatically: - **Node joins** (autoscaler scale-up, new machine pool) → a pod is created for it immediately, so agents come up alongside the first application pods. - **Node is cordoned** (`kubectl cordon`) → existing DaemonSet pods keep running; the controller tolerates the `node.kubernetes.io/unschedulable` taint by default so agents are not dropped from nodes an operator has marked off-limits. - **Node is drained** → `kubectl drain` refuses to proceed until you pass `--ignore-daemonsets`, precisely because evicting an agent is pointless: the controller would immediately recreate it. - **Node object is deleted** → its pods are garbage-collected with it. ## What it is for The test is: *does one copy of this need to exist per machine, because it consumes something machine-scoped?* Canonical cases: - **Log collection** — the agent reads `/var/log/containers` on the host, so it must be on every node. - **Node metrics** — node-exporter reads `/proc` and `/sys` of that host. - **Networking** — CNI plugins install binaries and program routes/iptables per node. - **Storage** — CSI *node* plugins mount volumes on the node (the CSI *controller* half is a Deployment; that split is a good detail to volunteer). - **Security/compliance agents** — kernel-level or file-integrity monitoring per host. These pods usually need elevated access: `hostPath` mounts, `hostNetwork: true`, `hostPID`, or specific capabilities. That is expected for infrastructure agents but makes DaemonSets a privileged surface worth reviewing carefully. ## What it is not for A frequent misuse is "I want a copy of my API near every node, so I'll use a DaemonSet". That makes your capacity a function of cluster size: scale the cluster for a batch job and you have just multiplied your database connections. Application replicas should be driven by load (a Deployment, optionally with an autoscaler). Node-locality for latency is a service-routing concern (topology-aware routing), not a reason to pin one replica per node. The other misuse is caching or state: a DaemonSet pod dies with its node and is not rescheduled elsewhere, so it offers no durability guarantees. Stable identity and per-replica storage are StatefulSet territory. ## Interview framing State the invariant in one line — "one pod per matching node, count derived from the node set" — then name two or three real agents, then mention the three levers that define "matching": `nodeSelector`/affinity to narrow, tolerations to reach tainted nodes, and the update strategy for rollouts. That shows you have operated one rather than just read the definition.

  • What happens to DaemonSet pods when you run kubectl drain on a node?
    The drain refuses to start unless you pass `--ignore-daemonsets`, and with that flag the DaemonSet pods are simply left running while other pods are evicted. The reason is that evicting them is futile: the DaemonSet controller would recreate a pod on that node immediately, since the node still exists and still matches. Agents therefore keep shipping logs and metrics during the drain, which is usually what you want.
  • How does a DaemonSet pod get onto a node — does it go through kube-scheduler?
    Yes, since Kubernetes 1.12. The controller creates a pod carrying a required node affinity on `metadata.name` for the target node, and the default scheduler binds it. Consequently DaemonSet pods obey priority, preemption and resource checks, and will sit in `Pending` with normal scheduling events if the node has no allocatable room. Before 1.12 the controller set `nodeName` directly and bypassed the scheduler.

A Deployment is like hiring N staff for a company and letting HQ decide which office they sit in. A DaemonSet is like the rule "every office must have exactly one fire warden" — open a new office and a warden appears; close it and the post disappears.

saying these in an interview costs you the question

  • Saying you set replicas on a DaemonSet to control how many nodes it covers — there is no replicas field.
  • Claiming a DaemonSet pod is rescheduled to another node if its node dies; it is deleted with the node, and coverage is restored only when a matching node exists.
  • Using a DaemonSet to scale a stateless application 'so there is one near every node', coupling application capacity to cluster size.
  • Assuming DaemonSet pods bypass the scheduler and therefore always fit on a node — they can be Pending like any other pod.
  • Being surprised that kubectl drain aborts, instead of knowing --ignore-daemonsets is expected.

context

open as a page

In Kubernetes, describe what a Deployment, a ReplicaSet and a Pod each own, and walk through what the cluster actually does when you change the container image in a Deployment's pod template.

level: juniorimportance: must knowfreq 80%

basics

~20 s

A Deployment manages ReplicaSets; a ReplicaSet manages Pods. Changing the pod template creates a new ReplicaSet, which is scaled up while the old one is scaled down. Old ReplicaSets stay at zero replicas so you can roll back.

open as a page

How do you watch, pause and roll back an in-flight Kubernetes Deployment rollout from the command line, and what limits how far back a rollback can go?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Use kubectl rollout status to watch, rollout pause/resume to hold it, rollout history to list revisions and rollout undo (optionally --to-revision) to go back. How far back you can go is bounded by revisionHistoryLimit, default 10.

open as a page

In a Kubernetes Pod spec, what does the `initContainers` list do, and what happens if one of those containers exits with a non-zero status?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Init containers run to completion one at a time, in list order, before any app container starts. A non-zero exit means the kubelet retries that init container, or fails the Pod if restartPolicy is Never. Later containers never start.

open as a page

What is a Kubernetes Job, how does it differ from a Deployment, and which values may the `restartPolicy` field of a Job's Pod template take?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A Job runs Pods until a set number of them terminate successfully, then stops. A Deployment keeps a set of Pods running forever, restarting them whenever they exit. A Job's Pod template must use restartPolicy Never or OnFailure — Always is rejected.

open as a page

In Kubernetes, what exactly is a Pod, what do the containers inside one share, and what is the 'pause' (sandbox) container for?

level: juniorimportance: must knowfreq 85%

basics

~20 s

A Pod is Kubernetes' smallest schedulable unit: one or more containers placed on the same node, sharing one network namespace (one Pod IP, one port space, localhost), IPC and mounted volumes. A tiny 'pause' container holds those namespaces open.

open as a page

In a Kubernetes Pod spec, what is the difference between livenessProbe and readinessProbe, and what action does the platform take when each one fails?

level: juniorimportance: must knowfreq 85%

basics

~20 s

A failing livenessProbe makes the kubelet restart that container. A failing readinessProbe leaves the container running but removes the Pod from Service endpoints so it stops receiving traffic. Liveness asks 'is it broken beyond recovery?', readiness asks 'can it serve right now?'.

open as a page

In Kubernetes, how do you build a blue-green deployment from two Deployments and one Service, and how do cutover and rollback work?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Run blue and green as two Deployments whose pods differ by one label, such as slot, and point one Service at blue. Cutover patches the Service selector to green; rollback patches it back while blue is still running.

open as a page

In a Kubernetes Pod specification, what is the difference between a container's resource requests and its resource limits, and what does each one actually control?

level: juniorimportance: must knowfreq 80%

basics

~20 s

A request is what the scheduler reserves to place the Pod on a node. A limit is the hard ceiling the node enforces at runtime: exceed a CPU limit and the container is throttled, exceed a memory limit and it is OOMKilled.

open as a page

In a Kubernetes Deployment, which changes start a new rollout, and how does `kubectl rollout restart` trigger one?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Only a change to the Deployment's pod template starts a rollout; scaling, strategy edits and edits inside a referenced ConfigMap do not. kubectl rollout restart writes a timestamp annotation into the pod template, which causes a normal rolling update.

open as a page

When would you choose a Kubernetes StatefulSet over a Deployment, and what guarantees does the StatefulSet controller add?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Use a StatefulSet when replicas are not interchangeable. It gives each Pod a stable ordinal name (db-0, db-1) that survives rescheduling, a stable per-Pod DNS identity, and ordered, one-at-a-time creation, scaling and updates. A Deployment gives none of that.

open as a page

A Kubernetes DaemonSet that ships node logs runs on all ordinary worker nodes, but no pod ever appears on the control-plane nodes or on a pool that was marked with a custom taint. Why, and what do you change?

level: middleimportance: must knowfreq 52%

basics

~20 s

Those nodes carry taints, and your pod template has no matching toleration, so pods are repelled. Add tolerations for each taint — for example the control-plane NoSchedule taint — or a blanket toleration with operator Exists to accept every taint on any node.

open as a page

Explain what maxSurge and maxUnavailable do in a Kubernetes Deployment's RollingUpdate strategy, what their default values are, and when you would choose the Recreate strategy instead.

level: middleimportance: must knowfreq 74%

basics

~20 s

maxSurge is how many pods may exist above the desired count during a rollout; maxUnavailable is how many of the desired pods may be unavailable. Both default to 25%. Recreate deletes all old pods before creating new ones, accepting downtime.

open as a page

Kubernetes allows `restartPolicy: Always` on an entry inside a Pod's `initContainers` list. What does that setting change about how the container behaves, and which problems of the older "just add a second container to the Pod" sidecar pattern does it solve?

level: middleimportance: must knowfreq 55%

basics

~20 s

It makes that entry a native sidecar: it starts in init order but keeps running for the Pod's whole life instead of running to completion. Later init containers and the app containers wait for it, it does not block Job completion, and it shuts down after the app containers.

open as a page

In a Kubernetes Job spec, what do the `backoffLimit`, `activeDeadlineSeconds` and `ttlSecondsAfterFinished` fields control, and what happens when each of them is reached?

level: middleimportance: must knowfreq 50%

basics

~20 s

backoffLimit (default 6) caps retries; exceeding it fails the Job with reason BackoffLimitExceeded. activeDeadlineSeconds caps wall-clock runtime; exceeding it kills running Pods and fails the Job with DeadlineExceeded. ttlSecondsAfterFinished deletes the finished Job and its Pods after that many seconds.

open as a page

In a Kubernetes Job spec, how do the `completions` and `parallelism` fields interact, and what changes when `completionMode` is set to `Indexed`?

level: middleimportance: must knowfreq 55%

basics

~20 s

completions is how many Pods must succeed; parallelism is how many may run at once. The Job keeps up to parallelism Pods running until completions successes accumulate. Indexed mode gives each Pod a fixed index 0..completions-1, so each does a distinct shard.

open as a page

How do the `schedule` and `concurrencyPolicy` fields of a Kubernetes CronJob work, and what do the values `Allow`, `Forbid` and `Replace` each do when the previous run is still going?

level: middleimportance: must knowfreq 55%

basics

~20 s

schedule is a five-field cron expression; at each firing the CronJob controller creates a Job. concurrencyPolicy decides what happens if the previous Job is still running: Allow (default) starts another anyway, Forbid skips the new run, Replace kills the running Job and starts the new one.

open as a page

Walk through the Kubernetes Pod phases (Pending, Running, Succeeded, Failed, Unknown) and explain what the Pod-level restartPolicy values Always, OnFailure and Never do when a container exits.

level: middleimportance: must knowfreq 70%

basics

~20 s

Pending = accepted but not all containers running (scheduling, image pull, init). Running = bound to a node with at least one container running. Succeeded/Failed = all containers terminated, all zero exit codes or not. Unknown = node unreachable. restartPolicy tells the kubelet whether to restart exited containers in place.

open as a page

What problem does the Kubernetes startupProbe solve, and how does it interact with the liveness and readiness probes while it is still running?

level: middleimportance: must knowfreq 60%

basics

~20 s

It protects slow-starting containers: while a startupProbe is defined and has not yet succeeded, the kubelet suspends the liveness and readiness probes and the container is not ready. Once it succeeds it never runs again. Its budget is failureThreshold x periodSeconds.

open as a page

Kubernetes assigns every Pod a Quality of Service class of Guaranteed, Burstable, or BestEffort. How is that class derived from the Pod spec, and where does it actually change behaviour?

level: middleimportance: must knowfreq 62%

basics

~20 s

Guaranteed: every container sets CPU and memory requests and limits, and request equals limit. BestEffort: no container sets any request or limit. Anything in between is Burstable. The class drives eviction order under node pressure and OOM kill priority.

open as a page

For a 7-replica Kubernetes Deployment with default 25% maxSurge and maxUnavailable, what pod-count bounds hold mid-rollout, and how does rounding behave at small replica counts?

level: middleimportance: must knowfreq 60%

basics

~20 s

maxSurge rounds up and maxUnavailable rounds down. With 7 replicas, 25% becomes 2 surge and 1 unavailable: at most 9 pods and at least 6 available. At 1 to 3 replicas the defaults become surge 1, unavailable 0.

open as a page

How does an individual Pod in a Kubernetes StatefulSet get a stable DNS name, and what role does a headless Service play in that?

level: middleimportance: must knowfreq 55%

basics

~10 s

The StatefulSet names a governing Service in spec.serviceName. That Service is headless (clusterIP: None), so DNS returns Pod addresses rather than a virtual IP, and each Pod resolves at <pod-name>.<service>.<namespace>.svc.cluster.local.

open as a page

When a Kubernetes pod is deleted, why can it still receive new requests after SIGTERM, and how does a preStop sleep prevent the resulting errors?

level: middleimportance: must knowfreq 66%

basics

~10 s

Deletion starts two unordered paths: the kubelet stops the container while endpoint controllers, kube-proxy and ingress controllers withdraw the pod from routing. A preStop sleep delays SIGTERM until routing has caught up.

open as a page

Describe what happens between 'kubectl delete pod' and the container disappearing in Kubernetes: which signal is sent, what terminationGracePeriodSeconds controls, and why in-flight requests can still fail during that window.

level: seniorimportance: must knowfreq 65%

basics

~20 s

The API server sets a deletionTimestamp and a grace period (default 30s). In parallel the Pod is removed from Service endpoints, and the kubelet runs any preStop hook then sends SIGTERM to PID 1 of each container; anything alive when the grace period ends gets SIGKILL. Requests still fail because endpoint removal propagates asynchronously.

open as a page

A service's Kubernetes livenessProbe calls an endpoint that verifies the database connection. Why is that dangerous, and what would you do instead?

level: seniorimportance: must knowfreq 55%

basics

~20 s

A shared dependency failing makes every replica fail liveness simultaneously, so the kubelet restarts the whole fleet - which does not fix the database, destroys warm state, and creates a thundering herd on recovery. Liveness should test only local, unrecoverable failure; dependency checks belong in readiness, selectively.

open as a page

How do you choose periodSeconds, timeoutSeconds, failureThreshold and initialDelaySeconds for a Kubernetes probe? Work through what those numbers mean for detection time and for false restarts under load.

level: seniorimportance: must knowfreq 55%

basics

~20 s

Worst-case reaction time is roughly initialDelaySeconds + periodSeconds x failureThreshold, plus timeoutSeconds. Defaults are period 10s, timeout 1s, failureThreshold 3. Keep liveness generous and cheap so latency spikes do not restart healthy Pods; keep readiness tighter, since removing traffic is reversible.

open as a page

A container exceeds its CPU limit, and another container exceeds its memory limit. Describe what the Linux kernel does in each case, and how you would detect each from outside the container.

level: seniorimportance: must knowfreq 55%

basics

~20 s

CPU is compressible: the kernel's CFS bandwidth controller stalls the container's threads until the next 100 ms period, visible as throttled-seconds metrics and latency spikes. Memory is not: the cgroup OOM killer terminates the process, exit code 137, reason OOMKilled.

open as a page

What does `kubectl delete pod --grace-period=0 --force` actually do in Kubernetes, and when is it safe to use?

level: juniorimportance: should knowfreq 46%

basics

~20 s

Force deletion removes the Pod object from the API server at once, without waiting for the kubelet to confirm the containers stopped. The processes may keep running on the node, so use it only when the pod is known to be dead.

open as a page

You need a Kubernetes DaemonSet's pods to run only on the subset of nodes that carry GPUs rather than on every node in the cluster. How do you scope it, and what happens if a node's labels change afterwards?

level: middleimportance: should knowfreq 50%

basics

~20 s

Put a nodeSelector, or a requiredDuringSchedulingIgnoredDuringExecution node affinity, in the DaemonSet's pod template. The controller then creates pods only on matching nodes. If a node stops matching, the DaemonSet controller deletes its pod; if a node starts matching, it gets one.

open as a page

What are the Kubernetes container lifecycle hooks postStart and preStop, when does the kubelet run each, and what guarantees do they give?

level: middleimportance: should knowfreq 40%

basics

~20 s

They are per-container callbacks run by the kubelet: postStart fires right after container creation, concurrently with the entrypoint, and blocks the container from being Running until it returns; preStop fires just before SIGTERM on shutdown. Handlers are exec, httpGet or sleep, delivered at least once.

open as a page

showing 1–30 of 49