skip to content

In Kubernetes, what do `kubectl cordon`, `kubectl drain` and `kubectl uncordon` each do when you take a node out of service?

level: juniorimportance: must knowfreq 72%

answer

  1. one field on the Node
  2. placement versus pods in place
  3. eviction, then controllers recreate
  4. SchedulingDisabled in get nodes
  5. no automatic rebalance afterwards

basics

~20 s

Cordon marks a node unschedulable so no new pods land there. Drain cordons it and then evicts its pods so their controllers recreate them elsewhere. Uncordon makes the node schedulable again without moving any pods back.

solid answer

~30 s

`kubectl cordon` sets `spec.unschedulable: true` on the Node, so the scheduler stops placing pods there, but running pods are untouched. `kubectl drain` does that cordon and then removes the node's pods through the Eviction API, so disruption budgets are respected and each pod shuts down gracefully; controllers such as ReplicaSets recreate the pods on other nodes. It skips mirror (static) pods and refuses to continue past DaemonSet pods, `emptyDir` pods or bare pods unless you pass the matching flag. `kubectl uncordon` clears the flag once maintenance is done. Nothing migrates back: the node only receives pods created from then on.

code

bash · 4 lines
bash
kubectl cordon node-a-17
kubectl get node node-a-17
kubectl drain node-a-17 --ignore-daemonsets --delete-emptydir-data --timeout=11m
kubectl uncordon node-a-17

go deeper

for a junior

Recall the three commands in order and the one-line effect of each: cordon blocks new pods, drain empties the node, uncordon reopens it.

for a middle

Explain that cordon is just spec.unschedulable plus a mirrored taint, that drain uses eviction and relies on controllers to recreate pods, and why DaemonSet pods stay.

for a senior

Show operational awareness: a refused or timed-out drain leaves the node cordoned, bare pods are lost for good, and uncordon never rebalances load.

for a principal

Frame cordon-drain-uncordon as the primitive every upgrade, image rotation and hardware process is built on, and argue for automating it with capacity checks and health gates.

## Three commands, three different jobs Taking a Kubernetes **node** out of service for maintenance (a kernel patch, a kubelet upgrade, a hardware swap) is a three-step procedure, and each step is a separate `kubectl` command with a separate effect. | Command | What it changes | Effect on pods already running | Effect on new pods | |---|---|---|---| | `kubectl cordon NODE` | sets `spec.unschedulable: true` on the Node object | none, they keep running | the scheduler stops placing them there | | `kubectl drain NODE` | cordons first, then removes the node's pods | evicted (or deleted) so their controllers recreate them elsewhere | none can land, because the node is cordoned | | `kubectl uncordon NODE` | sets `spec.unschedulable` back to `false` | none | the node is eligible again | ## Cordon: a fence, not an eviction **Cordoning** only flips one field on the Node object. `kubectl get nodes` then shows the node as `Ready,SchedulingDisabled`. Two things follow from that single field: - The scheduler's `NodeUnschedulable` filter rejects the node for any pod that does not tolerate the `node.kubernetes.io/unschedulable` taint. - The node lifecycle controller in kube-controller-manager mirrors the field as a **taint**, `node.kubernetes.io/unschedulable:NoSchedule`. Pods that tolerate that taint can still be scheduled there. The DaemonSet controller adds that toleration to every DaemonSet pod, which is why a cordoned node keeps its log shipper and CNI agent. Nothing that is already running is touched: `NoSchedule` affects placement only, never pods in place. Cordon on its own is useful when you want a node to "bleed out" naturally, for example while you investigate a flaky disk without disturbing the pods on it. ## Drain: cordon plus eviction **Draining** does the cordon for you and then empties the node: 1. It marks the node unschedulable, exactly like `kubectl cordon`. 2. It lists every pod on the node and checks them. DaemonSet-managed pods, pods with `emptyDir` volumes and pods with no owning controller make the drain stop with an error unless you pass the matching flag (`--ignore-daemonsets`, `--delete-emptydir-data`, `--force`). **Mirror pods** (the API view of static pods from `/etc/kubernetes/manifests`) are always skipped, because the API server cannot delete them. 3. For the remaining pods it calls the pod's **`eviction` subresource** (the Eviction API) rather than a plain `DELETE`, so disruption budgets are consulted. Each eviction starts ordinary graceful termination, using the pod's own `terminationGracePeriodSeconds` unless `--grace-period` overrides it. 4. It waits until every evicted pod is actually gone, then prints `node/<name> drained`. The replacements are created by the pods' controllers (ReplicaSet, StatefulSet, Job) and scheduled onto other nodes. Drain does not "move" a pod; it removes it and relies on a controller to recreate it. A pod with no controller simply ends. A detail that surprises people: if the pre-check fails, **the node stays cordoned** even though no pod was evicted, because the cordon happened in step 1. ## Uncordon: open the gate again `kubectl uncordon NODE` clears `spec.unschedulable`, and the node lifecycle controller removes the taint. **Nothing moves back.** The pods that were evicted are happily running elsewhere, and the scheduler only uses the node for pods created from now on. The node fills up again as rollouts, scale-ups and restarts happen; forcing a rebalance needs a separate tool (a descheduler) or a rollout restart. ## A typical session Picture a 64-node, two-zone managed cluster that runs a feature-flag evaluation service. One node needs an OS patch: ```bash kubectl cordon node-a-17 kubectl drain node-a-17 --ignore-daemonsets --delete-emptydir-data --timeout=11m # patch and reboot the machine kubectl uncordon node-a-17 ``` Running `cordon` before `drain` is redundant, since drain cordons anyway, but it is common in runbooks: it fences the node early, for example while you wait for the maintenance window to open, and makes the intent visible in `kubectl get nodes` to anyone else on call. If you drain several nodes, cordon them all first. Otherwise a pod evicted from the first node can be scheduled onto the second and be evicted again a few minutes later. ## Why interviewers ask - It separates people who have operated a cluster from people who have only deployed to one. - The follow-ups expose real understanding: why DaemonSet pods remain, why a bare pod needs `--force`, why uncordon does not rebalance. - Upgrades, node-image rotation and hardware work all start with this sequence, so every later operations question assumes it.

  • Why do DaemonSet pods keep running, and can even be scheduled, on a cordoned Kubernetes node?
    Cordoning makes the node lifecycle controller add the `node.kubernetes.io/unschedulable:NoSchedule` taint, and the DaemonSet controller gives every DaemonSet pod a toleration for it. The scheduler's unschedulable check lets through any pod that tolerates that taint. That is deliberate: per-node agents such as the CNI plugin or a log shipper must keep running while the node is being emptied, and `kubectl drain` leaves them in place.
  • Does `kubectl drain` kill pods immediately, or do they get a graceful shutdown?
    They get a graceful shutdown. Each eviction starts normal pod termination: the pod receives SIGTERM and has its own `terminationGracePeriodSeconds` before SIGKILL. `kubectl drain --grace-period` can override that value, and the default of -1 means the pod's own setting is used. Drain then waits until the pods are actually gone before reporting the node as drained.
  • After `kubectl uncordon`, the node stays nearly empty for hours. Is something broken?
    No. The scheduler only places pods when they are created; it never moves running pods. The node fills up as rollouts, scale-ups and restarts create new pods. If you need the load rebalanced sooner, use a descheduler or a rollout restart of the affected workloads, accepting the extra disruption that brings.

Cordoning is closing a restaurant's door to new guests while the diners already seated finish their meals; draining is asking those diners to move to the restaurant next door. Reopening the door does not bring anyone back.

saying these in an interview costs you the question

  • Cordoning a node evicts or restarts the pods already running on it
  • kubectl drain deletes DaemonSet pods when --ignore-daemonsets is passed
  • Uncordoning moves the evicted pods back to the node
  • Drain migrates a running pod live to another node
  • A failed drain automatically uncordons the node again
  • Drain kills pods instantly without a termination grace period