skip to content

Explain the mechanism by which `kubectl drain` and a PodDisruptionBudget interact — which API is involved, what response signals a blocked eviction, and which pod removals bypass the whole mechanism.

level: seniorimportance: must knowfreq 46%

answer

  1. cordon → filter → POST pods/eviction
  2. 429 TooManyRequests = budget exhausted
  3. drain retries; hang ≠ error
  4. budget recovers only when replacement is Ready
  5. delete pod / --disable-eviction / kubelet pressure bypass it

basics

~20 s

drain cordons the node, then calls the pods/eviction subresource for each pod. The disruption controller checks the matching PDB and returns HTTP 429 TooManyRequests when disruptionsAllowed is zero; drain retries until it succeeds or times out. Direct pod deletion, kubelet node-pressure eviction, and scheduler preemption all bypass eviction and ignore PDBs.

solid answer

~60 s

`kubectl drain` is a client-side loop, not a server-side operation: 1. It **cordons** the node — sets `spec.unschedulable: true` so no new pods land there. 2. It lists the node's pods, skipping mirror/static pods, and (with the right flags) DaemonSet pods and pods with local storage. 3. For each remaining pod it POSTs an `Eviction` object to the pod's **`pods/eviction`** subresource. The API server hands that to the **disruption controller**, which finds PDBs whose selector matches the pod and checks `status.disruptionsAllowed`. If it is zero, the request fails with **HTTP 429 TooManyRequests** and the message *"Cannot evict pod as it would violate the pod's disruption budget"*. `kubectl drain` retries indefinitely (or until `--timeout`), which is why a blocked drain looks like a hang rather than an error. When the eviction is allowed, it becomes an ordinary graceful delete: `preStop` hook, `SIGTERM`, `terminationGracePeriodSeconds`, then `SIGKILL`. The PDB governs *whether*, not *how gracefully*. Bypasses: `kubectl delete pod`, `--disable-eviction` on drain, kubelet node-pressure eviction, and scheduler preemption — none consult a PDB.

code

bash · 5 lines
bash
kubectl drain node-7 --ignore-daemonsets --delete-emptydir-data --timeout=15m

# evicting pod prod/api-6d9c7f5b8-2xq4t
# error when evicting pod "api-6d9c7f5b8-2xq4t" (will retry after 5s):
#   Cannot evict pod as it would violate the pod's disruption budget.

go deeper

for a junior

Know that drain cordons then evicts, and that a PDB can refuse an eviction.

for a middle

Name the pods/eviction subresource and the 429 response, and explain the retry loop.

for a senior

Explain budget recovery on Ready replacements, why a slow drain may be correct, and enumerate every bypass path.

for a principal

Discuss the cooperation-not-enforcement nature of eviction and what guardrails (admission policy, upgrade tooling, --disable-eviction hygiene) a platform needs as a result.

## drain is a client, not a server feature There is no "drain" API. `kubectl drain` is a helper that composes existing primitives, which matters because anything else can implement the same loop — and the cluster autoscaler, managed node-pool upgraders and various operators do exactly that. The sequence: 1. **Cordon** — PATCH the Node with `spec.unschedulable: true`. Existing pods keep running; the scheduler stops placing new ones. (`kubectl cordon` alone stops there.) 2. **Filter** — drain refuses to proceed if it finds pods it cannot safely handle unless you opt in: DaemonSet-managed pods need `--ignore-daemonsets` (they would be immediately recreated on the same node anyway), pods with `emptyDir` need `--delete-emptydir-data`, and unmanaged bare pods need `--force` because nothing will recreate them. 3. **Evict** — POST an `Eviction` (`policy/v1`) to `/api/v1/namespaces/<ns>/pods/<name>/eviction` for each remaining pod. ## What the Eviction API does The eviction subresource is a **policy-checked delete**. On receiving it, the API server asks the disruption controller: do any PDBs select this pod, and would removing it push `currentHealthy` below `desiredHealthy`? - **Allowed** → the pod is deleted gracefully. `preStop` runs, the container gets `SIGTERM`, endpoints are removed in parallel, and after `terminationGracePeriodSeconds` any survivor gets `SIGKILL`. The PDB's `disruptionsAllowed` is decremented and only recovers when a replacement pod becomes **Ready**. - **Denied** → **HTTP 429 TooManyRequests**, with a `DisruptionBudget` cause. `kubectl drain` prints `error when evicting pod "x" (will retry after 5s): Cannot evict pod as it would violate the pod's disruption budget` and loops. - **Matched by two PDBs** → the eviction is refused outright; overlapping budgets are treated as unresolvable rather than intersected. The 429 is deliberate: it is a *retryable* status, not a permanent failure. The contract is "not now", and a well-behaved client waits for replacements to become Ready and tries again. That is the pacing mechanism that turns a node drain into a controlled rolling handover. ## The recovery loop Because the budget only recovers on Ready replacements, drain speed is bounded by the workload's startup time, not by the drain command. A service with a 90-second readiness warm-up and `maxUnavailable: 1` drains one pod every ~90 seconds. On a node hosting 30 such pods, a drain that "hangs for an hour" may be working exactly as designed. Checking `kubectl get pdb` repeatedly and watching `disruptionsAllowed` oscillate between 0 and 1 distinguishes "progressing slowly" from "stuck". ## What bypasses PDBs entirely - **`kubectl delete pod`** — hits the pod resource, not the eviction subresource. No budget check. - **`kubectl drain --disable-eviction`** — explicitly deletes instead of evicting; the documented escape hatch when you have decided the budget must be ignored, and the reason PDBs are a cooperation mechanism rather than an enforcement boundary. - **kubelet node-pressure eviction** — under memory/disk pressure the kubelet kills pods locally to save the node. It never asks the API server's disruption controller. - **Scheduler preemption** — a higher-priority pending pod causes lower-priority pods to be deleted to free room. Preemption makes a best-effort attempt to respect PDBs but is explicitly not guaranteed by them. - **Node deletion / VM termination** — nothing to evict. ## Related knobs worth naming - **`unhealthyPodEvictionPolicy: AlwaysAllow`** on the PDB lets running-but-not-Ready pods be evicted even when the budget is exhausted. Without it, a workload that is failing readiness pins `disruptionsAllowed` at zero and blocks drains precisely when the pods are useless anyway. - **`terminationGracePeriodSeconds` and `preStop`** determine whether an allowed eviction is *safe* for in-flight requests. A PDB with no graceful shutdown still drops connections — the budget controls concurrency of disruption, not correctness of shutdown. - **`--pod-selector`** on drain lets you evict a subset, useful when triaging which workload is blocking. ## Interview-grade summary Drain = cordon + repeated eviction calls. The Eviction API is the only place a PDB is enforced. 429 means "budget exhausted, retry". Anything that deletes pods without going through eviction ignores the budget completely — which is both the mechanism's simplicity and its limitation.

  • A drain has been retrying for 20 minutes on a node with 30 pods. How do you tell whether it is progressing or genuinely stuck?
    Watch `kubectl get pdb` and the node's pod count over a few minutes. If `disruptionsAllowed` keeps cycling 0 → 1 and the pod count on the node is falling, it is pacing correctly and is bounded by replacement readiness time. If `disruptionsAllowed` is pinned at 0 and no pods leave the node, something is preventing replacements from becoming Ready — that is a stuck drain.
  • What actually happens to a pod once its eviction is permitted?
    It becomes a normal graceful deletion: the pod is removed from Service endpoints, its `preStop` hook runs, containers receive SIGTERM, and after `terminationGracePeriodSeconds` anything still alive is SIGKILLed. The PDB decides whether the eviction may happen at all; it has no influence on shutdown behavior, so connection draining still depends on the pod's own grace period and hooks.
  • Why is HTTP 429 the chosen status for a blocked eviction rather than 403 or 409?
    429 TooManyRequests signals a rate/temporary condition — the request is legitimate but cannot be served right now. That tells clients to back off and retry, which is exactly the pacing behavior a drain needs. A 403 would imply a permanent authorization failure and would make well-behaved clients give up instead of waiting for capacity to recover.

saying these in an interview costs you the question

  • Thinking drain is a single server-side API call rather than a cordon plus eviction loop
  • Reading a retrying drain as an error rather than as budget-paced progress
  • Believing kubectl delete pod is subject to the PDB
  • Assuming kubelet node-pressure eviction or preemption honors PDBs
  • Expecting the budget to recover as soon as a replacement pod is Running rather than Ready

context