You must drain 5 nodes of a 64-node, two-zone Kubernetes cluster running a feature-flag evaluation service, and new nodes take 19 minutes to join. How do you run that maintenance safely?
answer
- capacity before the first eviction
- requests, per zone
- cordon the batch up front
- serial drains with a health gate
- timeout leaves node cordoned
basics
~20 sConfirm the other nodes can absorb the evicted pods' requests, adding capacity before the 19-minute gap bites. Cordon the whole batch, drain one node at a time with a timeout, wait until the service is whole, then continue.
solid answer
~40 sFirst I check capacity: sum the requests of the pods on the five nodes and compare them with free allocatable on the other 59, per zone. If it falls short, I add nodes and wait out the 19 minutes before any eviction. Then I cordon all five so evicted pods don't land on a node I'm about to drain. I drain them one at a time, alternating zones, with `--ignore-daemonsets`, a deliberate `--delete-emptydir-data`, and a `--timeout` so automation can't hang. Between nodes I wait until the flag service's Deployment reports every replica available. If a drain times out, the node stays cordoned and I investigate: usually replacements are `Pending` for capacity, a pod is stuck terminating, or a budget can't be met. I don't reach for `--force` or `--disable-eviction` by reflex.
code
bash · 4 linesfor n in node-a-23 node-b-08 node-a-31 node-b-52 node-a-44; do kubectl cordon "$n"; done
kubectl drain node-a-23 --ignore-daemonsets --delete-emptydir-data --timeout=13m
kubectl rollout status deployment/flag-eval -n flags --timeout=19m
kubectl uncordon node-a-23go deeper
Know the basic order: check there is room, cordon, drain, do the work, uncordon, and never drain several nodes at once without thinking.
Explain why requests rather than usage decide where evicted pods fit, and why a slow node scale-up turns a missing capacity check into Pending pods.
Show a runbook with a capacity check per zone, batch cordon, serial bounded drains, a health gate between nodes, and a clear diagnosis path for a drain that times out.
Argue for automating this in node-lifecycle tooling with spare-capacity headroom and budgets set per service, so no one hand-runs escape-hatch flags during a window.
## The shape of the problem The cluster has 64 nodes across two zones and runs a feature-flag evaluation service that every other service calls on each request. Five nodes need maintenance, and the node group needs **19 minutes** to bring a new node to `Ready`. That delay is the whole problem: any pod evicted without somewhere to go sits `Pending` for up to 19 minutes, and a service that loses a large share of its replicas at once starts failing callers. A safe plan answers three questions before touching a node: where will the evicted pods land, how many can leave at once, and what happens when a drain does not finish. ## Step 1: prove the capacity exists Eviction only moves work if the remaining nodes can absorb it. Before starting: - List what runs on the target nodes, for example with `kubectl get pods -A --field-selector spec.nodeName=node-a-23 -o wide`. - Sum their **requests** (the scheduler places on requests, not on usage) and compare with the free allocatable on the other 59 nodes, **per zone** if the service uses zone topology spread or zone-pinned volumes. - If the shortfall is real, add nodes **first** and wait out the 19 minutes before the first eviction, rather than letting evictions trigger the scale-up. Pre-provisioning turns a 19-minute outage risk into a 19-minute wait that nobody notices. ## Step 2: cordon the whole batch, drain one node at a time 1. **Cordon all five** target nodes up front. Otherwise a pod evicted from the first node may be scheduled onto the second, and then evicted again a few minutes later. 2. **Drain one node**, with explicit flags and a bound: ```bash kubectl drain node-a-23 --ignore-daemonsets --delete-emptydir-data --timeout=13m ``` 3. **Wait until the service is whole again** before the next node: the Deployment reports all replicas available (`kubectl rollout status` or its `availableReplicas`), not just "drain returned". 4. Alternate zones, so one zone never carries the whole loss. 5. Do the maintenance, **uncordon**, and confirm the node is `Ready` before moving on. If the service's replicas are unevenly packed, look at how many live on each target node. A node holding 4 of 14 flag-evaluation pods is a bigger event than one holding none; draining it when the service is already short is the moment to wait. ## Step 3: treat a drain that does not finish as a signal `kubectl drain` without `--timeout` waits **forever**. In automation that means a job that silently hangs. With `--timeout` the command gives up and exits non-zero, and three facts follow: - The node is **still cordoned**; drain never undoes its own cordon. - Pods already evicted are gone and their replacements are running or `Pending` elsewhere. - Remaining pods are listed on stderr ("There are pending pods in node …"). The usual reasons a drain does not finish, and the check for each: | Symptom | Likely cause | What to check | |---|---|---| | repeated "will retry after 5s" eviction errors | a disruption budget cannot be satisfied, often because replacements are `Pending` | `kubectl get pods` for `Pending` replicas and their scheduling events | | an evicted pod sits in `Terminating` | a long grace period, a slow shutdown, or a finalizer | the pod's `terminationGracePeriodSeconds` and `metadata.finalizers` | | pods stay `Terminating` for a long time on a `NotReady` node | the kubelet is not there to confirm the shutdown | `--skip-wait-for-delete-timeout` so drain stops waiting for them | | drain exits at once with "cannot delete …" | a DaemonSet pod, `emptyDir` pod or bare pod | the flag named in the message, decided on purpose | The first row is the one tied to the 19-minute delay: evictions that wait for replacements are only as fast as the capacity behind them. ## Step 4: decide the escape hatches in advance `--disable-eviction` and `--force` both exist, and both remove protection. Write down, before the window, who may use them and when. Deleting pods past a budget turns a slow maintenance into a partial outage of the flag service, which then fails every caller. ## What makes this a senior answer - Capacity is checked **before** the first eviction, in requests and per zone. - Drains are **bounded** and **serial**, with a health gate between nodes. - A timed-out drain is investigated, not retried blindly or bypassed. - The runbook is written once and reused for upgrades, image rotation and hardware work.
- Why cordon all five nodes before draining the first one?A drained pod goes wherever the scheduler finds room, and an uncordoned target node looks like room. Without the batch cordon, a flag-evaluation pod evicted from the first node can land on the third and be evicted again minutes later. Each extra eviction is another cold start for callers and another chance of running short. Cordoning the batch sends replacements only to nodes that will stay.
- `kubectl drain --timeout=13m` exits with "global timeout reached". What state is the cluster in, and what do you do next?The node is still cordoned, the pods already evicted are gone, and the ones listed as pending are still on the node. I look at why they stayed: `Pending` replacements waiting for capacity, a pod stuck `Terminating` because of a long grace period or a finalizer, or a budget that cannot be met. I fix the cause and re-run the drain. If maintenance is postponed, I uncordon the node deliberately.
- One target node is `NotReady` and its evicted pods have shown `Terminating` for 40 minutes. How does `kubectl drain` help?With the kubelet gone, nothing confirms that the containers stopped, so the pods stay `Terminating` and drain keeps waiting for them. `kubectl drain --skip-wait-for-delete-timeout=300` tells drain to stop waiting for pods whose deletion started more than 300 seconds ago, so the drain can finish. It does not force-delete them. A StatefulSet pod stuck like that is replaced only once its pod object is really removed, so confirm the machine is down before anyone force-deletes it.
saying these in an interview costs you the question
- Drain all five nodes in parallel to finish the window faster
- Evictions will trigger new nodes, so no capacity check is needed
- Capacity can be judged from current CPU usage instead of requests
- A drain that returned without error means the service is healthy again
- If a drain hangs, rerun it with --disable-eviction and --force
- A timed-out drain uncordons the node automatically