When a cloud provider signals that a Kubernetes spot node will be reclaimed, what should happen before the machine disappears, and what makes it happen?
answer
- someone must read the notice
- cordon, then Eviction API
- grace period versus notice length
- shutdownGracePeriod defaults to 0s
- kubelet shutdown skips budgets
basics
~20 sAn interruption handler catches the provider's notice, cordons the node and evicts its pods through the Eviction API, so replacements schedule elsewhere. Each pod's shutdown must fit inside the notice, or it is killed when the machine goes.
solid answer
~40 sKubernetes itself does not read a provider's reclaim notice. A component does it for you: a node termination handler DaemonSet, a controller consuming the provider's interruption events, or a provisioner such as Karpenter that handles interruptions natively. On the notice, it cordons the node (sets `spec.unschedulable`) and evicts every pod through the Eviction API, much like `kubectl drain`. Each evicted pod gets SIGTERM and its `terminationGracePeriodSeconds` (default 30). Its controller creates a replacement, and the node autoscaler adds capacity if needed. The whole sequence must fit inside the notice, which is often about two minutes. A pod whose grace period is longer is killed when the machine vanishes. As a fallback, the kubelet's graceful node shutdown can terminate pods if the reclaim triggers an orderly OS shutdown, but only when `shutdownGracePeriod` is set.
code
yaml · 4 linesapiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
shutdownGracePeriod: 45s
shutdownGracePeriodCriticalPods: 15sgo deeper
Remember the order: something catches the notice, cordons the node, evicts the pods, and the controllers recreate them elsewhere.
Explain the time budget: detection plus eviction plus terminationGracePeriodSeconds must fit inside the notice, and why the kubelet fallback skips budgets.
Show you have seen blocked evictions and long grace periods turn a clean drain into a hard kill, and how you would detect it.
Weigh handler options, a DaemonSet, a queue controller or a native provisioner, against fleet size, detection latency and who owns the component.
## The problem the notice solves When a provider reclaims a spot machine, it usually sends an **interruption notice** first: a short warning, often around two minutes, readable from the instance's metadata service or from the provider's event stream. Without action, the node simply stops. The kubelet stops posting status, the Node turns `NotReady`, and after the taint-based eviction timeout its pods are deleted and recreated elsewhere, minutes after they actually died. Requests in flight fail, and load balancers keep sending traffic to dead endpoints until they notice. The goal is to turn that **involuntary** loss into an orderly **voluntary** one while the machine still exists. ## What makes it happen Kubernetes has no built-in reader for provider notices, so something outside the core does it: - **A node termination handler DaemonSet** on each spot node polls the instance metadata and acts for its own node. - **A queue-based controller** subscribes to the provider's interruption and rebalance events for the whole fleet. - **A node provisioner with native interruption handling**, such as Karpenter, which also launches replacement capacity right away. Whichever you use, the Kubernetes-side sequence is the same. ## The sequence inside the notice 1. **Cordon.** Mark the node `spec.unschedulable: true` (what `kubectl cordon` does), often with a taint too, so no new pods land on it. 2. **Evict.** Call the **Eviction API** for each pod, as `kubectl drain` does. This honours PodDisruptionBudgets, and an eviction the budget blocks is refused with HTTP 429. 3. **Terminate gracefully.** Each evicted pod enters termination. Endpoints are removed, the container gets SIGTERM, and the kubelet waits up to the pod's `terminationGracePeriodSeconds` (default **30**) before SIGKILL. 4. **Replace.** The owning controller (ReplicaSet, StatefulSet, Job) creates new pods. If no node has room, they go `Pending` and the node autoscaler provisions capacity. 5. **Reclaim.** The provider takes the machine. Anything still running dies with it. ## The time budget | Consumer of the notice | Typical cost | |---|---| | Handler detection latency | a few seconds (poll interval or queue delay) | | Eviction calls and PDB retries | seconds, or unbounded if a budget is exhausted | | Pod graceful shutdown | up to `terminationGracePeriodSeconds` each | | Replacement scheduling and startup | happens elsewhere, but capacity must exist | The hard rule: **grace period plus detection plus eviction time must fit inside the notice.** A pod with `terminationGracePeriodSeconds: 300` on a node with two minutes' warning gets SIGTERM, then is killed with the machine long before 300 seconds pass. No Kubernetes setting stretches the provider's deadline. A PDB that blocks eviction does not save the pod either: the reclaim happens anyway, now without a graceful shutdown. ## The kubelet fallback: graceful node shutdown If the reclaim itself triggers an orderly OS shutdown, the kubelet's **graceful node shutdown** can still terminate pods cleanly. The `GracefulNodeShutdown` feature gate is Beta and on by default, but it does nothing until you set `shutdownGracePeriod` in the KubeletConfiguration. Both it and `shutdownGracePeriodCriticalPods` default to `0s`. When active, the kubelet holds a systemd inhibitor lock, then kills regular pods first and critical pods last. Each pod gets the smaller of its own grace period and the time its shutdown phase allows. The kubelet kills pods **directly**, not through the Eviction API, so **PodDisruptionBudgets are not consulted**. Treat this as a safety net, not the primary path. ## Loose ends - **DaemonSet pods** are skipped by drain (`--ignore-daemonsets`) and die with the node, which is usually fine. - **Pods using `emptyDir`** need `--delete-emptydir-data` in a manual drain; their scratch data is lost either way. - **StatefulSet pods** on a node that vanished without a clean drain can stay stuck while the Node object lingers; the `node.kubernetes.io/out-of-service` taint tells Kubernetes the node is really gone. - In-container shutdown handling (preStop hooks, SIGTERM handling) is a separate subject; here it only matters that it finishes in time. ```bash # Roughly what a handler does, by hand, on notice for node spot-node-a17 kubectl cordon spot-node-a17 kubectl drain spot-node-a17 --ignore-daemonsets --delete-emptydir-data \ --grace-period=45 --timeout=100s ```
- The handler's eviction for one pod is refused with 429 because its PodDisruptionBudget is exhausted. What happens at reclaim time?The handler keeps retrying until the machine disappears. The pod then dies without a graceful shutdown, and the budget has protected nothing, because it only gates voluntary evictions. The lesson is to size budgets so a single spot node's drain always fits, and to keep workloads whose budget cannot absorb that off spot.
- With shutdownGracePeriod: 45s and shutdownGracePeriodCriticalPods: 15s, how long do regular pods get?The first 30 seconds are for regular pods and the last 15 for critical pods. A regular pod gets the smaller of its own `terminationGracePeriodSeconds` and those 30 seconds, so a pod asking for 60 is cut to 30.
saying these in an interview costs you the question
- Kubernetes reads the provider's reclaim notice by itself
- The kubelet waits for the full grace period before the machine goes
- A PodDisruptionBudget stops the provider from reclaiming the node
- Graceful node shutdown works out of the box with no kubelet config
- Pods migrate live to another node on a reclaim