How do low-priority placeholder pods running the pause image give a Kubernetes cluster headroom for pods displaced by spot reclaims?
answer
- provisioning slower than the notice
- preempt the placeholders first
- Pending placeholders trigger refill
- expendable cutoff defaults to -10
- zero grace for fast preemption
basics
~10 sPlaceholder pods with a very low PriorityClass reserve spare capacity. Displaced pods preempt them and start at once, and the now-Pending placeholders make the Cluster Autoscaler add a node to restore the headroom.
solid answer
~40 sNode provisioning takes minutes, which is longer than a reclaim notice. A placeholder Deployment runs the `pause` image with CPU and memory requests sized to the headroom you want, under a PriorityClass below every real workload, conventionally `-10`. When an interruption handler evicts spot pods, their replacements (priority 0 or higher) cannot find room, so kube-scheduler preempts placeholders and the real pods start in seconds. The evicted placeholders go `Pending`, and the Cluster Autoscaler provisions a node to fit them, restoring the buffer. The value matters: pods below the Cluster Autoscaler's `--expendable-pods-priority-cutoff` (default `-10`) are expendable and never trigger scale-up. So `-10` works and `-11` silently disables the refill. Set `terminationGracePeriodSeconds: 0` on placeholders so preemption frees space immediately, and keep the headroom off spot nodes that could be reclaimed in the same wave.
code
yaml · 25 linesapiVersion: apps/v1
kind: Deployment
metadata:
name: spot-headroom
spec:
replicas: 4
selector:
matchLabels:
app: spot-headroom
template:
metadata:
labels:
app: spot-headroom
spec:
priorityClassName: spot-headroom
terminationGracePeriodSeconds: 0
nodeSelector:
karpenter.sh/capacity-type: on-demand
containers:
- name: reserve
image: registry.k8s.io/pause:3.10.2
resources:
requests:
cpu: "2"
memory: 4Gigo deeper
Recall the idea: low-priority dummy pods hold room, and real pods push them out when they need space.
Explain the chain: preemption frees room, placeholders go Pending, and the Cluster Autoscaler adds a node, with the cutoff that makes the refill work.
Size the buffer to a reclaim wave, match placeholder shapes to real pods, and keep placeholders on capacity that survives the wave.
Treat headroom as an explicit insurance premium and compare its idle cost with the brief degraded capacity it removes.
## The gap headroom fills When spot nodes are reclaimed, their pods need somewhere to go **now**. A node autoscaler reacts to `Pending` pods, but a new machine takes minutes to boot, join and become Ready. For those minutes the displaced replicas sit Pending while the survivors carry the load. **Headroom** is spare capacity kept free ahead of time so displaced pods start at once. Free capacity does not stay free by itself. Autoscalers remove underused nodes, and other workloads fill any gap. **Placeholder pods** reserve it explicitly. ## How the placeholder pattern works 1. **A low PriorityClass.** Create a PriorityClass whose value is below every real workload's (real pods use 0 or higher). The Cluster Autoscaler FAQ's recipe uses `-10`. 2. **A pause Deployment.** Run replicas of the tiny `pause` container, which just sleeps, with `resources.requests` sized to the headroom, for example 4 replicas of 2 CPU and 4Gi. 3. **Displacement.** Spot pods are evicted, and their controllers create replacements. The replacements find no free room, but they outrank the placeholders. 4. **Preemption.** kube-scheduler picks placeholder victims, deletes them, and sets `nominatedNodeName` on the waiting pod. Once the victims are gone, the real pods bind and start. 5. **Refill.** The placeholders' controller recreates them. Now they are `Pending`, so the Cluster Autoscaler adds a node that fits them, and the buffer is back. ## The two settings people get wrong | Setting | Right value | What goes wrong otherwise | |---|---|---| | PriorityClass `value` | `-10` (at or above the Cluster Autoscaler cutoff) | Below the cutoff, Pending placeholders never trigger scale-up, so headroom is used once and never restored | | Placeholder `terminationGracePeriodSeconds` | `0` | Preemption waits for victims to terminate, so the default 30 seconds delays every displaced pod | The Cluster Autoscaler flag `--expendable-pods-priority-cutoff` defaults to **-10**. Pods with priority **below** it are expendable: they do not trigger scale-up, and they do not block scale-down. A value of exactly -10 is not below the cutoff, so it still triggers scale-up. ## Sizing and placing the buffer - **Size it to a reclaim wave**, not to one pod: roughly the requests of the largest group of spot nodes you expect to lose together. - **Keep the headroom on capacity that survives the wave.** Placeholders on spot nodes vanish in the same reclaim they were meant to absorb, so place them on on-demand nodes, or at least in other pools. - **Match the shapes.** A buffer made of 1-CPU placeholders cannot fit a 6-CPU pod even if the total is enough; preemption frees space only on one node at a time. - **Scale it with the cluster** if a fixed size stops making sense. The Cluster Autoscaler FAQ pairs placeholders with a proportional autoscaler for this. - **Watch the cost.** Headroom is paid-for idle capacity, the premium you pay to make reclaims invisible. ## Failure modes to watch - **Silent drain of the buffer.** If refills stop, for example because the class drops below the cutoff or a node-group maximum is reached, the placeholders sit `Pending` forever. Alert on Pending placeholder pods that stay Pending past a few minutes. - **Scale-down fights.** Placeholders at or above the cutoff are not expendable, and their requests count toward node utilization, so the Cluster Autoscaler does not see the nodes holding them as idle. That is the point, but it also means the buffer is billed continuously. - **Wrong victims.** If some real workload runs at a priority below the placeholders, displaced pods can preempt it instead. Keep the placeholder class the lowest in the cluster. ## Where it fits Priority and preemption mechanics, and node provisioning itself, are separate subjects. Here the point is the combination: preemption gives instant room, and the autoscaler restores it afterwards. A just-in-time provisioner can also react to Pending placeholders, but check how it treats low-priority pods before relying on the same pattern. ```yaml apiVersion: scheduling.k8s.io/v1 kind: PriorityClass metadata: name: spot-headroom value: -10 globalDefault: false preemptionPolicy: Never description: Placeholder pods reserving room for displaced spot workloads ```
- Would placeholder pods help on a 5-node edge cluster in a retail store with no node autoscaler?Mostly not. The refill half needs an autoscaler, so a preempted buffer is never restored. What remains is a reservation that any higher-priority pod can take for any reason, not only for failover. On a fixed cluster, it is usually better to size the five nodes so that four can carry the critical workloads, and use PriorityClasses to decide what gets dropped.
- Why does the example PriorityClass set preemptionPolicy: Never?It is a guard rather than a necessity. With `Never`, a Pending placeholder waits in the queue instead of trying to preempt anything. At -10 there is usually nothing lower to preempt, but if someone later adds an even lower class, the placeholders will not start evicting it.
Placeholder pods are coats saved on seats at a busy concert: when a friend arrives they take the seat at once, and the coat's owner goes looking for another seat, which is the refill.
saying these in an interview costs you the question
- Any negative priority still triggers Cluster Autoscaler scale-up
- Placeholder pods should run on the spot nodes they protect
- Preemption is instant regardless of the victim's grace period
- Total free CPU is enough, whatever the placeholder shape
- Headroom is free because pause containers use no resources