skip to content

Spot & Interrupted Capacity

Spot and preemptible nodes are cheap because the cloud can reclaim them at a couple of minutes' notice. Surviving that means tainting the pool, draining on the interruption signal, and keeping single-replica or stateful workloads off it entirely.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

5

In a Kubernetes cluster with a spot node pool, which workloads belong on those nodes, and how do you keep everything else off them?

level: juniorimportance: must knowfreq 62%

answer

  1. cheap because reclaimable
  2. unit of loss vs redundancy
  3. taint repels, label attracts
  4. toleration is permission only

basics

~20 s

Spot nodes suit replicated, stateless or checkpointing work that can lose a node at short notice. Label the pool by capacity type and taint it, so only pods that tolerate the taint and select the label land there.

solid answer

~40 s

Spot or preemptible nodes are cheap because the provider can take them back with only a short warning. Good tenants are workloads that survive losing a node: stateless Deployments with several replicas, idempotent queue workers, retryable CI builds and batch Jobs that checkpoint. Keep single-replica services, quorum or single-copy stateful workloads, and pods that need minutes to shut down on on-demand capacity. The pool is fenced in two halves. A `NoSchedule` taint on every spot node (for example `example.com/spot=true:NoSchedule`) keeps out any pod that lacks a matching toleration. A capacity-type label (such as `karpenter.sh/capacity-type=spot`) lets the pods that should run there require it through `nodeSelector` or node affinity. A toleration only *permits* a node, so you need both halves to actually *place* a pod there.

code

bash · 2 lines
bash
kubectl taint nodes spot-node-a17 example.com/spot=true:NoSchedule
kubectl get nodes -L karpenter.sh/capacity-type

go deeper

for a junior

Remember the two halves: a NoSchedule taint keeps pods off spot unless they tolerate it, and a capacity-type label lets the right pods ask for spot.

for a middle

Explain why a toleration is permission rather than attraction, and show the nodeSelector or node affinity that completes the fence.

for a senior

Classify workloads by unit of loss versus redundancy and by shutdown time, and catch blanket tolerations that quietly open the pool to everything.

for a principal

Treat the fence as platform policy: who may tolerate the spot taint, how that is enforced at admission, and what that saves against the risk it accepts.

## Why spot capacity is different Cloud providers sell spare machines as **spot** or **preemptible** capacity at a steep discount. The catch is that the provider can **reclaim** them whenever it needs the capacity back, with only a short notice, often around a couple of minutes and sometimes less. To Kubernetes, a reclaimed spot node is a Node that disappears, usually cleanly announced but always on the provider's schedule, never yours. The design question is therefore: **which workloads can lose a node at any moment without users noticing?** Everything else must be kept away from the pool. ## What belongs on spot | Workload shape | On spot? | Why | |---|---|---| | Stateless Deployment with many replicas behind a Service | Yes | Losing one or two replicas is absorbed by the rest | | Queue consumers with idempotent processing | Yes | An unacknowledged message is simply redelivered | | Batch `Job` that checkpoints progress | Yes | A restarted pod resumes from the checkpoint | | CI runners | Yes | A failed build is retried | | Single-replica service | No | One reclaim is a full outage | | Single-copy or quorum stateful workload | No | Losing a member can lose data or quorum | | Pod needing many minutes of `terminationGracePeriodSeconds` | No | The machine is gone before shutdown finishes | The general rule: a workload belongs on spot when its **unit of loss** (one node's worth of pods) is smaller than its **redundancy**, and its shutdown fits inside the notice. ## Fencing the pool: label plus taint There are two mechanisms, and each does one job: - **A taint repels.** A taint with effect `NoSchedule` on every spot node means the scheduler will not place a pod there unless the pod carries a matching **toleration**. This protects the unaware majority: a new team's Deployment never lands on spot by accident. - **A label attracts.** A **capacity-type label** on each node says what kind of capacity it is. Node provisioners set one automatically; Karpenter, for example, labels its nodes `karpenter.sh/capacity-type` with values such as `spot` and `on-demand`. With node groups you apply your own label. A pod that *should* run on spot requires that label through `nodeSelector` or `requiredDuringSchedulingIgnoredDuringExecution` node affinity. ## Why both halves are needed A toleration is **permission, not preference**. A pod that tolerates the spot taint can land on spot nodes *or* on any untainted on-demand node, so a batch Deployment you meant to run cheaply may quietly fill expensive capacity. The reverse mistake is also common: a pod with only a `nodeSelector` for the spot label and no toleration stays `Pending`, because the taint still repels it. So the usual shapes are: 1. **Spot-only work** (batch, CI): toleration plus a required node selector or node affinity on the spot label. 2. **Spot-allowed work** (replicated stateless services): a toleration, optionally with *preferred* node affinity toward spot, and a spread rule so not every replica sits on spot. 3. **Everything else**: no toleration, so the taint keeps it off. ## Common mistakes - Labelling the pool but not tainting it, so every workload in the cluster can drift onto spot. - Tainting with `PreferNoSchedule`, which is only a soft hint; under pressure, unaware pods still land on spot. - Putting a DaemonSet-style agent's tolerations on application pods by copy-paste, for example a blanket `operator: Exists` toleration with no key, which tolerates every taint in the cluster. - Treating spot as fine for anything with two replicas, when both replicas can sit on spot nodes that are reclaimed in the same wave. ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: ledger-reindex spec: replicas: 6 selector: matchLabels: app: ledger-reindex template: metadata: labels: app: ledger-reindex spec: nodeSelector: karpenter.sh/capacity-type: spot tolerations: - key: example.com/spot operator: Equal value: "true" effect: NoSchedule containers: - name: worker image: registry.example.com/ledger-reindex:1.4.2 ```

  • Why is a toleration alone not enough to run a batch Deployment on the spot pool?
    A toleration only lifts the taint's ban; it does not attract the pod. The scheduler still scores untainted on-demand nodes as candidates, so the pods may fill on-demand capacity instead. To pin them to spot, add a `nodeSelector` or required node affinity on the capacity-type label. To merely prefer spot, use preferred node affinity with a weight.
  • A GPU training job wants cheap interruptible GPUs. What changes compared with CPU batch work on spot?
    The fence stacks. The GPU pool usually has its own taint, so a spot GPU node carries both the accelerator taint and the spot taint, and the pod must tolerate both and select both labels. The bigger change is the workload: one lost GPU node can cost hours of training, so frequent checkpoints to durable storage matter more than on CPU pools.

A spot pool is a standby seat on a flight: fine for a flexible traveller, wrong for someone who must arrive. The taint is the gate agent turning away passengers without a standby ticket; the label is the traveller asking for standby in the first place.

saying these in an interview costs you the question

  • A toleration makes the pod run on the tainted spot nodes
  • Any workload with two replicas is safe on spot
  • Labelling the spot pool is enough to keep other workloads away
  • PreferNoSchedule reliably keeps unaware pods off spot
  • Spot nodes are just slower on-demand nodes
open as a page

When a cloud provider signals that a Kubernetes spot node will be reclaimed, what should happen before the machine disappears, and what makes it happen?

level: middleimportance: must knowfreq 58%

basics

~20 s

An interruption handler catches the provider's notice, cordons the node and evicts its pods through the Eviction API, so replacements schedule elsewhere. Each pod's shutdown must fit inside the notice, or it is killed when the machine goes.

open as a page

A payments-authorization API on Kubernetes runs 7 replicas, 5 on spot nodes. One reclaim wave killed 4 at once and cut its 13-minute node drain short. How would you redesign its placement?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Guarantee an on-demand floor that can serve peak alone, let only the surplus replicas use spot, spread both across zones, and cut shutdown time to fit the reclaim notice, because a disruption budget cannot stop a reclaim.

open as a page

As the owner of a shared Kubernetes platform, how would you decide how much capacity runs on spot nodes, and what rules would every team have to follow?

level: principalimportance: should knowfreq 34%

basics

~20 s

Let the workload mix set the spot share, not the discount: only interruption-tolerant work goes there, critical services keep an on-demand floor that serves peak alone, spot pools are diversified, and admission policy enforces who may tolerate them.

open as a page

How do low-priority placeholder pods running the pause image give a Kubernetes cluster headroom for pods displaced by spot reclaims?

level: middleimportance: nice to knowfreq 30%

basics

~10 s

Placeholder pods with a very low PriorityClass reserve spare capacity. Displaced pods preempt them and start at once, and the now-Pending placeholders make the Cluster Autoscaler add a node to restore the headroom.

open as a page