skip to content

In a Kubernetes cluster with a spot node pool, which workloads belong on those nodes, and how do you keep everything else off them?

level: juniorimportance: must knowfreq 62%

answer

  1. cheap because reclaimable
  2. unit of loss vs redundancy
  3. taint repels, label attracts
  4. toleration is permission only

basics

~20 s

Spot nodes suit replicated, stateless or checkpointing work that can lose a node at short notice. Label the pool by capacity type and taint it, so only pods that tolerate the taint and select the label land there.

solid answer

~40 s

Spot or preemptible nodes are cheap because the provider can take them back with only a short warning. Good tenants are workloads that survive losing a node: stateless Deployments with several replicas, idempotent queue workers, retryable CI builds and batch Jobs that checkpoint. Keep single-replica services, quorum or single-copy stateful workloads, and pods that need minutes to shut down on on-demand capacity. The pool is fenced in two halves. A `NoSchedule` taint on every spot node (for example `example.com/spot=true:NoSchedule`) keeps out any pod that lacks a matching toleration. A capacity-type label (such as `karpenter.sh/capacity-type=spot`) lets the pods that should run there require it through `nodeSelector` or node affinity. A toleration only *permits* a node, so you need both halves to actually *place* a pod there.

code

bash · 2 lines
bash
kubectl taint nodes spot-node-a17 example.com/spot=true:NoSchedule
kubectl get nodes -L karpenter.sh/capacity-type

go deeper

for a junior

Remember the two halves: a NoSchedule taint keeps pods off spot unless they tolerate it, and a capacity-type label lets the right pods ask for spot.

for a middle

Explain why a toleration is permission rather than attraction, and show the nodeSelector or node affinity that completes the fence.

for a senior

Classify workloads by unit of loss versus redundancy and by shutdown time, and catch blanket tolerations that quietly open the pool to everything.

for a principal

Treat the fence as platform policy: who may tolerate the spot taint, how that is enforced at admission, and what that saves against the risk it accepts.

## Why spot capacity is different Cloud providers sell spare machines as **spot** or **preemptible** capacity at a steep discount. The catch is that the provider can **reclaim** them whenever it needs the capacity back, with only a short notice, often around a couple of minutes and sometimes less. To Kubernetes, a reclaimed spot node is a Node that disappears, usually cleanly announced but always on the provider's schedule, never yours. The design question is therefore: **which workloads can lose a node at any moment without users noticing?** Everything else must be kept away from the pool. ## What belongs on spot | Workload shape | On spot? | Why | |---|---|---| | Stateless Deployment with many replicas behind a Service | Yes | Losing one or two replicas is absorbed by the rest | | Queue consumers with idempotent processing | Yes | An unacknowledged message is simply redelivered | | Batch `Job` that checkpoints progress | Yes | A restarted pod resumes from the checkpoint | | CI runners | Yes | A failed build is retried | | Single-replica service | No | One reclaim is a full outage | | Single-copy or quorum stateful workload | No | Losing a member can lose data or quorum | | Pod needing many minutes of `terminationGracePeriodSeconds` | No | The machine is gone before shutdown finishes | The general rule: a workload belongs on spot when its **unit of loss** (one node's worth of pods) is smaller than its **redundancy**, and its shutdown fits inside the notice. ## Fencing the pool: label plus taint There are two mechanisms, and each does one job: - **A taint repels.** A taint with effect `NoSchedule` on every spot node means the scheduler will not place a pod there unless the pod carries a matching **toleration**. This protects the unaware majority: a new team's Deployment never lands on spot by accident. - **A label attracts.** A **capacity-type label** on each node says what kind of capacity it is. Node provisioners set one automatically; Karpenter, for example, labels its nodes `karpenter.sh/capacity-type` with values such as `spot` and `on-demand`. With node groups you apply your own label. A pod that *should* run on spot requires that label through `nodeSelector` or `requiredDuringSchedulingIgnoredDuringExecution` node affinity. ## Why both halves are needed A toleration is **permission, not preference**. A pod that tolerates the spot taint can land on spot nodes *or* on any untainted on-demand node, so a batch Deployment you meant to run cheaply may quietly fill expensive capacity. The reverse mistake is also common: a pod with only a `nodeSelector` for the spot label and no toleration stays `Pending`, because the taint still repels it. So the usual shapes are: 1. **Spot-only work** (batch, CI): toleration plus a required node selector or node affinity on the spot label. 2. **Spot-allowed work** (replicated stateless services): a toleration, optionally with *preferred* node affinity toward spot, and a spread rule so not every replica sits on spot. 3. **Everything else**: no toleration, so the taint keeps it off. ## Common mistakes - Labelling the pool but not tainting it, so every workload in the cluster can drift onto spot. - Tainting with `PreferNoSchedule`, which is only a soft hint; under pressure, unaware pods still land on spot. - Putting a DaemonSet-style agent's tolerations on application pods by copy-paste, for example a blanket `operator: Exists` toleration with no key, which tolerates every taint in the cluster. - Treating spot as fine for anything with two replicas, when both replicas can sit on spot nodes that are reclaimed in the same wave. ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: ledger-reindex spec: replicas: 6 selector: matchLabels: app: ledger-reindex template: metadata: labels: app: ledger-reindex spec: nodeSelector: karpenter.sh/capacity-type: spot tolerations: - key: example.com/spot operator: Equal value: "true" effect: NoSchedule containers: - name: worker image: registry.example.com/ledger-reindex:1.4.2 ```

  • Why is a toleration alone not enough to run a batch Deployment on the spot pool?
    A toleration only lifts the taint's ban; it does not attract the pod. The scheduler still scores untainted on-demand nodes as candidates, so the pods may fill on-demand capacity instead. To pin them to spot, add a `nodeSelector` or required node affinity on the capacity-type label. To merely prefer spot, use preferred node affinity with a weight.
  • A GPU training job wants cheap interruptible GPUs. What changes compared with CPU batch work on spot?
    The fence stacks. The GPU pool usually has its own taint, so a spot GPU node carries both the accelerator taint and the spot taint, and the pod must tolerate both and select both labels. The bigger change is the workload: one lost GPU node can cost hours of training, so frequent checkpoints to durable storage matter more than on CPU pools.

A spot pool is a standby seat on a flight: fine for a flexible traveller, wrong for someone who must arrive. The taint is the gate agent turning away passengers without a standby ticket; the label is the traveller asking for standby in the first place.

saying these in an interview costs you the question

  • A toleration makes the pod run on the tainted spot nodes
  • Any workload with two replicas is safe on spot
  • Labelling the spot pool is enough to keep other workloads away
  • PreferNoSchedule reliably keeps unaware pods off spot
  • Spot nodes are just slower on-demand nodes