As the owner of a shared Kubernetes platform, how would you decide how much capacity runs on spot nodes, and what rules would every team have to follow?
answer
- inventory before discount
- correlated reclaims set the floor
- admission policy guards the taint
- diversify pools, fall back
- churn is the hidden cost
basics
~20 sLet the workload mix set the spot share, not the discount: only interruption-tolerant work goes there, critical services keep an on-demand floor that serves peak alone, spot pools are diversified, and admission policy enforces who may tolerate them.
solid answer
~50 sI would start from the workload inventory, not the price. Batch, CI, queue consumers and the surplus replicas of stateless services qualify, and that sets the ceiling for the spot share. For each critical service, such as a payments-authorization API, the on-demand floor must serve peak if *every* spot node disappears, because reclaims are correlated. The rules I would set: only approved namespaces may tolerate the spot taint, enforced by an admission policy. No single-replica or single-copy stateful workloads on spot. Shutdown must fit the notice. Disruption budgets must be sized so one spot node's drain always succeeds. On the platform side: diversify spot pools across instance shapes and zones, fall back to on-demand when spot is unavailable, run interruption handling as a platform component, and run game days that drain all spot at once. The cost is complexity and churn, so I would measure savings against reclaim-driven incidents.
go deeper
Know that not every workload belongs on spot, and that critical services keep some replicas on on-demand capacity.
Explain why a spread rule is not a floor, and why disruption budgets must allow one spot node's drain.
Design the per-service floor, headroom and zone spread, and prove them with a drain of all spot nodes.
Own the trade: derive the spot share from the workload inventory, enforce it at admission, and weigh savings against churn, complexity and incident cost.
## Start from the workloads, not the discount The tempting plan is "move 70% to spot and save". The honest plan starts with an **inventory** that classifies each workload by how it survives losing a node: | Class | Spot policy | Examples | |---|---|---| | Interruptible by design | Spot by default | CI, batch with checkpoints, idempotent queue consumers | | Replicated stateless | Surplus on spot, floor on on-demand | Web APIs, a payments-authorization API | | Stateful with replication and quorum awareness | Case by case, usually not | Replicated caches, some data systems | | Singletons, single-copy state, long shutdowns | Never | Leader-only jobs, single-instance databases | The spot share is an **output** of this table. If only 35% of requested CPU is honestly interruptible, the target is 35%, whatever the finance slide says. ## The floor rule Reclaims are **correlated**: the provider takes back capacity from a pool, often many machines at once. So for every service that matters, the platform's rule is: - **The on-demand floor must carry peak with all spot gone.** A spread rule that balances replicas is not a floor. Pin the floor with its own Deployment, or with required node affinity. - **Only surplus rides spot**, which is where HPA-driven bursts belong. - **Headroom** (low-priority placeholder pods on on-demand nodes) shortens the gap while replacements provision. This ties the spot plan to capacity planning. The floor is effectively your N+1 thinking, applied to a capacity type instead of a zone. ## Rules every team follows 1. **The spot taint is tolerated only by approved namespaces or workloads**, enforced at admission by a policy engine, not by code review. 2. **No single-replica Deployments and no single-copy StatefulSets on spot**, enforced by the same policy. 3. **`terminationGracePeriodSeconds` must fit the notice with slack**. Anything that needs longer stays on on-demand. 4. **Disruption budgets must let one spot node drain**. A budget that blocks a drain turns a graceful eviction into a hard kill at reclaim time. 5. **Spread across zones** as well as capacity types, with every node carrying the capacity-type label that spread rules rely on. ## Platform-side obligations - **Diversify spot pools** across several instance shapes and zones, so one pool's reclaim wave takes a fraction of the fleet. - **Fallback to on-demand** when spot is unavailable. A provisioner that can choose either capacity type, or a Cluster Autoscaler using the `priority` expander over node groups, keeps Pending pods from waiting on spot that never comes. - **Interruption handling is a platform component** with an owner, monitoring and alerts. Individual teams should not each reinvent it. - **Observability**: reclaim counts per pool, time from notice to pods Ready elsewhere, and evictions refused by budgets during drains. - **Game days**: cordon and drain every spot node in a staging cluster, or in production during a quiet window, and verify the floors hold. ## What the policy costs - **Churn.** Every reclaim restarts pods, cold caches and warm-up paths, and restarts connection pools. Services with long warm-up pay this repeatedly. - **Complexity.** Two Deployments per critical service, admission policies, a handler to operate, and another failure mode in every incident review. - **Idle premium.** Headroom and on-demand floors are paid for whether or not a reclaim happens. - **Hidden coupling.** A spot shortage across a whole region can push everything onto on-demand at once, so the on-demand quota and budget must cover that case. ## Deciding and revisiting Publish the rules with the numbers behind them: spot share, floor per tier and headroom size. Review them quarterly against reclaim-driven incidents and realised savings. If reclaims keep causing customer-visible errors for a tier, move that tier off spot. The saving was never worth an incident budget spent on it.
- Finance asks for 70% of compute on spot, but your inventory says 35% is interruptible. What do you do?Show the inventory and the floor arithmetic: what would run on spot beyond 35%, and which incidents that invites. Offer the real levers instead, such as making more workloads interruptible (idempotency, checkpointing, shorter shutdown) and right-sizing requests. Each workload re-classified raises the honest share.
- How would you verify that the on-demand floors actually hold?Run a game day that cordons and drains every spot node together, under representative load, and watch error rates, latency and Pending pods for each critical service. Track notice-to-ready times and budget-refused evictions from real reclaims, too. A floor that has never lost all its spot peers has not been proven.
saying these in an interview costs you the question
- Set the spot share from the discount, then fit workloads to it
- Topology spread across capacity types guarantees enough on-demand replicas
- Reclaims are independent events, so diversification adds nothing
- Teams can each run their own interruption handling
- Spot savings come with no operational cost