A payments-authorization API on Kubernetes runs 7 replicas, 5 on spot nodes. One reclaim wave killed 4 at once and cut its 13-minute node drain short. How would you redesign its placement?
answer
- notice shorter than the drain
- floor sized to peak alone
- two Deployments, one Service
- budgets cannot stop reclaims
- maxSkew 1 splits 7 as 4/3
basics
~20 sGuarantee an on-demand floor that can serve peak alone, let only the surplus replicas use spot, spread both across zones, and cut shutdown time to fit the reclaim notice, because a disruption budget cannot stop a reclaim.
solid answer
~50 sFirst, the 13-minute drain comes from `terminationGracePeriodSeconds: 780`, which can never fit a notice of about two minutes, so those pods get killed mid-shutdown. Cut shutdown to well under the notice, or keep the pods off spot. Second, nothing guaranteed a floor: the scheduler put 5 of 7 on spot, and one wave took 4. If peak needs 4 replicas, run two Deployments behind one Service: `payments-auth-ondemand` with 4 replicas pinned to on-demand, and `payments-auth-spot` with 3 that tolerate the spot taint. Losing every spot node then leaves 4. A `topologySpreadConstraint` on the capacity-type label with `maxSkew: 1` gives only a 4/3 split in either direction, not a guaranteed floor. Add zone spread for both, and a PodDisruptionBudget over the shared label. Remember that the budget gates evictions only: a reclaim is involuntary and ignores it.
code
yaml · 34 linesapiVersion: apps/v1
kind: Deployment
metadata:
name: payments-auth-spot
spec:
replicas: 3
selector:
matchLabels:
app: payments-auth
pool: spot
template:
metadata:
labels:
app: payments-auth
pool: spot
spec:
terminationGracePeriodSeconds: 75
nodeSelector:
karpenter.sh/capacity-type: spot
tolerations:
- key: example.com/spot
operator: Exists
effect: NoSchedule
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: payments-auth
pool: spot
containers:
- name: api
image: registry.example.com/payments-auth:3.8.1go deeper
Know that a spot reclaim can take several replicas at once, and that some replicas must live on on-demand capacity.
Explain why maxSkew 1 across two capacity types still leaves a floor of only 3 of 7, and why a disruption budget cannot stop a reclaim.
Size the on-demand floor from peak need, fix the shutdown path to fit the notice, and prove it with a game-day drain of all spot nodes.
Decide whether critical request paths may use spot at all, and what the saving on the 3 surplus replicas is worth against the operational complexity.
## Reading the incident Two separate failures combined: 1. **Shutdown longer than the notice.** A 13-minute drain means the pods declare `terminationGracePeriodSeconds: 780`, perhaps to finish in-flight authorization batches. A reclaim notice of about two minutes cannot hold that. The interruption handler evicted the pods, they started shutting down, and the machine vanished about 11 minutes early. In-flight work was lost without a clean handoff. 2. **No capacity-type floor.** Seven replicas with a spot toleration and nothing else let the scheduler put 5 on spot. The provider reclaims capacity in correlated waves, so 4 went at once and 3 replicas were left to carry peak. The disruption budget did not help, and it could not have. A PodDisruptionBudget only gates evictions requested through the **Eviction API**. A reclaim is an **involuntary** disruption. At most, the budget made the handler's evictions wait, which only meant less graceful time before the kill. ## Fix 1: make shutdown fit the notice - Authorization requests are short. The pod needs to stop receiving traffic, finish the requests in flight, and exit, which takes seconds rather than minutes. - Move long-running work (settlement batches, reconciliation) out of the request path into a queue consumer with idempotent processing, which is a natural spot tenant. - Set `terminationGracePeriodSeconds` to something like 60-90 seconds, leaving slack for detection and eviction inside the notice. - If some component truly needs 13 minutes to drain, it does not belong on spot at all. ## Fix 2: a guaranteed on-demand floor Compare the placement tools: | Approach | What it guarantees | Weakness | |---|---|---| | Toleration only | Nothing | The scheduler may put most replicas on spot | | Preferred node affinity to on-demand | A preference | Under pressure, replicas still land on spot | | Spread over capacity type, `maxSkew: 1` | A 4/3 split of 7, in either direction | The 4 can be on spot; the floor is 3 | | Two Deployments, one pinned on-demand | An exact floor | Two objects to scale and roll out | If peak needs **4** replicas, the only arrangement that survives losing *all* spot is to pin 4 to on-demand. So run `payments-auth-ondemand` (4 replicas, required node affinity on the on-demand label, no spot toleration) and `payments-auth-spot` (3 replicas, spot toleration and spot selector). Both carry `app: payments-auth`, which the Service and the budget select. An HPA, if any, scales the spot Deployment, so bursts ride cheap capacity while the floor stays fixed. ## Fix 3: spread both halves across zones - Add a `topologySpreadConstraint` on `topology.kubernetes.io/zone` to each Deployment, so a zone outage and a reclaim wave do not stack. - If you spread on the capacity-type label instead, **every** eligible node must carry it. With `whenUnsatisfiable: DoNotSchedule`, the PodTopologySpread filter rejects nodes missing the key. - `nodeTaintsPolicy` defaults to `Ignore`. A pod that does *not* tolerate the spot taint still counts spot nodes as a domain with zero pods, and can stay `Pending` on skew. Set `nodeTaintsPolicy: Honor` when spread domains include tainted nodes. ## Fix 4: the budget, sized for one node's drain A PodDisruptionBudget with `minAvailable: 4` over `app: payments-auth` lets the handler evict the 3 spot replicas without blocking, while stopping a careless voluntary drain of the on-demand floor. Budget semantics belong elsewhere; here, size the budget so that one spot node's drain always fits. ## Checking the result 1. `kubectl get pods -l app=payments-auth -o wide` and map each node to its capacity-type label. 2. Drain one spot node by hand in a staging cluster and time shutdown against the notice. 3. Cordon and drain *all* spot nodes during a game day and confirm the 4 on-demand replicas hold peak load. ```yaml apiVersion: policy/v1 kind: PodDisruptionBudget metadata: name: payments-auth spec: minAvailable: 4 selector: matchLabels: app: payments-auth ```
- Why not just spread all 7 replicas over the capacity-type label with maxSkew: 1?Over two domains it yields a 4/3 split, and the scheduler may put the 4 on spot. Losing all spot then leaves 3, below the peak need of 4. Spread balances the replicas but guarantees no floor, and it also needs every node labelled and `nodeTaintsPolicy` handled.
- When a scale-down removes replicas from a single mixed Deployment, how could you make spot replicas go first?Set the `controller.kubernetes.io/pod-deletion-cost` annotation lower on spot pods. On scale-down, the ReplicaSet controller prefers to delete pods with lower cost. It is best-effort and needs something to keep the annotation current, which is why two Deployments are usually simpler.
- A pod without the spot toleration stays Pending with a topology spread message, although on-demand nodes have room. What is the likely cause?The spread counts spot nodes as a domain because `nodeTaintsPolicy` defaults to `Ignore`. That domain holds zero matching pods, so the skew check fails on every on-demand node. Setting `nodeTaintsPolicy: Honor` drops nodes whose taints the pod does not tolerate from the calculation.
saying these in an interview costs you the question
- A PodDisruptionBudget protects replicas from a spot reclaim
- A long terminationGracePeriodSeconds makes spot shutdown safer
- Topology spread over capacity type guarantees an on-demand floor
- Preferred node affinity keeps enough replicas on on-demand under pressure
- Reclaims are independent, so five spot replicas rarely go together