For a 7-replica Kubernetes Deployment with default 25% maxSurge and maxUnavailable, what pod-count bounds hold mid-rollout, and how does rounding behave at small replica counts?
answer
- opposite rounding directions
- ceiling on total, floor on available
- available means Ready plus a delay
- small pools need one spare slot
- both zero resolves to one
basics
~20 smaxSurge rounds up and maxUnavailable rounds down. With 7 replicas, 25% becomes 2 surge and 1 unavailable: at most 9 pods and at least 6 available. At 1 to 3 replicas the defaults become surge 1, unavailable 0.
solid answer
~40 sThe RollingUpdate strategy turns percentages into pod counts with opposite rounding: `maxSurge` rounds **up** and `maxUnavailable` rounds **down**. For 7 replicas, 25% is 1.75, so surge is 2 and unavailable is 1. The Deployment may then run at most 9 pods and must keep at least 6 *available*, meaning Ready for at least `minReadySeconds`. At 1, 2 or 3 replicas the defaults give surge 1 and unavailable 0, so the rollout adds one pod and removes nothing until that pod is available. It therefore needs room to schedule one extra pod. Validation rejects both fields set to a literal 0. If both resolve to 0 only through rounding, the controller uses an unavailable budget of 1.
code
yaml · 29 linesapiVersion: apps/v1
kind: Deployment
metadata:
name: transcoder
namespace: media
spec:
replicas: 7
minReadySeconds: 45
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 0
maxUnavailable: 1
selector:
matchLabels:
app: transcoder
template:
metadata:
labels:
app: transcoder
spec:
containers:
- name: worker
image: registry.example.com/transcoder:4.12.3
resources:
requests:
memory: 2662Mi
limits:
memory: 2662Migo deeper
Know that maxSurge caps the total number of pods and maxUnavailable sets how many must stay available, and that both default to 25%.
Do the rounding out loud: surge rounds up, unavailable rounds down. Walk a 7-replica rollout and explain why small pools default to one surge and zero unavailable.
Tie the budgets to real capacity: surge pods that cannot be scheduled, quota rejections, and when to use maxSurge 0 with maxUnavailable 1 on a packed cluster.
Set platform defaults that balance rollout speed, the capacity reserved for surges, and the availability dip each service can tolerate, instead of relying on 25/25 everywhere.
## Two budgets on the RollingUpdate strategy A Deployment's default `spec.strategy.type` is `RollingUpdate`. It has two budgets under `spec.strategy.rollingUpdate`, and each can be an absolute integer or a percentage of `spec.replicas`. **Both default to 25%.** - **`maxSurge`** caps the **total** number of pods: `replicas + maxSurge`. - **`maxUnavailable`** sets the minimum number of **available** pods: `replicas - maxUnavailable`. - **Available** means the pod's Ready condition has been true for at least `minReadySeconds` (default 0). A pod that has been created but is not yet Ready counts toward the total and not toward availability. ## Rounding: surge up, unavailable down Percentages are converted with **opposite rounding**. `maxSurge` rounds **up**, so a percentage that is not zero always allows at least one extra pod. `maxUnavailable` rounds **down**, so the rollout never removes more capacity than the percentage allows. With the 25%/25% defaults: | replicas | 25% raw | maxSurge | maxUnavailable | max pods | min available | |---|---|---|---|---|---| | 1 | 0.25 | 1 | 0 | 2 | 1 | | 2 | 0.5 | 1 | 0 | 3 | 2 | | 3 | 0.75 | 1 | 0 | 4 | 3 | | 4 | 1 | 1 | 1 | 5 | 3 | | 5 | 1.25 | 2 | 1 | 7 | 4 | | 7 | 1.75 | 2 | 1 | 9 | 6 | | 13 | 3.25 | 4 | 3 | 17 | 10 | For **three replicas or fewer**, the defaults quietly mean *add one pod and remove none until that pod is available*. The rollout never drops below full capacity, but it cannot move at all unless the cluster can schedule one extra pod. ## Walking a 7-replica rollout Take a pool of 7 video-transcoding workers with the defaults, so surge 2 and unavailable 1. When the template changes, the Deployment controller works through several reconcile passes: 1. It creates the new ReplicaSet and scales it to 2, bringing the total to 9. Seven old pods are available and only 6 are required, so it also scales the old ReplicaSet down to 6. 2. The total is now 8 against a ceiling of 9, so it scales the new ReplicaSet up to 3. 3. As new pods become available, it scales the old ReplicaSet down and the new one up, and at every moment there are **no more than 9 pods** and **no fewer than 6 available**. 4. The rollout is complete when all 7 replicas belong to the new ReplicaSet and are available, and the old ReplicaSet is at 0 (kept for rollback). The arithmetic counts pods by the ReplicaSets' desired replica counts. A pod that is still terminating has already been subtracted, so the capacity actually running can briefly differ from what the table suggests. ## Capacity: where the surge pods actually land Every surge pod needs somewhere to run. Suppose each transcoding worker requests and is limited to `2662Mi` of memory (about 2.6 GiB), and the 64-node, two-zone cluster is packed with little headroom. The 2 surge pods can then stay `Pending`. The minimum-available rule keeps the old pods in place, so the rollout stalls instead of reducing capacity. The choices are: - **`maxSurge: 0`, `maxUnavailable: 1`**: replace pods one at a time, trading one worker's capacity for needing no spare room; - **keep headroom** or let a node autoscaler add capacity (that mechanism is owned by the scaling tree); - **mind quotas**: a namespace ResourceQuota counts surge pods too, and a rejected pod create shows up as a `ReplicaFailure` condition on the Deployment. ## Edge rules the API enforces - Setting both fields to a literal zero is rejected: validation reports that `maxUnavailable` may not be 0 when `maxSurge` is 0. - If both **resolve** to zero only through rounding, for example `maxSurge: 0` and `maxUnavailable: 25%` on 3 replicas, validation passes, and the controller uses an unavailable budget of **1** so the rollout can still move. - `maxUnavailable` may not exceed 100%. `maxSurge` may. - `maxUnavailable: 100%` is **not** the same as `Recreate`. Recreate scales the old pods to zero and waits until they have stopped running before it creates new ones, whereas RollingUpdate has no such wait.
- When is neither budget setting enough, so you switch the Deployment to the Recreate strategy?Switch when two versions must never run at the same time: they share an exclusive resource, such as a ReadWriteOnce volume, or they speak incompatible message or schema formats. Recreate scales the old ReplicaSet to zero, waits until those pods have stopped running, then creates the new pods, which means a planned outage. You must also remove the rollingUpdate block, because validation forbids it when type is Recreate.
- How does minReadySeconds change what these bounds mean in practice?The availability floor counts available pods, not just Ready ones. With minReadySeconds set to 45, a new worker must stay Ready for 45 seconds before the controller counts it and removes another old pod. A worker that passes readiness and crashes within that window never counts as available, so the rollout does not advance past it. The cost is a slower rollout.
saying these in an interview costs you the question
- Both maxSurge and maxUnavailable percentages round to the nearest whole pod.
- A 25% maxUnavailable on 3 replicas lets one pod go down.
- maxUnavailable: 100% behaves exactly like the Recreate strategy.
- A rolling update never needs spare cluster capacity.
- A pod counts as available the moment its container starts.