Explain what maxSurge and maxUnavailable do in a Kubernetes Deployment's RollingUpdate strategy, what their default values are, and when you would choose the Recreate strategy instead.
answer
- surge = above desired; unavailable = below desired
- both default 25%, surge rounds up, unavailable rounds down
- both zero = illegal
- available = ready + minReadySeconds
- Recreate for singletons and RWO volumes
basics
~20 smaxSurge is how many pods may exist above the desired count during a rollout; maxUnavailable is how many of the desired pods may be unavailable. Both default to 25%. Recreate deletes all old pods before creating new ones, accepting downtime.
solid answer
~60 sBoth are bounds evaluated against `spec.replicas`, and both accept an absolute number or a percentage (default `25%` each). - **`maxSurge`** — the extra capacity allowed above the desired count. With 4 replicas and 25%, at most 5 pods may exist at once. Percentages round **up**. - **`maxUnavailable`** — how many of the desired replicas may be unavailable at once. With 4 and 25%, at least 3 must stay available. Percentages round **down**, and both cannot be zero simultaneously — that would make progress impossible. "Available" means ready plus ready for `minReadySeconds`, so readiness quality drives rollout correctness. Tuning: `maxUnavailable: 0, maxSurge: 1` never loses capacity but needs headroom for one extra pod and rolls slowly. `maxUnavailable: 25%, maxSurge: 0` fits a fixed resource budget but runs degraded. **Recreate** scales the old ReplicaSet to zero, waits for termination, then scales the new one up. It guarantees no two versions run simultaneously — the right choice for singletons, an RWO volume that only one pod can mount, or a schema change two versions cannot share — at the cost of a real outage window.
code
yaml · 21 linesapiVersion: apps/v1
kind: Deployment
metadata:
name: web
spec:
replicas: 6
minReadySeconds: 15
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 2
selector:
matchLabels: { app: web }
template:
metadata:
labels: { app: web }
spec:
containers:
- name: app
image: registry.example.com/web:2.4.0go deeper
State the two definitions correctly, know the 25% defaults, and know Recreate causes downtime.
Add rounding rules, the illegal both-zero case, the availability definition including minReadySeconds, and concrete tuning pairs.
Reason about capacity headroom, downstream connection limits during surge, RWO volume constraints, and how readiness quality determines whether the knobs mean anything.
Set organisation-wide defaults per workload class, budget cluster headroom for surge across many services, and decide where built-in strategies must be replaced by canary or blue-green delivery.
## The two bounds A rolling update is a constrained search: the Deployment controller repeatedly asks "may I add a pod to the new ReplicaSet?" and "may I remove one from the old?", answering with two ceilings measured against `spec.replicas`. **`maxSurge`** caps total pods above desired. `replicas: 10, maxSurge: 25%` allows up to 13 pods (12.5 rounded **up**) to exist at once. Surge is what lets a rollout add capacity before removing any, so throughput never drops. **`maxUnavailable`** caps how many desired replicas may be unavailable. `replicas: 10, maxUnavailable: 25%` means at least 8 must be available at all times (2.5 rounded **down** to 2). The asymmetric rounding is deliberate: both rules err toward more capacity. They cannot both be 0 — with no surge headroom and no tolerance for unavailability there is no legal move, and the API rejects it. ## What "available" means Availability, not existence, drives the loop. A pod counts when it is **Ready** and has been Ready for `minReadySeconds` (default 0). Consequences worth stating in an interview: - Without a readiness probe, a pod counts as available as soon as its container starts, so a slow-starting service can have its old pods removed while the new ones cannot serve — a self-inflicted outage that looks like a Kubernetes bug. - `minReadySeconds` guards against processes that pass readiness and then crash seconds later; the rollout stalls instead of sweeping the fleet. - Removing an old pod also requires graceful termination to avoid dropping in-flight requests; the rolling strategy governs *counts*, not connection draining. ## Choosing values The decision is capacity headroom versus degradation tolerance: - **`maxUnavailable: 0, maxSurge: 1`** — the conservative production default for a service with a tight capacity budget per pod. You never serve with fewer than N ready pods; you need room in the cluster for one extra pod, and the rollout is serial and slow. - **`maxUnavailable: 0, maxSurge: 25%`** — same guarantee, faster, at the cost of temporarily running ~125% of the pods (and their connection pools, licences, and database connections — a real constraint when a downstream limits connections). - **`maxSurge: 0, maxUnavailable: 1`** — for workloads that cannot exceed a fixed pod count: constrained node pools, a licence limit, or a fixed shard count. You run degraded during the rollout, so N must have enough margin. - **Aggressive percentages (50%+)** — fine for large stateless fleets where speed matters and capacity is elastic. Remember both are evaluated per Deployment, not per zone or node, and neither is aware of topology; spreading is a scheduling concern (topology spread constraints), not a rollout knob. ## Recreate `strategy.type: Recreate` ignores both fields. The controller scales the old ReplicaSet to 0, waits for every pod to terminate, then scales the new one up. There is a genuine outage between those steps, proportional to termination plus startup time. You choose it when two versions must never run simultaneously: - **Singleton processes** — a leader-less scheduler or migration runner where two instances would double-execute work. - **ReadWriteOnce volumes** — a PVC that only one node can mount; a surged pod on a different node would be stuck `Pending` forever, silently hanging the rollout. This is the most common accidental discovery of Recreate's necessity. - **Incompatible schema or protocol versions** — when the new version writes a format the old cannot read and you have not done an expand/contract migration. - **Non-production environments** — where the simplicity and lower resource use beat availability. A good answer names the RWO-volume case explicitly, because it turns Recreate from a nostalgic default into an engineering decision. ## Verifying During a rollout, `kubectl get rs` shows both generations with their DESIRED/CURRENT/READY counts, and their sum against `replicas` demonstrates the two bounds in action. `kubectl describe deployment` prints `RollingUpdateStrategy: 25% max unavailable, 25% max surge` and events narrating each scale step, which is the quickest way to confirm the strategy actually in force versus the one you believe you applied.
- What happens if you set both maxSurge and maxUnavailable to 0?The API server rejects the Deployment. With no allowance to create an extra pod and no allowance for any desired pod to be unavailable, the controller has no legal move and the rollout could never progress, so the combination is invalid by validation rather than a runtime hang.
- A Deployment mounts a ReadWriteOnce PersistentVolumeClaim and uses the default rolling update. What goes wrong?The rollout surges a new pod before removing the old one, and if that pod is scheduled to a different node it cannot attach the RWO volume, so it stays Pending or stuck in ContainerCreating while the old pod holds the mount. The rollout stalls until the progress deadline. Recreate — or a StatefulSet with the appropriate semantics — is the correct choice for exclusive volumes.
saying these in an interview costs you the question
- Swapping the two definitions — describing maxSurge as pods allowed to be down.
- Believing maxUnavailable counts pods that are not Running, rather than pods that are not Available (ready plus minReadySeconds).
- Claiming zero-downtime is guaranteed by maxUnavailable: 0 alone, with no readiness probe or graceful shutdown.
- Assuming rolling updates work with ReadWriteOnce volumes.
- Not knowing both fields default to 25%, or asserting they are absolute counts only.