A Kubernetes PodDisruptionBudget accepts either minAvailable or maxUnavailable. How do the two differ in practice, how are percentages resolved, and when does each choice go wrong?
answer
- exactly one of minAvailable / maxUnavailable
- desiredHealthy = min, or expectedPods − maxUnavailable
- percentages resolve against expectedPods and round up
- maxUnavailable tracks a shrinking fleet; integer minAvailable does not
- maxUnavailable: 0 and minAvailable: 1-of-1 block drains forever
basics
~20 sminAvailable sets an absolute floor of healthy pods; maxUnavailable sets how many may be missing relative to the expected replica count. Percentages are resolved against the controller's replica count and round up. minAvailable as an integer is dangerous when replicas shrink; maxUnavailable: 0 and minAvailable equal to the replica count both block every drain.
solid answer
~60 sBoth express the same constraint from opposite ends, and exactly one may be set. - **`minAvailable: 3`** — at least 3 pods must stay healthy. The floor is fixed regardless of how many replicas exist. - **`maxUnavailable: 1`** — at most 1 may be down, computed as `expectedPods − maxUnavailable`. The floor moves with the fleet. Percentages resolve against `status.expectedPods`, taken from the owning controller's `scale` subresource, and **round up**. `minAvailable: 50%` of 5 replicas requires 3 healthy. Because `maxUnavailable` also rounds up, a small percentage on a small fleet can still permit a disruption where the integer form would not — so state the percentage in terms of what it implies for your smallest expected replica count. Rules of thumb: prefer **`maxUnavailable`** for horizontally scaled stateless services — it stays correct when an HPA scales the Deployment down, whereas `minAvailable: 5` on a fleet that shrinks to 5 silently blocks all maintenance. Prefer **`minAvailable`** for quorum systems, where "at least 2 of 3 etcd members" is the literal requirement. Anti-patterns: `maxUnavailable: 0`, and `minAvailable: 1` on a single-replica Deployment — both mean drains never complete.
code
yaml · 19 linesapiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: web
spec:
maxUnavailable: 1
selector:
matchLabels:
app: web
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: etcd
spec:
minAvailable: 2
selector:
matchLabels:
app: etcdgo deeper
Explain the two fields in plain terms and that only one may be set.
Give the desiredHealthy arithmetic, percentage rounding, and why maxUnavailable survives replica-count changes.
Choose per workload class — maxUnavailable for autoscaled stateless, minAvailable for quorum systems — and call out the never-satisfiable anti-patterns.
Set a fleet-wide convention and admission checks so budgets stay satisfiable, and treat 'needs every replica at peak' as a capacity problem rather than a PDB problem.
## Two spellings of one constraint A PodDisruptionBudget must set exactly one of: - `minAvailable` — the number (or percentage) of pods matching the selector that must remain **healthy**. - `maxUnavailable` — how many may be unhealthy or absent at once. Internally the controller reduces both to `desiredHealthy` and compares it to `currentHealthy`: ``` minAvailable form: desiredHealthy = minAvailable maxUnavailable form: desiredHealthy = expectedPods − maxUnavailable disruptionsAllowed = currentHealthy − desiredHealthy (floored at 0) ``` `expectedPods` comes from the pods' owning controller via its `scale` subresource. This is why `maxUnavailable` (and percentage forms generally) require the pods to be owned by something scalable — a Deployment/ReplicaSet/StatefulSet. A PDB selecting bare pods can work with an integer `minAvailable` but cannot compute a percentage-of-replicas target. ## The behavioral difference The distinction that matters is what happens **when the replica count changes**. Suppose a Deployment runs 10 replicas. - With `minAvailable: 8`, two pods may be disrupted. If an HPA scales the Deployment down to 8 overnight, `currentHealthy` is 8 and `desiredHealthy` is 8 — `disruptionsAllowed` becomes **0**. Node maintenance silently stops working at exactly the time of day when maintenance is cheapest. Nothing in the PDB looks wrong; the number simply no longer matches reality. - With `maxUnavailable: 2`, `desiredHealthy` is always `expectedPods − 2`. At 10 replicas it allows 2 disruptions; at 8 replicas it still allows 2. The budget tracks the fleet. This is the main reason `maxUnavailable` is the better default for autoscaled stateless services. Conversely, quorum-based systems have an absolute requirement: a 3-node etcd or ZooKeeper ensemble needs 2 members alive, full stop. `minAvailable: 2` states exactly that and does not drift if someone scales the StatefulSet. Use the form that matches the actual invariant. ## Percentages and rounding Both forms accept a string percentage, resolved against `expectedPods` and **rounded up**: - `minAvailable: 50%` of 5 → `ceil(2.5)` = 3 must stay healthy → 2 disruptions allowed. - `maxUnavailable: 10%` of 15 → `ceil(1.5)` = 2 may be down. Rounding up makes `minAvailable` percentages stricter and `maxUnavailable` percentages more permissive. Always sanity-check a percentage at the *smallest* replica count the workload can reach, because that is where it behaves most surprisingly — `maxUnavailable: 10%` on a 2-replica Deployment still resolves to 1, i.e. half the fleet. ## Anti-patterns **`maxUnavailable: 0`.** Reads like maximum safety; means `desiredHealthy == expectedPods`, so no pod may ever be evicted. Every drain hangs, node upgrades stall cluster-wide, and eventually an operator deletes the PDB during a maintenance window — leaving the workload with *no* protection at all. If a workload genuinely cannot tolerate any pod loss, the honest answer is that it cannot run on a cluster that patches its nodes; fix the application, not the budget. **`minAvailable: 1` with one replica.** Same effect: the single pod can never be evicted. Either accept brief downtime (`maxUnavailable: 1`, or no PDB at all), or run two replicas. **`minAvailable: 100%`.** Identical to `maxUnavailable: 0`. **A selector broader than intended.** PDB selectors match labels, not controllers. `app: web` matching two Deployments creates one shared budget whose arithmetic nobody reasons about correctly; worse, a pod matched by *two* PDBs cannot be evicted at all. Scope selectors as tightly as the controller's own selector. **Setting `minAvailable` to the replica count** to "guarantee availability" — this is the same trap dressed differently and is a common review finding. ## Choosing a number Start from the capacity question: how many replicas does the service need to carry peak load with acceptable latency? If 10 replicas serve peak and 8 suffice, `maxUnavailable: 2` is defensible. If every replica is needed at peak, the budget is not the problem — the fleet is undersized for maintenance, and the answer is more replicas, not a stricter PDB. A budget that is never satisfiable is worse than none, because it converts a planned rolling upgrade into a stuck one.
- An HPA can scale a Deployment between 4 and 20 replicas. Why is `minAvailable: 10` a bad budget for it?Whenever the HPA scales below or to 10 replicas, currentHealthy equals or falls under desiredHealthy, so disruptionsAllowed drops to zero and no node holding those pods can be drained. The budget silently becomes an outright block during low-traffic periods. `maxUnavailable: 2` (or a percentage) stays meaningful across the whole 4–20 range.
- A team proposes `maxUnavailable: 0` for a payment service. How do you respond?It forbids every eviction, so node upgrades and autoscaler consolidation stall indefinitely on any node running that service, and the usual outcome is someone deleting the PDB mid-incident. If the service truly cannot lose a single pod, the real fix is graceful shutdown with connection draining plus enough replicas that losing one is invisible — then a normal budget such as `maxUnavailable: 1` is safe.
saying these in an interview costs you the question
- Setting both minAvailable and maxUnavailable in one PDB
- Treating maxUnavailable: 0 as a best practice for critical services
- Using an integer minAvailable equal to the current replica count on an autoscaled workload
- Assuming percentages round down
- Writing a selector broader than the owning controller's, so one budget spans several workloads