skip to content

What is a Kubernetes PodDisruptionBudget, and which kinds of pod loss does it actually protect against?

level: middleimportance: must knowfreq 60%

answer

  1. policy/v1, selector + minAvailable | maxUnavailable
  2. enforced at pods/eviction, 429 on violation
  3. voluntary only — drain, autoscaler, upgrades
  4. involuntary: node crash, node-pressure, preemption, delete pod
  5. healthy = Ready; disruptionsAllowed is the live number

basics

~20 s

A PodDisruptionBudget declares the minimum availability a set of pods must keep during voluntary disruptions — evictions from node drains, cluster-autoscaler scale-down, upgrades. The API server rejects an eviction that would break it. It cannot stop involuntary loss: node crashes, kernel panics, kubelet node-pressure eviction, or a direct pod delete.

solid answer

~50 s

A **PodDisruptionBudget (PDB)** is a `policy/v1` object with a label selector and either `minAvailable` or `maxUnavailable`. It is a constraint enforced at the **Eviction API**: when a client calls `pods/eviction`, the disruption controller checks whether removing that pod would violate the budget and returns **HTTP 429 TooManyRequests** if it would. That scope is the whole point. Disruptions split into two classes: - **Voluntary** — deliberate actions that go through eviction: `kubectl drain` for a node upgrade, cluster-autoscaler consolidating nodes, node-pool rotation, some operators. PDBs govern these. - **Involuntary** — node hardware failure, kernel panic, VM termination, out-of-resource kubelet eviction, and higher-priority preemption. A PDB has **no effect** on any of them; you survive those with replica count, spread across failure domains, and health checking. A PDB also does not stop `kubectl delete pod`, which bypasses the eviction subresource entirely. Its status fields — `disruptionsAllowed`, `currentHealthy`, `desiredHealthy`, `expectedPods` — are what you actually read during an incident.

code

yaml · 11 lines
yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: api
  namespace: prod
spec:
  minAvailable: 80%
  unhealthyPodEvictionPolicy: AlwaysAllow
  selector:
    matchLabels:
      app: api

go deeper

for a junior

Say what the object declares and that it applies to planned maintenance, not to crashes.

for a middle

Draw the voluntary/involuntary line explicitly, name the Eviction API as the enforcement point, and know that healthy means Ready.

for a senior

Add the failure modes — selectors spanning workloads, unhealthy pods zeroing the budget, unhealthyPodEvictionPolicy — and treat PDBs as a contract with the platform team.

for a principal

Position PDBs as the negotiated interface between application availability targets and the cluster's ability to patch itself, and set organizational defaults accordingly.

## The problem it solves Cluster maintenance means pods move. A node gets a new kernel, a node pool rolls to a new image, the cluster autoscaler consolidates two half-empty nodes into one. Without coordination, an automated drain can take out every replica of a service at once — each individual eviction looks harmless, and the aggregate is an outage. A PodDisruptionBudget is how a workload owner tells the cluster "you may take pods from me, but not below this line". ## Voluntary versus involuntary This distinction is the single idea an interviewer is testing. **Voluntary disruptions** are initiated by an actor that *asks* the cluster politely, through the Eviction API: - `kubectl drain` before a node upgrade or repair - cluster-autoscaler removing an underused node - node-pool upgrades in managed Kubernetes - descheduler-style rebalancing tools **Involuntary disruptions** are events nobody asked for and nothing can veto: - node hardware failure, kernel panic, network partition - the cloud provider reclaiming a spot/preemptible instance - **kubelet node-pressure eviction** under memory or disk pressure — the kubelet kills pods locally and never consults the API server's disruption controller - **scheduler preemption** of a lower-priority pod to make room for a higher-priority one - someone running `kubectl delete pod`, which deletes directly rather than evicting PDBs bind only the first list. Saying "a PDB protects my service from node failure" is the classic wrong answer; the protection against node failure is having enough replicas spread across enough nodes and zones. ## The object ```yaml apiVersion: policy/v1 kind: PodDisruptionBudget metadata: name: api spec: minAvailable: 80% selector: matchLabels: app: api ``` Key points about the shape: - The **selector is independent of any controller**. A PDB is not "attached" to a Deployment; it matches pods by label. Two Deployments sharing a label are covered by one budget, which is usually a bug — and a pod matched by two PDBs cannot be evicted at all, because that is treated as an error condition. - Exactly one of `minAvailable` / `maxUnavailable` may be set. - `spec.unhealthyPodEvictionPolicy` (GA in Kubernetes 1.31) chooses whether pods that are *running but not Ready* may be evicted even when the budget is exhausted: `IfHealthyBudget` (the default, conservative) or `AlwaysAllow`, which prevents the deadlock where broken pods can never be cleared off a node. ## "Healthy" means Ready The controller counts a pod toward `currentHealthy` only when its `Ready` condition is true. That has a practical consequence: a workload whose pods are failing readiness has `currentHealthy` below `desiredHealthy`, so `disruptionsAllowed` is 0 and every eviction is refused. A broken application therefore blocks node maintenance — often the real story behind a drain that will not finish. ## Reading its state ``` kubectl get pdb api NAME MIN AVAILABLE MAX UNAVAILABLE ALLOWED DISRUPTIONS AGE api 80% N/A 2 31d ``` `ALLOWED DISRUPTIONS` (`status.disruptionsAllowed`) is the live budget: how many pods may be evicted right now. Zero means the next drain will stall. `status.expectedPods` comes from the owning controller's replica count via its `scale` subresource, which is why a PDB over bare pods with a percentage target may be unable to compute anything. ## What a good answer includes - PDBs are enforced at eviction time by the API server, not by the scheduler and not by kubelet. - They express *availability during maintenance*, not durability. - They are a contract in both directions: too loose and upgrades cause outages; too strict and the platform team cannot patch nodes. A PDB that permanently allows zero disruptions is not "safe", it is an operational blocker that someone will eventually delete under time pressure. - Every replicated, availability-sensitive workload should have one; single-replica workloads need a deliberate decision, because `minAvailable: 1` on one replica blocks drains forever.

  • Does a PodDisruptionBudget protect a service when a node's hardware fails?
    No. That is an involuntary disruption — the pods are simply gone, with no eviction request for the disruption controller to reject. Surviving node failure comes from having enough replicas, spreading them across nodes and zones, and letting the controller reschedule. PDBs only constrain deliberate, API-mediated pod removals.
  • Why can `kubectl delete pod` remove a pod that a PDB would have protected?
    Because deletion targets the pod resource directly and never touches the `pods/eviction` subresource where the disruption controller enforces the budget. Only clients that use the Eviction API — `kubectl drain`, the cluster autoscaler, well-behaved operators — are subject to a PDB. This is why PDBs are a cooperation mechanism, not a security control.

A staffing rule that says at least four nurses must be on the ward: it governs who may be granted leave, and it says nothing about how many can call in sick.

saying these in an interview costs you the question

  • Claiming PDBs protect against node crashes or spot-instance reclamation
  • Thinking a PDB is bound to a Deployment rather than matching pods by label selector
  • Assuming kubelet node-pressure eviction and scheduler preemption honor PDBs — they do not
  • Believing a PDB prevents kubectl delete pod
  • Not knowing that a pod only counts as available when it is Ready

context