skip to content

A Kubernetes Deployment rollout has been stuck partway for ten minutes: some new pods are up, the old ones are still serving. How do you diagnose it, and what does the ProgressDeadlineExceeded condition on a Deployment mean?

level: seniorimportance: should knowfreq 55%

answer

  1. stuck = safe: maxUnavailable protects old pods
  2. chain: Deployment conditions → RS → pod → logs
  3. FailedCreate on RS = quota / admission / webhook
  4. ProgressDeadlineExceeded = status only, no auto-rollback
  5. PDBs block evictions, not rollouts

basics

~20 s

Look at the Deployment's conditions and events, then at the new ReplicaSet and its pods. ProgressDeadlineExceeded means the rollout made no progress within progressDeadlineSeconds (default 600); it is a status marker only — Kubernetes does not roll back for you.

solid answer

~60 s

Work down the ownership chain. 1. `kubectl rollout status` reports what it is waiting for; `kubectl describe deployment` shows the `Progressing` and `Available` conditions and the scale events. 2. `kubectl get rs` shows which generation is stuck: new ReplicaSet DESIRED > READY means the new pods are the problem; DESIRED not increasing at all means the controller is blocked from creating them. 3. `kubectl describe pod` on a new pod gives the verdict — `ImagePullBackOff` (bad tag or missing pull secret), `CrashLoopBackOff` (the app dies on start), `Pending` with a scheduling message (no capacity, untolerated taint, unsatisfiable affinity, unbound PVC), or Running-but-not-Ready (readiness never passes, often a missing dependency or a probe pointing at the wrong path). 4. `kubectl get events --sort-by=.lastTimestamp` and `kubectl logs --previous` fill in the rest. Also check ResourceQuota, which blocks pod creation with a clear event. `progressDeadlineSeconds` (default 600) sets when the controller stamps `Progressing=False, reason=ProgressDeadlineExceeded`. Nothing is rolled back automatically — old pods keep serving, which is the safe outcome; recovery is your `rollout undo` or roll-forward.

code

bash · 10 lines
bash
kubectl rollout status deployment/web --timeout=30s
kubectl describe deployment web | sed -n '/Conditions/,/Events/p'

kubectl get rs -l app=web
# web-6c8d9  DESIRED 3  CURRENT 3  READY 0   <- new gen unhealthy
# web-5b7f2  DESIRED 5  CURRENT 5  READY 5

kubectl describe pod -l pod-template-hash=6c8d9 | tail -30
kubectl logs -l pod-template-hash=6c8d9 --previous --tail=50
kubectl get events --sort-by=.lastTimestamp | tail -20

go deeper

for a junior

Know where to look — describe the deployment and the pods — and recognise ImagePullBackOff, CrashLoopBackOff and Pending as the usual culprits.

for a middle

Walk the chain deliberately, read the Progressing condition, and explain progressDeadlineSeconds and why the old pods keep serving.

for a senior

Separate creation failures from readiness failures, spot quota/admission/webhook causes, know PDBs are irrelevant here, and decide rollback versus roll-forward on recovery time.

for a principal

Make failure observable and bounded by policy: progress deadlines aligned to pipeline timeouts, automated undo in CD, capacity headroom so zero-unavailability strategies cannot deadlock, and alerting on Progressing=False.

## Why a stuck rollout is the safe failure A rolling update advances only while pods become **available**. If new pods never become available, the controller refuses to remove more old ones, because doing so would breach `maxUnavailable`. The result is a half-updated Deployment serving mostly from the old generation. That is a feature: a broken release degrades to "no release" rather than an outage. Say this first — it shows you understand the mechanism rather than just the commands. ## The diagnostic ladder Always move down the ownership chain: Deployment → ReplicaSet → Pod → container. **Deployment level.** `kubectl describe deployment web` shows two conditions. `Available` reflects whether enough replicas are up right now. `Progressing` narrates the rollout; when it flips to `False` with reason `ProgressDeadlineExceeded`, the controller has given up advancing. The events section lists each "Scaled up replica set …" step, and where it stops tells you how far the rollout got. **ReplicaSet level.** `kubectl get rs -l app=web` distinguishes two very different situations: - New RS shows DESIRED 3 / READY 0 → pods are being created but not becoming ready. The problem is inside the pod. - New RS shows DESIRED 1 and never grows, with events on the ReplicaSet like `FailedCreate` → pod creation itself is being refused. Causes: a `ResourceQuota` exhausted, a `LimitRange` rejecting the spec, a missing ServiceAccount, or admission control (Pod Security Admission rejecting a privileged spec, or a validating webhook that is down — a webhook whose backing service is unhealthy blocks pod creation cluster-wide and is a classic mystery). **Pod level.** `kubectl describe pod` names the cause almost verbatim: - `Pending` with `FailedScheduling` — insufficient CPU/memory, untolerated taint, unsatisfiable node affinity or topology spread, an unbound PVC, or (with `maxUnavailable: 0` and a full cluster) simply no room for the surge pod. This last one is a real trap: a zero-unavailability strategy on a cluster with no headroom deadlocks, because the controller may not free capacity by deleting an old pod first. - `ImagePullBackOff` / `ErrImagePull` — wrong tag, wrong registry, missing `imagePullSecrets`, or an architecture mismatch on a mixed fleet. - `CrashLoopBackOff` — read `kubectl logs <pod> --previous` and check the exit code; `137` is OOM-kill (limit too low for the new version), a config parse error is the other frequent cause. - `Running` but `0/1 READY` — readiness never passes. Either the app genuinely cannot serve (a dependency it waits on is unreachable, a migration is blocking startup) or the probe is misconfigured for the new version — changed port, changed path, or a timeout shorter than real startup. ## ProgressDeadlineExceeded precisely `spec.progressDeadlineSeconds` (default 600) is the time the Deployment may go without *progress* — progress meaning a change in the number of available updated replicas. When exceeded, the controller sets `Progressing=False` with reason `ProgressDeadlineExceeded`. It does **not** stop trying, does not delete pods, and does not roll back; Kubernetes has no automatic rollback for Deployments. Its concrete value is that `kubectl rollout status` then exits non-zero, so a CD pipeline can detect failure and call `kubectl rollout undo` itself. Set it shorter than your pipeline timeout so failure is reported rather than hanging, and longer than the honest worst-case startup of your slowest pod, or healthy releases will be flagged as failures. ## Two things people get wrong **PodDisruptionBudgets do not block a Deployment rollout.** A rolling update deletes pods directly through the API; PDBs constrain the *eviction* API used by drains and the descheduler. A PDB that is impossible to satisfy will block node drains and cluster autoscaler scale-down, not your release. Getting this right is a strong senior signal. **Deleting the stuck pods rarely helps.** The ReplicaSet recreates them with the same broken spec. Fix the template, or undo. ## Recovery Decide between rolling back (`kubectl rollout undo`, fastest, restores the last known-good template) and rolling forward (when the bug is in configuration or an external dependency and undo would not fix it). Either way, note that undo is a rolling update with the same knobs, so recovery time equals rollout time. If you are mid-incident and the old pods are still serving fine, you already have the safe state — pause the rollout, then decide deliberately.

  • Does Kubernetes automatically roll back a Deployment when ProgressDeadlineExceeded is set?
    No. The condition is purely informational: the controller marks Progressing=False and keeps trying, while the old pods continue serving. Its practical value is that kubectl rollout status then exits non-zero, so your CD pipeline can detect the failure and run kubectl rollout undo explicitly. Automatic, metric-driven rollback requires progressive-delivery tooling layered on top.
  • A rollout with maxUnavailable: 0 is stuck with the new pod Pending and FailedScheduling on a full cluster. Why can it never resolve itself?
    With maxUnavailable at zero, the controller is not permitted to delete an old pod to make room, and with no free capacity the surge pod cannot be scheduled — a deadlock. It resolves only when capacity appears, whether from cluster autoscaler scale-up, another workload releasing resources, or you temporarily allowing some unavailability.

saying these in an interview costs you the question

  • Claiming Kubernetes rolls a Deployment back automatically when the progress deadline is exceeded.
  • Blaming a PodDisruptionBudget for a blocked rollout — PDBs gate evictions, not the Deployment controller's pod deletions.
  • Deleting the failing pods repeatedly instead of fixing or reverting the pod template.
  • Ignoring the ReplicaSet layer and so missing FailedCreate causes such as quota or a failing admission webhook.
  • Not distinguishing 'pods not created' from 'pods created but never ready', which have completely different causes.

context