A CI job gates a Kubernetes transcoding Deployment's release on `kubectl rollout status`, and workers take about 7 minutes to turn Ready. How do you set progressDeadlineSeconds, minReadySeconds and --timeout so the gate fails correctly?
answer
- longest stall, not total time
- four kinds of progress
- ready for N seconds first
- timeout zero waits forever
- reports failure, never rolls back
basics
~20 sprogressDeadlineSeconds measures the longest gap between progress events, not the whole rollout, so it must exceed one pod's time to Ready plus margin. minReadySeconds catches early crashes, and a --timeout above the total rollout time is the backstop.
solid answer
~50 sThe Deployment controller updates the `Progressing` condition's `lastUpdateTime` whenever it sees progress: new pods created, old pods scaled down, or the ready or available count rising. It sets reason `ProgressDeadlineExceeded` only when **no** progress has happened for `progressDeadlineSeconds` (default 600). A worker that needs about 430 s to become Ready leaves thin margin under 600, so I'd set 900. `minReadySeconds`, say 45, means a worker must stay Ready that long to count as available, which catches workers that crash just after passing readiness. It must be lower than the deadline. `kubectl rollout status` exits non-zero with `exceeded its progress deadline`, but its `--timeout` defaults to 0, meaning wait forever. I'd set it above the whole rollout, about 24 minutes for 7 workers, for example `--timeout=40m`. Neither the command nor the controller rolls back, so the pipeline must run `kubectl rollout undo` itself.
code
bash · 8 linesset -euo pipefail
kubectl set image deployment/transcoder worker=registry.example.com/transcoder:4.12.3 -n media
if ! kubectl rollout status deployment/transcoder -n media --timeout=40m; then
kubectl describe deployment transcoder -n media
kubectl rollout undo deployment/transcoder -n media
kubectl rollout status deployment/transcoder -n media --timeout=40m
exit 1
figo deeper
Know that kubectl rollout status waits for a Deployment rollout to finish and exits with an error when the rollout exceeds its progress deadline.
Explain what counts as progress, that the deadline measures the longest stall, and how minReadySeconds changes when a pod counts as available.
Size the deadline from real startup times, add a --timeout backstop, pin the revision, and make the pipeline run the undo, since nothing rolls back on its own.
Decide which failures should be caught by the rollout gate and which by later health signals, and set organisation-wide deadline defaults for slow-starting workloads.
## What counts as progress The Deployment controller keeps a **`Progressing`** condition in `status.conditions`. Its `lastUpdateTime` is refreshed whenever the controller sees progress: - the new ReplicaSet has more pods (new pods were created); - the old ReplicaSets have fewer pods; - the number of **ready** pods rose; - the number of **available** pods rose. `progressDeadlineSeconds` (default **600**) is compared with **the time since the last progress**, not the time since the rollout started. It limits the longest stall, and a rollout can legitimately take far longer in total. | Progressing reason | Status | Meaning | |---|---|---| | `ReplicaSetUpdated` | True | Progress seen recently | | `NewReplicaSetAvailable` | True | Rollout complete | | `ProgressDeadlineExceeded` | False | No progress for `progressDeadlineSeconds` | | `DeploymentPaused` | Unknown | Paused; progress is not estimated | When the deadline passes, the controller **only reports** it. It keeps reconciling, it does **not** roll back, and the partly finished state stays as it is until someone acts. ## Sizing the three knobs for slow workers Take 7 video-transcoding workers with the default surge 2 / unavailable 1 and a startup time of about 430 s (model load plus codec warm-up). 1. **`minReadySeconds: 45`.** A worker must stay Ready for 45 s before it counts as available. A worker that passes its readiness probe and is then OOM-killed at its `2662Mi` limit on its first real job, within those 45 s, never counts as available, so the rollout does not move past it. The API requires `progressDeadlineSeconds` to be **greater** than `minReadySeconds`. 2. **`progressDeadlineSeconds: 900`.** The longest normal gap is creation to Ready, about 430 s, followed by 45 s to available. That fits under 600, but an image pull onto a fresh node from the 64-node pool can use up the remaining margin and fail a healthy release. 900 s leaves room, and a pod that never becomes Ready still fails the gate within about 15 minutes. 3. **`kubectl rollout status --timeout=40m`.** The rollout moves in waves: about 3 new pods, then 6, then 7, and each wave waits roughly 475 s. That is about 24 minutes in total. The timeout is a backstop above that figure, not a copy of the deadline. ## What kubectl rollout status actually checks On each update it receives, the command: 1. waits until `status.observedGeneration` has caught up with `metadata.generation` (`Waiting for deployment spec update to be observed...`); 2. **fails** with `deployment "transcoder" exceeded its progress deadline` if the Progressing reason is `ProgressDeadlineExceeded`; 3. otherwise reports the updated, terminating and available counts until all replicas are updated and available, then prints `successfully rolled out` and exits 0. Points that matter for a gate: - `--timeout` defaults to **0, meaning never**. Without it, a stall that the deadline does not catch leaves the CI job hanging until the runner kills it. - `--revision=N` pins the watch to one revision and fails if another rollout replaces it, so a concurrent deploy cannot make your gate pass. - Readiness **flapping** can register as progress. A crash-looping worker that is briefly Ready on each restart raises the ready count, which can push the deadline back. That is another reason to keep the timeout backstop. ## What the gate should do on failure If a bad image never pulls, the new pods sit in `ImagePullBackOff`, no ready or available count rises, and about 900 s after the last progress the condition changes and the command exits 1. Meanwhile, one old worker has already been removed (unavailable 1), so the pool is running at 6 of 7. The pipeline should: - run `kubectl rollout undo deployment/transcoder` and then gate on `rollout status` again; - keep the failing ReplicaSet's events and pod descriptions for the post-mortem; - page someone if the undo also stalls, because at that point capacity is below target. A release tool's wait flags follow the same principle, but they belong to that tool's own tree.
- Why must progressDeadlineSeconds be larger than minReadySeconds on a Deployment?After a pod becomes Ready, the next progress event is that pod becoming available, which takes minReadySeconds. If the deadline were shorter, every healthy rollout would pass its deadline while waiting out that window. The API server rejects the combination with 'must be greater than minReadySeconds'.
- What happens to the progress deadline if someone pauses the Deployment during the release?While spec.paused is true, the controller sets the Progressing condition to Unknown with reason DeploymentPaused and does not estimate progress, so the deadline cannot fire. On resume the condition is refreshed and the countdown starts again. A CI gate watching the rollout will simply wait, so the --timeout backstop is what ends a job that someone forgot to resume.
saying these in an interview costs you the question
- progressDeadlineSeconds is the maximum total duration of a rollout.
- ProgressDeadlineExceeded makes the Deployment controller roll back automatically.
- kubectl rollout status gives up after ten minutes by default.
- minReadySeconds delays the start of the readiness probe.
- Setting the deadline to the startup time exactly is safe.