How do you watch, pause and roll back an in-flight Kubernetes Deployment rollout from the command line, and what limits how far back a rollback can go?
answer
- revisions = old ReplicaSets at zero
- status blocks + non-zero exit → CI gate
- undo creates a NEW revision, never rewinds
- revisionHistoryLimit default 10; 0 = no rollback
- change-cause annotation; --record is gone
basics
~20 sUse kubectl rollout status to watch, rollout pause/resume to hold it, rollout history to list revisions and rollout undo (optionally --to-revision) to go back. How far back you can go is bounded by revisionHistoryLimit, default 10.
solid answer
~60 sThe command family is `kubectl rollout` against `deployment/<name>`: - **`status`** — blocks until the rollout completes or the progress deadline is exceeded, and returns non-zero on failure, which makes it the right gate in a pipeline (pair it with `--timeout`). - **`history`** — lists revisions; `--revision=N` prints that revision's pod template. - **`undo`** — reverts to the previous revision, or `--to-revision=N` for a specific one. Rollback is itself a normal rolling update and produces a *new* revision number rather than rewinding the counter. - **`pause` / `resume`** — freeze the rollout so you can make several edits and release them as one rollout, or hold a half-updated state while investigating. - **`restart`** — patches a template annotation to force a fresh rollout with the same image. The limit is `revisionHistoryLimit` (default 10): older zero-replica ReplicaSets are garbage-collected, so their revisions can no longer be restored. Setting it to 0 removes rollback entirely. Remember that `undo` restores only the **pod template**. ConfigMaps, Secrets, Services, CRDs and database state are untouched.
code
bash · 13 lineskubectl annotate deployment/web kubernetes.io/change-cause="web 2.4.0 - checkout fix"
kubectl set image deployment/web app=registry.example.com/web:2.4.0
kubectl rollout status deployment/web --timeout=10m || kubectl rollout undo deployment/web
kubectl rollout history deployment/web
kubectl rollout history deployment/web --revision=7
kubectl rollout undo deployment/web --to-revision=6
kubectl rollout pause deployment/web
kubectl set resources deployment/web -c app --limits=memory=1Gi
kubectl rollout resume deployment/web
kubectl rollout restart deployment/web # cycle pods after a Secret changego deeper
Know the four everyday commands — status, history, undo, restart — and that revisionHistoryLimit bounds how far back you can go.
Explain that revisions are ReplicaSets, that undo appends a revision, and how pause/resume batches template edits into one rollout.
Wire rollout status with a timeout into CI as a release gate, keep change-cause meaningful, pin digests so revisions are reproducible, and know undo covers only compute.
Define rollback as a system property: history retention versus deploy frequency, immutable image references, and coordinating config, schema and traffic reverts alongside the pod template.
## Revisions are ReplicaSets Every rollout command is sugar over the ReplicaSet chain. Each revision of the pod template is a ReplicaSet; the current one holds your replicas and the previous ones sit at zero. "Rolling back" means scaling an old ReplicaSet up and the current one down — the same rolling update machinery, in reverse. That is why an undo respects `maxSurge`/`maxUnavailable` and takes about as long as the original rollout, which matters when you are deciding whether to roll back or roll forward under pressure. ## The commands **`kubectl rollout status deployment/web`** blocks and streams progress ("Waiting for deployment \"web\" rollout to finish: 2 of 6 updated replicas are available..."), exiting 0 on success and non-zero when the deployment exceeds `progressDeadlineSeconds`. In CI always add `--timeout=10m` so a stuck rollout fails the job rather than hanging the runner. **`kubectl rollout history deployment/web`** lists revision numbers with a CHANGE-CAUSE column. That column is populated from the `kubernetes.io/change-cause` annotation on the Deployment, which you set yourself (`kubectl annotate deployment/web kubernetes.io/change-cause="image 2.4.0"` or in the manifest). The old `--record` flag that filled it automatically is deprecated and removed in recent kubectl versions, so unless your tooling sets the annotation, expect `<none>` everywhere. `--revision=3` shows that revision's full pod template, which is how you confirm what you are about to restore. **`kubectl rollout undo deployment/web`** reverts to the immediately preceding revision; `--to-revision=3` targets a specific one. Two details matter: the restored template becomes a *new* revision number (so history grows forward, never rewinds), and because the controller recognises a previously used template hash it reuses that old ReplicaSet rather than creating a duplicate. **`kubectl rollout pause` / `resume`** stops the controller from acting on template changes. The classic use is batching: pause, change the image, change resources, change env, resume — one rollout instead of three. The other use is incident containment: pause a rollout that is halfway through so the current mix of old and new pods is frozen while you investigate. Note that a paused Deployment ignores subsequent template edits until resumed, which surprises people who apply a fix and see nothing happen. **`kubectl rollout restart deployment/web`** writes a `kubectl.kubernetes.io/restartedAt` timestamp into the pod template. Since the template changed, a normal controlled rolling update follows — the correct way to cycle pods after a Secret or ConfigMap change, rather than deleting pods by hand. ## What bounds rollback `spec.revisionHistoryLimit` (default 10) caps how many zero-replica ReplicaSets are retained; beyond it, the oldest are garbage-collected and those revisions are gone. Two failure modes: - **Set to 0** — no history at all, `rollout undo` fails. Occasionally chosen to reduce object count in very large clusters; a poor trade for services you may need to revert. - **Churn** — a busy service that deploys many times a day can push a known-good revision out of the window within hours. If your recovery plan is "roll back to Friday's build", the real recovery mechanism must be re-deploying that image tag from your source of truth, not `undo`. ## What undo does not restore This is the highest-value caveat and a frequent follow-up. `undo` restores only the Deployment's pod template. It does **not**: - revert ConfigMap or Secret contents the pods read; - revert Service, Ingress, NetworkPolicy or CRD changes shipped with the release; - undo database migrations or any external side effect; - restore an image tag whose contents were overwritten in the registry — if you deploy mutable tags, rolling back to `v2` may pull today's `v2`, not last week's. Pin digests so a revision is genuinely reproducible. Which is why rollback is a property of your whole delivery system, and `kubectl rollout undo` is only the compute half of it.
- After kubectl rollout undo, why does kubectl rollout history show a higher revision number rather than returning to the old one?Because undo is implemented as a normal forward rollout that reapplies an older pod template. The revision counter only ever increases; the restored template simply becomes the newest revision, and the controller reuses the existing ReplicaSet whose template hash matches. History is an append-only log of what was deployed, not a pointer you rewind.
- Your team sets revisionHistoryLimit to 0 to reduce clutter. What capability do you lose?All rollback via kubectl rollout undo, because the previous ReplicaSets are deleted as soon as they reach zero replicas and there is no stored template to restore. Recovery then depends entirely on re-deploying a known-good manifest and image from source control or your CD system, which is slower and assumes the pipeline is healthy at the moment you need it.
saying these in an interview costs you the question
- Thinking rollout undo rewinds the revision counter instead of appending a new revision.
- Expecting the CHANGE-CAUSE column to populate itself; it comes from the kubernetes.io/change-cause annotation now that --record is gone.
- Believing rollout undo also reverts ConfigMaps, Secrets, Services or database migrations.
- Deleting pods manually to restart a service instead of using kubectl rollout restart.
- Applying a fix to a paused Deployment and expecting it to roll out before resume.