skip to content

Rollouts and Rollback

Driving a Deployment update: RollingUpdate versus Recreate, what surge and unavailable budgets do to capacity mid-deploy, and when progressDeadlineSeconds calls the rollout failed. 'Your rollout is stuck at 3/5, what do you check?' is the standard follow-up.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

4

In a Kubernetes Deployment, which changes start a new rollout, and how does `kubectl rollout restart` trigger one?

level: juniorimportance: must knowfreq 68%

answer

  1. only one part of the spec counts
  2. hash label on the ReplicaSet
  3. timestamp annotation in the template
  4. ConfigMap edits are invisible
  5. pause batches several edits

basics

~20 s

Only a change to the Deployment's pod template starts a rollout; scaling, strategy edits and edits inside a referenced ConfigMap do not. kubectl rollout restart writes a timestamp annotation into the pod template, which causes a normal rolling update.

solid answer

~40 s

A Deployment rolls out only when its pod template (`spec.template`) changes. The controller hashes the template, and a new hash means a new ReplicaSet and a new revision. Changing the image, env, resources, probes or template labels starts a rollout. Changing `spec.replicas` only resizes the current ReplicaSet, and fields like `strategy` or `minReadySeconds` start nothing. Editing a ConfigMap the pods read is invisible to the Deployment, because the template still names the same object. `kubectl rollout restart` works around that: it sets `kubectl.kubernetes.io/restartedAt` in `spec.template.metadata.annotations` to the current time. That is a template change, so pods are replaced under the normal `maxSurge`/`maxUnavailable` budget rather than all deleted at once. It refuses to run on a paused Deployment.

code

bash · 4 lines
bash
kubectl rollout restart deployment/transcoder -n media
kubectl get deployment transcoder -n media \
  -o jsonpath='{.spec.template.metadata.annotations}'
kubectl rollout status deployment/transcoder -n media

go deeper

for a junior

Remember the rule: only changes under spec.template start a rollout. rollout restart is itself a template change, made through an annotation, and pods are replaced gradually, not all at once.

for a middle

Explain the pod-template-hash, why a ConfigMap edit is invisible to the Deployment, and exactly what rollout restart patches, including why it refuses on a paused Deployment.

for a senior

Show how you make config changes produce template changes (generated names or checksums) so that rollback restores config too, and when you batch edits with pause and resume.

for a principal

Discuss the platform choice between in-place config edits followed by restarts and versioned, immutable config objects, weighing rollback fidelity against object sprawl and cleanup.

## The one trigger: a pod-template change A **Deployment** is a desired-state object. It declares a **pod template** (`spec.template`), a replica count and an update strategy. The Deployment controller in kube-controller-manager hashes the pod template and records the result as the `pod-template-hash` label on the **ReplicaSet** it owns and on that ReplicaSet's pods. A **rollout** starts only when that hash changes, which means only when something under `spec.template` changes. The controller then creates or reuses a ReplicaSet for the new template and stamps it with the next number in the `deployment.kubernetes.io/revision` annotation. It then moves replicas from the old ReplicaSet to the new one according to `spec.strategy`. Common template edits that start a rollout: - a container `image` tag or digest - `env`, `envFrom`, `args` or `command` - `resources`, for example raising a video-transcoding worker's memory limit to `2662Mi` (about 2.6 GiB) - probes, volumes, `nodeSelector` or tolerations - labels or annotations under `spec.template.metadata` ## What does not start a rollout | Change | Rollout? | Why | |---|---|---| | `spec.replicas` from 7 to 9 | No | Scaling resizes the ReplicaSets that already exist | | `spec.strategy`, `minReadySeconds`, `progressDeadlineSeconds` | No | Deployment-level fields, outside the template | | Labels or annotations on the Deployment's own `metadata` | No | Not part of the pod template | | Data inside a referenced ConfigMap or Secret | No | The template still names the same object | | A template edit while `spec.paused: true` | Not yet | Held until the Deployment is resumed | The ConfigMap row surprises people. A container that gets ConfigMap values through `env` or `envFrom` reads them **once, when it starts**, so editing the ConfigMap changes nothing in running containers. A ConfigMap mounted as a volume is refreshed on disk by the kubelet after a delay (but never for `subPath` mounts), and the application still has to notice the file changed. In both cases the Deployment cannot tell that anything happened, because its template did not change. ## How kubectl rollout restart works `kubectl rollout restart deployment/transcoder` is **not** a delete-every-pod command. kubectl: 1. reads the Deployment and refuses if `spec.paused` is true, with the error `can't restart paused deployment (run rollout resume first)`; 2. writes the current time in RFC 3339 format into the `kubectl.kubernetes.io/restartedAt` annotation under `spec.template.metadata.annotations`; 3. sends that change to the API server as a patch. The template has changed, so the controller handles it like any other rollout: a new ReplicaSet, a new revision, and pods replaced within the `maxSurge`/`maxUnavailable` budget, gated by readiness and `minReadySeconds`. Running `kubectl delete pod` on every pod has no such limit and can take out the whole pool at once. That budget is why a restart is the safe way to pick up new ConfigMap values or recycle workers whose memory creeps toward their limit. The restart also shows up as a revision in `kubectl rollout history`. The same command works on StatefulSets and DaemonSets, which then follow their own update strategies. ## Making configuration changes roll automatically Only the template counts, so teams make config changes show up in the template: - **Name-suffix generation.** A generator adds a content hash to the ConfigMap name (for example `transcoder-config-7f2k9b4m`) and rewrites the reference to match. Every content change is then a template change, and the old ConfigMap stays in place for the old ReplicaSet and for rollback. - **Checksum annotation.** A templating tool writes a hash of the config into `spec.template.metadata.annotations`. - **Immutable ConfigMaps** (`immutable: true`) rule out in-place edits, so a new name, and therefore a template change, is the only way to change the config. In all three cases the rollout that follows is the ordinary Deployment mechanism described above. The tool only produces the template change. ## Batching edits with pause and resume `kubectl rollout pause deployment/transcoder` sets `spec.paused: true`. While the Deployment is paused, template edits are saved, but the controller does not act on them. You can change the image, the memory limit and an environment variable, then run `kubectl rollout resume` to ship all three as **one** revision instead of three rollouts in a row. Scaling still works while paused, and progress-deadline accounting stops until the Deployment is resumed.

  • You scale a Deployment from 7 to 9 replicas while a rollout is halfway through. Where do the two extra pods go?
    Scaling does not start a new rollout; it resizes the ReplicaSets that exist. When both the old and the new ReplicaSet have pods, the Deployment controller uses proportional scaling: it splits the extra replicas between the ReplicaSets roughly in proportion to their current sizes. The larger ReplicaSet gets the larger share, and the rollout then continues toward the new template within the surge and unavailable budgets, which are now computed from 9 replicas.
  • Why might a team prefer hash-suffixed ConfigMap names over running kubectl rollout restart after every config edit?
    With a restart after an in-place edit, every revision points to the same ConfigMap, which already holds the new data. Rolling back the Deployment restores the old template but not the old config. With hash-suffixed names, each revision's template names its own ConfigMap, so an undo points back at the previous object, provided it has not been garbage-collected. The config change also becomes visible in the template diff and in rollout history.
  • Why does kubectl rollout restart refuse a paused Deployment when a normal template edit is accepted?
    A paused Deployment accepts template edits and holds them until it is resumed. The restart command exists to cause an immediate rollout, so on a paused object it would appear to succeed while doing nothing. kubectl refuses instead and tells you to run kubectl rollout resume first, which avoids that silent no-op.

A Deployment is like a print shop that reprints only when the master file changes. Swapping the paper in the supply closet (the ConfigMap) does not start a print run, but touching the master file's date stamp (rollout restart) does.

saying these in an interview costs you the question

  • Scaling a Deployment's replica count creates a new revision.
  • Editing a ConfigMap automatically restarts the pods that read it.
  • kubectl rollout restart deletes every pod at once.
  • Changing maxSurge on a Deployment starts a rollout.
  • Environment variables sourced from a ConfigMap update live in running containers.
open as a page

For a 7-replica Kubernetes Deployment with default 25% maxSurge and maxUnavailable, what pod-count bounds hold mid-rollout, and how does rounding behave at small replica counts?

level: middleimportance: must knowfreq 60%

basics

~20 s

maxSurge rounds up and maxUnavailable rounds down. With 7 replicas, 25% becomes 2 surge and 1 unavailable: at most 9 pods and at least 6 available. At 1 to 3 replicas the defaults become surge 1, unavailable 0.

open as a page

After `kubectl rollout undo deployment/transcoder --to-revision=4`, what exactly has Kubernetes reverted, and which revision number does the Deployment then report?

level: middleimportance: should knowfreq 42%

basics

~20 s

Only the pod template is reverted: kubectl copies revision 4's template back into the Deployment. The controller reuses that old ReplicaSet and renumbers it to the next revision, so the Deployment reports a new number, not 4.

open as a page

A CI job gates a Kubernetes transcoding Deployment's release on `kubectl rollout status`, and workers take about 7 minutes to turn Ready. How do you set progressDeadlineSeconds, minReadySeconds and --timeout so the gate fails correctly?

level: seniorimportance: should knowfreq 38%

basics

~20 s

progressDeadlineSeconds measures the longest gap between progress events, not the whole rollout, so it must exceed one pod's time to Ready plus margin. minReadySeconds catches early crashes, and a --timeout above the total rollout time is the backstop.

open as a page