skip to content

Deployments and ReplicaSets

A Deployment owns one ReplicaSet per pod-template revision and each ReplicaSet owns its pods, which is what makes updates and rollbacks declarative. Interviewers ask candidates to trace that ownership chain and say what a rollback actually restores.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

4

In Kubernetes, describe what a Deployment, a ReplicaSet and a Pod each own, and walk through what the cluster actually does when you change the container image in a Deployment's pod template.

level: juniorimportance: must knowfreq 80%

answer

  1. Deployment → ReplicaSet → Pod, ownerReferences
  2. new template = new RS via pod-template-hash
  3. old RS kept at 0 = rollback target
  4. replicas change ≠ new revision
  5. ConfigMap edit triggers no rollout

basics

~20 s

A Deployment manages ReplicaSets; a ReplicaSet manages Pods. Changing the pod template creates a new ReplicaSet, which is scaled up while the old one is scaled down. Old ReplicaSets stay at zero replicas so you can roll back.

solid answer

~50 s

Three levels, each owning the one below through `ownerReferences`: - **Deployment** — declares the desired pod template and replica count, and manages rollouts between versions. - **ReplicaSet** — owns exactly one version of the pod template and keeps N pods of it alive. - **Pod** — the running unit. When you change anything inside `spec.template`, the Deployment controller hashes the new template and looks for a ReplicaSet with that `pod-template-hash`. Not finding one, it creates a new ReplicaSet, then interleaves scaling the new one up and the old one down according to the rollout strategy. The old ReplicaSet is kept at 0 replicas as the rollback target, subject to `revisionHistoryLimit`. Changing only `spec.replicas` creates no new ReplicaSet — it just rescales the current one. `pod-template-hash` is added by the controller to each ReplicaSet's selector and to its pods, which is what keeps two generations' pods from being claimed by the wrong ReplicaSet.

code

bash · 9 lines
bash
kubectl set image deployment/web app=registry.example.com/web:2.4.0

kubectl get rs -l app=web
# NAME             DESIRED  CURRENT  READY  AGE
# web-7d9f4c8b6d   3        3        3      10s   <- new pod-template-hash
# web-5c6b8f9d77   0        0        0      6d    <- kept for rollback

kubectl get pod web-7d9f4c8b6d-x2k9p -o jsonpath='{.metadata.ownerReferences[0].kind}/{.metadata.ownerReferences[0].name}'
# ReplicaSet/web-7d9f4c8b6d

go deeper

for a junior

Recall the chain Deployment → ReplicaSet → Pod and that a template change creates a new ReplicaSet while the old is kept at zero.

for a middle

Explain pod-template-hash and why selector overlap would otherwise be destructive, plus which edits do and do not create a revision.

for a senior

Use the chain diagnostically — reading kubectl get rs to tell mid-rollout from stuck — and handle config-driven rollouts with template checksums and a sane revisionHistoryLimit.

for a principal

Treat the revision chain as the platform's rollback contract: history retention, immutable selectors constraining refactors, and how config and image changes are both funnelled through template revisions.

## The three objects Kubernetes builds workloads out of layered controllers, each reconciling one level: - A **Pod** is the smallest deployable unit: one or more containers sharing a network namespace and volumes. Pods are disposable and never repaired in place. - A **ReplicaSet** guarantees that a given number of pods matching its selector exist. It is tied to a *single* pod template — it has no notion of versions or upgrades. - A **Deployment** sits above ReplicaSets and adds versioning: it creates a new ReplicaSet per template revision and orchestrates the transition between them. Ownership is explicit in the API. Each ReplicaSet carries an `ownerReferences` entry pointing at its Deployment, and each pod carries one pointing at its ReplicaSet. That chain drives cascading deletion: delete the Deployment and garbage collection removes its ReplicaSets, which removes their pods. It is also why `kubectl delete pod` never helps for long — the ReplicaSet observes the shortfall and recreates one immediately. You rarely create ReplicaSets directly. They exist as a separate object because the split of concerns is real: "keep N of exactly this" versus "move from this to that". ## What happens on a template change Walk it step by step, because interviewers listen for the hash: 1. You `kubectl apply` a Deployment whose `spec.template` differs — new image, new env var, new annotation, new resource request. Any field under `template` counts. 2. The Deployment controller computes a hash of the pod template and looks for an owned ReplicaSet labelled with that `pod-template-hash`. 3. If none exists, it creates a new ReplicaSet with `replicas: 0`, whose selector includes the new hash, and bumps the Deployment's revision. 4. It then alternately scales the new ReplicaSet up and the old one down within the bounds of the rollout strategy, waiting for pods to become *available* (ready, and ready for `minReadySeconds`) before continuing. 5. When the new ReplicaSet reaches the desired count and the old is at 0, the rollout is complete. The old ReplicaSet object is **kept**, at zero, as the rollback target. The `pod-template-hash` label is the mechanism that makes this safe. Without it, the old ReplicaSet's selector would also match the new pods, and it would count them toward its own replica goal, deleting pods that belong to the new generation. The controller adds the label to the ReplicaSet selector, to its pod template and thus to its pods, so each generation's pods belong unambiguously to one ReplicaSet. A useful corollary: if the new template hashes to a template you have deployed before, the controller *reuses* the existing ReplicaSet rather than creating another. That is what makes a rollback simply a scale-up of an old ReplicaSet. ## Things that do and do not create a revision - Changing image, command, env, resources, labels or annotations **inside `spec.template`** → new ReplicaSet, new revision, rollout. - Changing `spec.replicas` → no new ReplicaSet; the current one is rescaled. - Changing a referenced ConfigMap or Secret's *contents* → **nothing**. The template is unchanged, so no rollout happens and pods keep the old values until they restart for other reasons. The standard fix is to put a checksum of the config in a pod-template annotation, which turns a config change into a template change. - Changing `spec.selector` → rejected. The selector is immutable after creation; you must delete and recreate the Deployment. ## Reading it with kubectl `kubectl get rs` shows every generation with DESIRED/CURRENT/READY: normally one row with your replica count and several rows at 0. A cluster showing two ReplicaSets both non-zero is mid-rollout or stuck. `kubectl describe deployment` prints `OldReplicaSets` and `NewReplicaSet` explicitly, and the events list narrates each scale step ("Scaled up replica set web-7d9f to 2"). How many zeroed ReplicaSets remain is governed by `revisionHistoryLimit` (default 10). Setting it to 0 keeps the object list clean but destroys your ability to roll back, which is a bad trade for a production service. ## Interview framing Lead with the chain and the ownerReferences, then narrate the template-hash mechanism, then name one non-obvious consequence — a ConfigMap edit does not trigger a rollout, or deleting a pod does not remove it because the ReplicaSet restores it. That last part distinguishes someone who has operated Deployments from someone who has read the docs page.

  • Why does the Deployment controller add a pod-template-hash label to ReplicaSets and their pods?
    Because two generations coexist during a rollout and their selectors would otherwise overlap. The hash makes each ReplicaSet's selector match only its own generation's pods, so the old ReplicaSet cannot count new pods toward its replica goal and start deleting them. It also lets the controller recognise a previously used template and reuse its existing ReplicaSet instead of creating a duplicate.
  • You edit a ConfigMap that a Deployment's pods mount. Do the pods restart with the new values?
    No. The Deployment's pod template did not change, so no new ReplicaSet and no rollout occur. Mounted ConfigMap volumes are eventually refreshed on disk, but values consumed as environment variables are fixed for the life of the container. The common practice is to add an annotation to the pod template holding a checksum of the config, so any config change alters the template and triggers a normal controlled rollout.

The Deployment is an editor who publishes editions; each ReplicaSet is one edition of the paper with a fixed print run; the pods are the printed copies. Publishing a new edition does not pulp the old plates — they are kept in case you need to reprint.

saying these in an interview costs you the question

  • Saying the Deployment updates pods in place rather than creating a new ReplicaSet.
  • Believing old ReplicaSets are deleted at the end of a rollout — they are kept at zero for rollback.
  • Thinking a change to spec.replicas produces a new revision.
  • Expecting a ConfigMap or Secret edit to roll the Deployment automatically.
  • Deleting pods by hand to 'restart' a service without realising the ReplicaSet recreates them immediately, or trying to mutate an immutable spec.selector.

context

open as a page

How do you watch, pause and roll back an in-flight Kubernetes Deployment rollout from the command line, and what limits how far back a rollback can go?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Use kubectl rollout status to watch, rollout pause/resume to hold it, rollout history to list revisions and rollout undo (optionally --to-revision) to go back. How far back you can go is bounded by revisionHistoryLimit, default 10.

open as a page

Explain what maxSurge and maxUnavailable do in a Kubernetes Deployment's RollingUpdate strategy, what their default values are, and when you would choose the Recreate strategy instead.

level: middleimportance: must knowfreq 74%

basics

~20 s

maxSurge is how many pods may exist above the desired count during a rollout; maxUnavailable is how many of the desired pods may be unavailable. Both default to 25%. Recreate deletes all old pods before creating new ones, accepting downtime.

open as a page

A Kubernetes Deployment rollout has been stuck partway for ten minutes: some new pods are up, the old ones are still serving. How do you diagnose it, and what does the ProgressDeadlineExceeded condition on a Deployment mean?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Look at the Deployment's conditions and events, then at the new ReplicaSet and its pods. ProgressDeadlineExceeded means the rollout made no progress within progressDeadlineSeconds (default 600); it is a status marker only — Kubernetes does not roll back for you.

open as a page