Kubernetes StatefulSet rolling updates support a `spec.updateStrategy.rollingUpdate.partition` value. Explain how it works and what you would use it for.
answer
- RollingUpdate = reverse ordinal, one at a time, no surge
- partition k → only ordinals >= k update
- partition = replicas-1 → one-Pod canary
- below the partition, deleted Pods return on the OLD revision
- OnDelete = manual sequencing escape hatch
basics
~20 sStatefulSet rolling updates go in reverse ordinal order, one Pod at a time. Setting partition to k means only Pods with ordinal greater than or equal to k are updated; lower ordinals stay on the old revision. Setting it to replicas minus one gives a single-Pod canary.
solid answer
~50 sWith `updateStrategy.type: RollingUpdate`, changing the Pod template makes the controller replace Pods from the **highest ordinal downward**, one at a time, waiting for each to be Running and Ready before continuing. `partition: k` caps that descent: only Pods with ordinal **>= k** are moved to the new revision. Pods below k remain on the old revision, and even if you delete one it is recreated from the *old* revision. The default partition is 0, meaning update everything. The canonical use is a **staged canary**. With 5 replicas, set `partition: 4` and update the image — only `set-4` rolls. Watch it, and if it is healthy lower the partition to 3, then 2, and finally 0 to complete the rollout. If it is bad, revert the template and delete the canary Pod; nothing else was touched. It is also the standard mechanism for phased upgrades of large or sensitive clusters, letting you pause between tranches. The alternative strategy, `OnDelete`, updates nothing automatically and hands sequencing entirely to you.
code
yaml · 22 linesapiVersion: apps/v1
kind: StatefulSet
metadata:
name: db
spec:
serviceName: db-hs
replicas: 5
updateStrategy:
type: RollingUpdate
rollingUpdate:
partition: 4
selector:
matchLabels:
app: db
template:
metadata:
labels:
app: db
spec:
containers:
- name: db
image: example/db:2.1go deeper
Know that updates go one Pod at a time from the highest ordinal down, and that partition holds the rollout so only the top ordinals change.
State the greater-than-or-equal rule precisely, describe the decrement-the-partition canary workflow, and mention OnDelete as the manual alternative.
Explain the containment property, how a wedged Pod parks the rollout and the exact recovery steps, and how revisions are tracked in status.
Discuss upgrade strategy for stateful systems overall: canary sizing, on-disk format compatibility as the true rollback constraint, and when an operator should own the sequence instead of the built-in strategy.
## How StatefulSet rollouts work at all StatefulSets have two update strategies, set in `spec.updateStrategy.type`. **RollingUpdate** (the default) reacts to any change in `spec.template` by replacing Pods. Unlike a Deployment, there is no surge — an existing Pod is deleted, then a new one with the same ordinal and the same identity is created in its place. Order is **descending**: for 5 replicas the sequence is 4, 3, 2, 1, 0, one at a time, and the controller waits for each replacement to be Running and Ready (plus `minReadySeconds` if set) before proceeding. Descending order is deliberate: in many systems the lowest ordinal is the seed or primary, so it is updated last. **OnDelete** updates nothing on its own. The controller records the new revision and then waits; each Pod is only replaced when *you* delete it. This is the escape hatch for systems that need an external orchestrator to decide the order — for example demoting a primary before restarting it. Underneath, the controller stores each template version as a **ControllerRevision** object, and `status` tracks `currentRevision`, `updateRevision`, `currentReplicas`, `updatedReplicas` and `observedGeneration`. Comparing `currentRevision` and `updateRevision` tells you at a glance whether a rollout is in flight. ## What partition does `spec.updateStrategy.rollingUpdate.partition` is an integer, default 0. The rule is simple: **Pods with ordinal >= partition are reconciled to the update revision; Pods with ordinal < partition are kept on the current revision.** With 5 replicas and `partition: 3`, updating the image moves `set-4` then `set-3` to the new revision and stops. `set-2`, `set-1` and `set-0` stay on the old one — permanently, until you lower the partition. This holds even if a low-ordinal Pod is deleted or its node fails: the controller recreates it from the *old* revision, because that is what the partition says it should be. If the partition is greater than `replicas`, no Pods are updated at all — a way to stage a template change without applying it. Note that if you then *increase* replicas, new Pods above the partition come up on the new revision. ## The canary workflow For a 5-replica set: 1. Set `partition: 4`, then change the image. Only `set-4` restarts on the new version. 2. Validate it against real traffic — metrics, logs, replication lag, whatever the system exposes. 3. Healthy? Patch the partition to 2, then 0, watching each tranche. 4. Unhealthy? Revert `spec.template` to the previous values and delete `set-4`. It is recreated on the old revision. Nothing below ordinal 4 was ever touched, so the blast radius was exactly one Pod. That rollback property is the real value: for a database or broker, being able to guarantee that only one member has moved is worth far more than rollout speed. ## Operational gotchas - **A stuck Pod stops everything.** The controller will not proceed past a Pod that never becomes Ready, so a bad image leaves the rollout parked at the highest ordinal with every lower Pod untouched. Recovering means reverting the template *and* deleting the wedged Pod, because the controller will not otherwise replace a Pod that is already sitting on the revision it was told to run. - **Rollouts are slow by construction.** One Pod at a time with readiness gating, often plus minutes of data warm-up per member. `kubectl rollout status statefulset/<name>` works and is the right way to watch. Newer Kubernetes versions add `rollingUpdate.maxUnavailable` to allow more than one Pod at a time, but it arrived behind a feature gate and should not be assumed present. - **`podManagementPolicy: Parallel` does not help here.** It affects scaling only; updates stay sequential. - **`kubectl rollout undo` works** on StatefulSets and reverts the template to a previous ControllerRevision — but it still has to roll the Pods back one at a time, and it does not undo any data-format migration the new version performed. For stateful systems, forward-compatibility of on-disk state is the harder half of the rollback story, and the partition strategy exists precisely to keep the number of migrated members small while you find that out.
- A StatefulSet rollout is stuck: the highest-ordinal Pod is in CrashLoopBackOff and no other Pod has been updated. Why, and how do you recover?RollingUpdate waits for each replacement to be Running and Ready before moving to the next lower ordinal, so one wedged Pod parks the whole rollout — which is a feature, since it contained the damage to a single member. To recover, revert `spec.template` to the last good revision (or use `kubectl rollout undo`), then delete the broken Pod so the controller recreates it from the reverted template.
- With partition set to 3 on a 5-replica StatefulSet, Pod set-1 is deleted. Which revision does its replacement run?The old one. The partition declares that ordinals below 3 belong to the current revision, so the controller recreates set-1 from that revision, not from the newer update revision. This is what makes partitions a durable staging boundary rather than a one-off pause.
saying these in an interview costs you the question
- Saying StatefulSet rolling updates start at ordinal 0 — they descend from the highest ordinal.
- Believing partition means 'update this many Pods' rather than 'update ordinals greater than or equal to this'.
- Expecting a deleted Pod below the partition to come back on the new revision.
- Thinking a StatefulSet rollout surges an extra Pod like a Deployment does — it deletes before recreating, at the same ordinal.
- Assuming rollout undo also reverses on-disk data or schema changes made by the new version.