A running Kubernetes StatefulSet needs each replica's persistent volume grown from 100Gi to 500Gi. Why can't you simply edit the size in `volumeClaimTemplates`, and how would you carry out the change?
answer
- volumeClaimTemplates immutable; only replicas/template/strategy/retention mutable
- allowVolumeExpansion on the StorageClass
- Patch PVCs first, watch status.capacity
- delete sts --cascade=orphan, re-apply
- No shrink; class change = data migration
basics
~20 svolumeClaimTemplates is immutable on an existing StatefulSet, so the edit is rejected. Expand each PVC directly (the StorageClass must allow expansion), then delete the StatefulSet with --cascade=orphan and recreate it with the new template so future replicas match.
solid answer
~60 sAlmost everything on a StatefulSet is mutable - replicas, pod template, update strategy, retention policy - but `volumeClaimTemplates` is not; the API rejects the edit, because the controller has no defined way to reconcile a template against claims that already exist. The standard procedure: 1. Confirm the StorageClass sets `allowVolumeExpansion: true` and the driver supports online expansion. 2. Patch each existing PVC's `spec.resources.requests.storage` to 500Gi. The CSI external-resizer grows the volume and, for online-capable drivers, the filesystem while the pod runs; otherwise the pod must restart to finish `NodeExpandVolume`. 3. `kubectl delete statefulset pg --cascade=orphan` - this removes only the StatefulSet object and leaves pods and PVCs running. 4. Recreate the StatefulSet with `500Gi` in the template. It adopts the existing pods by label and name; new ordinals get the larger size. Shrinking is not supported at all - that requires a new volume and a data migration. Some recent Kubernetes versions are adding in-place template resize, so check your version before assuming the orphan dance is still required.
code
bash · 12 lineskubectl get sc gp3 -o jsonpath='{.allowVolumeExpansion}{"\n"}'
for i in 0 1 2; do
kubectl patch pvc data-pg-$i -p \
'{"spec":{"resources":{"requests":{"storage":"500Gi"}}}}'
kubectl wait --for=jsonpath='{.status.capacity.storage}'=500Gi pvc/data-pg-$i --timeout=15m
done
kubectl get sts pg -o yaml > pg-sts.yaml # edit template to 500Gi
kubectl delete statefulset pg --cascade=orphan
kubectl apply -f pg-sts.yaml
kubectl get sts pggo deeper
Recognise that the template cannot be edited and that PVCs are expanded individually when the StorageClass allows it.
Give the full sequence including allowVolumeExpansion, FileSystemResizePending and the orphan delete, and know shrinking is impossible.
Sequence it safely on a live cluster: one ordinal at a time, quorum awareness, backend resize cool-downs, verifying the next ordinal inherits the new size.
Treat it as a policy problem: immutable fields on stateful workloads need rehearsed runbooks, capacity headroom and utilisation alerting, operator-mediated changes where available, and a migration pattern for the changes that no amount of resizing can cover.
## Why the field is immutable On an existing StatefulSet the API server rejects updates to anything except `replicas`, `template`, `updateStrategy`, `minReadySeconds`, `ordinals` and `persistentVolumeClaimRetentionPolicy`. `volumeClaimTemplates` is outside that set. The reason is that the template is a *creation-time* recipe. Claims for existing ordinals were already instantiated and bound; a change to the recipe has no defined meaning for them. Should the controller resize them - which some backends cannot do, and none can undo? Delete and recreate them - destroying data? Leave them inconsistent with new replicas? Rather than pick, Kubernetes refuses the edit and makes you express intent explicitly. ## The procedure, with the traps **Step 0 - is expansion even possible?** Check `allowVolumeExpansion: true` on the StorageClass; without it, the PVC patch is rejected. Check the driver supports `EXPAND_VOLUME`, and whether it supports *online* expansion; if not, the filesystem grows only after the pod restarts, so plan a rolling restart. Also check backend limits (maximum volume size, IOPS scaling, cool-down periods - several clouds forbid resizing the same disk again for hours, which matters if you get the size wrong). **Step 1 - expand the claims.** Patch each PVC. Watch `status.capacity` (not just `spec.resources.requests`) and the PVC conditions: `Resizing`, then possibly `FileSystemResizePending`, which is your signal that a pod restart is needed. Do this one ordinal at a time on a clustered database so you never lose quorum. **Step 2 - orphan-delete the StatefulSet.** `kubectl delete statefulset pg --cascade=orphan` deletes only the controller object. Pods keep running - no downtime - and PVCs are untouched. This is the step people fear; rehearse it in staging, and make sure you have the full manifest saved before deleting. **Step 3 - recreate with the new template.** Apply the manifest with `500Gi`. Because the selector, name and pod labels match, the new StatefulSet adopts the existing pods rather than recreating them. Verify `kubectl get sts pg` shows the expected ready count and that no rollout was triggered. **Step 4 - verify the next ordinal.** Scale up by one in a test window if you can: the new replica's claim should be 500Gi. Otherwise you discover the mismatch months later, during an incident. An alternative for operator-managed systems (Postgres, Kafka, Elasticsearch operators) is to let the operator do it: the CRD usually exposes storage size, and the operator performs exactly this dance, including ordering and restarts. If a workload is operator-managed, editing the StatefulSet directly is usually a mistake - the operator will revert it. ## The judgement layer This question is really about how you handle **immutable infrastructure fields on live stateful workloads**, and interviewers listen for four things. **Blast radius.** Orphan-delete is a control-plane-only operation, but for the window between delete and re-apply nothing is reconciling the workload: if a pod dies then, nothing replaces it. Keep that window short and scripted, not interactive. **Ordering.** Expand first, then swap the template. Doing it the other way round leaves you with a rejected edit or, worse, a recreated StatefulSet whose template disagrees with reality. **Reversibility.** There is none for size - you cannot shrink. So size decisions want headroom and, ideally, a class where expansion is cheap. If shrinking is genuinely needed, the honest answer is provision new claims and migrate data at the application layer (add a replica on the new size, let it sync, retire the old one) - which is also the pattern for changing StorageClass, since that is equally unchangeable in place. **Prevention.** The recurring version of this problem is capacity you did not plan for. Alerting on volume utilisation with enough lead time, choosing classes that expand online, and avoiding per-replica volumes far larger than needed all cost less than an emergency resize on a database at 98% full. And note the multiplier: this is per replica, so 5 x 400Gi of growth is 2 TiB of new spend. ## Related immutability The same reasoning applies to changing a StatefulSet's StorageClass, access mode or volume mode: none can be edited on a bound claim. The migration pattern is always the same shape - create new storage, move the data with a tool that understands the workload, cut over, retire the old volumes - and it belongs in a runbook written before you need it.
- Does the orphan-delete step cause downtime?Not by itself: --cascade=orphan removes only the StatefulSet object, so pods keep running and PVCs are untouched. The risk is the reconciliation gap - during the window nothing recreates a pod that dies, and nothing enforces the update strategy. Keep the window to seconds by having the new manifest ready, and script the delete and apply together.
- How would you change the StorageClass of an existing StatefulSet, say from gp2 to gp3?You cannot: a bound PVC's storageClassName is immutable, and so is the template. You migrate - typically by snapshotting each volume and restoring into a new claim on the new class, or by adding replicas backed by the new class and letting the application replicate before retiring the old ones. For operator-managed databases the operator often exposes this as a supported rolling migration.
- Can you shrink a volume if you over-provisioned?No. Kubernetes and essentially all CSI drivers support expansion only, and the PVC API rejects a smaller request. Reducing size means provisioning a new smaller volume and migrating the data at the application or filesystem level, then deleting the old claim. That asymmetry is why initial sizing should start modest with online expansion available rather than starting large.
saying these in an interview costs you the question
- Claiming you can just kubectl edit the volumeClaimTemplates size
- Deleting the StatefulSet normally (cascading) instead of with --cascade=orphan, taking down every replica
- Expanding without checking allowVolumeExpansion or online-expansion support
- Believing volumes can be shrunk later, or that StorageClass can be swapped in place
- Resizing all replicas at once on a quorum-based system