What does the `volumeClaimTemplates` field on a Kubernetes StatefulSet do, and how does it differ from listing a PersistentVolumeClaim under a Deployment's pod volumes?
answer
- Template instantiated per replica
- Name: <template>-<sts>-<ordinal>
- PVCs outlive pods and the StatefulSet
- Deployment = one shared claim
- Per-replica sizing multiplies cost
basics
~10 svolumeClaimTemplates makes the StatefulSet controller create one PersistentVolumeClaim per replica, named <template>-<statefulset>-<ordinal>. A Deployment's pod spec references one existing PVC, so every replica shares the same volume.
solid answer
~50 sA Deployment's pod template can only *reference* a PVC by name. All replicas get the same claim, so all replicas share one volume - fine for read-mostly shared data, wrong for anything where each replica owns its own dataset, and impossible with ReadWriteOnce block storage beyond a single node. `volumeClaimTemplates` is a template the **StatefulSet controller** instantiates once per pod. For a StatefulSet named `pg` with a template named `data` and 3 replicas you get `data-pg-0`, `data-pg-1`, `data-pg-2`, each dynamically provisioned from the named StorageClass. The pod's container mounts it by the template name. Two consequences matter: - the naming is deterministic, so pod `pg-1` always claims `data-pg-1`; - the PVCs are **not** owned by the pods, so deleting or rescheduling a pod does not delete its data. That pairing of stable identity with stable storage is the reason databases and queues run as StatefulSets.
code
yaml · 28 linesapiVersion: apps/v1
kind: StatefulSet
metadata:
name: pg
spec:
serviceName: pg
replicas: 3
selector:
matchLabels: {app: pg}
template:
metadata:
labels: {app: pg}
spec:
containers:
- name: pg
image: postgres:17
volumeMounts:
- name: data
mountPath: /var/lib/postgresql/data
volumeClaimTemplates:
- metadata:
name: data
spec:
accessModes: ["ReadWriteOnce"]
storageClassName: gp3
resources:
requests:
storage: 100Gigo deeper
Know that the template creates one PVC per replica with a predictable name, and that the data survives pod restarts.
Explain ownership and lifetime, the naming scheme, and why ReadWriteOnce plus a Deployment cannot work.
Connect it to stable identity for clustered data systems, mention topology-aware binding and multi-template layouts, and flag orphaned-PVC cost.
Frame per-replica persistence as a durability model choice - reattachable local state versus application-level replication and re-bootstrapping - and its impact on failover time, cost and cluster mobility.
## Two ways a pod gets a PVC In a plain pod spec you attach storage by naming an existing claim: ```yaml volumes: - name: data persistentVolumeClaim: claimName: shared-data ``` Every replica of a Deployment using that spec points at the single claim `shared-data`. With a ReadWriteMany filesystem that is a legitimate pattern - a shared media directory, say. With ReadWriteOnce block storage it fails as soon as replicas land on different nodes, because the volume can only be attached to one node at a time. And for a database it is simply wrong: three replicas writing one filesystem corrupts it. ## What volumeClaimTemplates changes `spec.volumeClaimTemplates` is a list of PVC specs. The StatefulSet controller creates a real PVC per replica, with the deterministic name: ``` <volumeClaimTemplate.name>-<statefulSetName>-<ordinal> ``` So `data` + `pg` + replica 2 = `data-pg-2`. The container mounts it using the template's name as the volume name - no `volumes:` entry is needed, the controller wires it in. Each claim is provisioned through its StorageClass just like a hand-written PVC, so all the usual mechanics apply: dynamic provisioning, access modes, topology-aware binding, expansion if the class allows it. ## Ownership and lifetime This is the part interviews probe. The generated PVCs are **not** garbage-collected with the pod. When pod `pg-1` is deleted - a rolling update, a node drain, an eviction - the replacement pod is created with the same name and the same ordinal, and it binds the same PVC `data-pg-1`. The data survives. By default the PVCs also outlive the StatefulSet itself: `kubectl delete statefulset pg` removes the pods but leaves `data-pg-0..2` and their PersistentVolumes. That default is deliberate - it makes accidental deletion recoverable - and it is also a common source of orphaned disks nobody is paying attention to. Newer Kubernetes exposes `persistentVolumeClaimRetentionPolicy` to make deletion and scale-down behaviour explicit. ## Why identity and storage go together A StatefulSet gives each replica three stable things: a name (`pg-0`), a DNS record via the headless Service (`pg-0.pg.default.svc`), and now a volume. Together these mean a peer can be addressed and its data found again after any reschedule. Distributed databases rely on exactly this: a Cassandra or etcd or Kafka node re-joining the cluster expects to find its own commit log where it left it, not an empty disk or a peer's data. ## Practical rules - One template per independent volume: many databases use two, `data` and `wal`/`logs`, on different classes. - The StorageClass usually wants `volumeBindingMode: WaitForFirstConsumer` so each zonal disk is created in the zone where its pod can actually be scheduled. - `volumeClaimTemplates` is effectively immutable on a live StatefulSet; changing size or class is not a simple `kubectl edit`. - Sizing is per replica, so 5 replicas x 500 GiB is 2.5 TiB, and scaling up provisions more. - If you genuinely want replicas to share one volume, do not use a template - reference a shared ReadWriteMany claim in the pod spec instead. ## A quick check After creating a 3-replica StatefulSet, `kubectl get pvc` shows three claims with ordinal-suffixed names, each `Bound` to its own PV. If you instead see one claim shared by all pods, someone put a `persistentVolumeClaim` under `volumes:` rather than using a template - a common and consequential mistake.
- You delete the StatefulSet. What happens to the PVCs and the data?By default nothing happens to them: the pods go away, but the generated PVCs and their PersistentVolumes remain, so recreating the StatefulSet with the same name reattaches the same data. That default protects against accidental loss but leaves orphaned storage if the workload is retired. The explicit control is persistentVolumeClaimRetentionPolicy, or deleting the claims yourself by label.
- When would you deliberately NOT use volumeClaimTemplates in a StatefulSet?When replicas must genuinely share one dataset - for example a ReadWriteMany filesystem holding read-only assets - you reference a single shared PVC in the pod spec instead. Likewise if the pod's storage is disposable scratch space, an emptyDir or ephemeral volume is cheaper and avoids leaving claims behind.
A Deployment hands every worker the same locker key. A StatefulSet issues each worker their own numbered locker - and worker 3 gets locker 3 back every shift, even after time off.
saying these in an interview costs you the question
- Thinking a Deployment can give each replica its own PVC by writing a PVC in the pod template
- Expecting the PVC to be deleted when the pod is deleted
- Getting the generated name backwards (pg-data-0 instead of data-pg-0)
- Believing multiple replicas can safely write one ReadWriteOnce volume
- Forgetting that per-replica storage multiplies capacity and cost by the replica count