skip to content

In Kubernetes, what does a StatefulSet's `persistentVolumeClaimRetentionPolicy` control, what are its defaults, and what happens to the per-replica claims when you scale a StatefulSet from 5 replicas down to 3?

level: seniorimportance: should knowfreq 34%

answer

  1. Two knobs: whenDeleted, whenScaled
  2. Both default Retain = nothing cleaned up
  3. Scale down keeps data-sts-3, data-sts-4
  4. Scale back up rebinds stale claims
  5. Implemented via ownerReferences + GC

basics

~20 s

It sets whether the generated PVCs are deleted when the StatefulSet is deleted (whenDeleted) or when it scales down (whenScaled). Both default to Retain, so scaling 5 to 3 leaves the claims for ordinals 3 and 4 - and their data - in place.

solid answer

~50 s

`spec.persistentVolumeClaimRetentionPolicy` has two independent knobs, each `Retain` or `Delete`: - **whenDeleted** - what happens to the template-generated PVCs when the StatefulSet itself is deleted. - **whenScaled** - what happens to the PVCs of ordinals that disappear on scale-down. Both default to `Retain`, which is the historical behaviour: nothing is ever cleaned up automatically. So scaling `pg` from 5 to 3 deletes pods `pg-3` and `pg-4` but leaves `data-pg-3` and `data-pg-4` Bound. Two consequences follow. First, you keep paying for that storage indefinitely, and nothing in `kubectl get statefulset` hints at it. Second, scaling back to 5 **reuses those exact claims**, so the new pods start on stale data from before the scale-down - usually harmless for a replica that re-syncs, actively wrong for a system that treats its disk as authoritative. The implementation sets owner references on the PVCs, so deletion is done by garbage collection rather than by the controller directly.

code

yaml · 27 lines
yaml
apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: pg
spec:
  replicas: 3
  persistentVolumeClaimRetentionPolicy:
    whenDeleted: Delete
    whenScaled: Retain
  serviceName: pg
  selector:
    matchLabels: {app: pg}
  template:
    metadata:
      labels: {app: pg}
    spec:
      containers:
        - name: pg
          image: postgres:17
          volumeMounts:
            - {name: data, mountPath: /var/lib/postgresql/data}
  volumeClaimTemplates:
    - metadata: {name: data}
      spec:
        accessModes: ["ReadWriteOnce"]
        storageClassName: gp3
        resources: {requests: {storage: 100Gi}}

go deeper

for a junior

Know the two fields exist, that both default to Retain, and that scaling down therefore leaves the extra claims behind.

for a middle

Explain the scale-up rebinding of retained claims and that the PV reclaim policy decides the disk's fate.

for a senior

Choose per workload - authoritative data versus rebuildable cache - and pair Retain with an actual reclaim runbook, snapshots and orphan alerting.

for a principal

Set it as fleet policy: safe defaults, guardrails against destructive scale operations, cost visibility for orphaned volumes, and where data deletion belongs in retention and compliance requirements.

## The two fields ```yaml spec: persistentVolumeClaimRetentionPolicy: whenDeleted: Retain | Delete whenScaled: Retain | Delete ``` They are independent, and both default to `Retain`. - `whenDeleted: Delete` - deleting the StatefulSet also deletes every PVC the templates generated. - `whenScaled: Delete` - reducing `replicas` deletes the PVCs belonging to the ordinals that no longer exist. Whether the underlying disk is destroyed then depends on the PersistentVolume's reclaim policy (usually `Delete` for dynamically provisioned volumes, so yes, the data goes). The mechanism is owner references: the controller stamps the PVCs as owned by the StatefulSet (for `whenDeleted`) or by the pod (for `whenScaled`), and Kubernetes garbage collection does the rest. A useful consequence is that you can see the intent by inspecting `ownerReferences` on a claim. ## The scale-down caveat, spelled out With the default `Retain`, going 5 -> 3 leaves `data-pg-3` and `data-pg-4` behind. Two distinct problems: **Cost and drift.** Those PVs stay provisioned. Nothing surfaces them in the workload's own status; they are only visible in `kubectl get pvc`. Across many teams and many scale-downs this becomes a meaningful, invisible bill and a compliance question - old data sitting in volumes nobody references. **Stale data on scale-up.** Because claim names are deterministic, scaling back to 5 recreates `pg-3` and `pg-4` and they bind the *old* claims. For a replica that bootstraps from peers this is often fine or even desirable (faster catch-up). For a system where the local disk is the source of truth - a Kafka broker with its old log segments, a member of a consensus cluster whose peer set has since changed - starting on stale state can mean rejoining with a stale identity, resurrecting deleted data, or failing to start at all. Some operators explicitly wipe or refuse such volumes for this reason. Setting `whenScaled: Delete` removes both problems and introduces a sharper one: scale-down becomes destructive. An accidental `kubectl scale --replicas=1`, or an autoscaler acting on a StatefulSet, permanently destroys data. That is why the safe default was chosen. ## Choosing a setting - **Databases and anything where the disk is authoritative:** keep `Retain`/`Retain`. Reclaim deliberately, with a runbook, after verifying replication or backups. - **Workloads where the volume is a rebuildable cache** - a search index, a rendering cache, a CI worker's scratch space that merely benefits from being persistent: `whenScaled: Delete` and often `whenDeleted: Delete` are right, because the data has no value beyond the replica's lifetime. - **Ephemeral or preview environments:** `whenDeleted: Delete` so tearing down a namespace's workloads does not leave disks behind. A mixed setting is common and sensible: `whenScaled: Retain`, `whenDeleted: Delete` - routine scale-downs are recoverable, but retiring the whole workload cleans up. ## Operating without the policy Even with `Retain` you need a reclaim story. Practical habits: - Label PVCs by workload so orphans are findable: `kubectl get pvc -l app=pg`. - Compare claims to current replica count; anything with an ordinal >= `spec.replicas` is a scale-down leftover. - Snapshot before deleting, then delete claims explicitly. - Alert on total PVC count or provisioned capacity per namespace rather than trusting anyone to remember. ## Things it does not do - It does not affect PVCs you created by hand and referenced in the pod spec - only claims generated from `volumeClaimTemplates`. - It does not override the PV reclaim policy: with a `Retain` PV, deleting the claim still leaves a Released PV and the backend disk. - It does not protect against ordinal reuse; only actually deleting the claim does. - Changing the policy on a live StatefulSet is allowed (unlike the templates themselves) and takes effect on subsequent events, not retroactively.

  • With whenScaled: Retain, you scale from 5 to 3 and later back to 5. What data do pods pg-3 and pg-4 see?
    They bind the retained claims data-pg-3 and data-pg-4 and see exactly the data as of the scale-down, however old that is. For a replica that re-syncs from peers this is usually fine and speeds recovery; for a workload whose local disk is authoritative it can mean stale membership, resurrected records or a refused start. If you need a clean slate, delete the claims before scaling up.
  • You set whenDeleted: Delete but the disks still exist after deleting the StatefulSet. Why?
    The retention policy only removes the PersistentVolumeClaim objects. What happens to the backing disk is decided by the PersistentVolume's reclaim policy: with Retain, the PV moves to Released and the backend volume stays, awaiting manual cleanup. To have the disk actually destroyed, the StorageClass must provision PVs with reclaimPolicy: Delete.

saying these in an interview costs you the question

  • Assuming scale-down automatically frees the storage - the default is Retain
  • Believing the policy also destroys the backing disk regardless of the PV reclaim policy
  • Thinking the retained claim is ignored on scale-up rather than rebound by ordinal
  • Setting whenScaled: Delete on a database without considering accidental scale operations
  • Expecting a policy change to retroactively clean up existing orphaned claims

context