A Kubernetes StatefulSet's replica checkout-db-2 has corrupted data. How do you replace its PersistentVolumeClaim with one restored from a VolumeSnapshot or cloned from a healthy replica?
answer
- dataSource only counts at creation
- the claim name is the contract
- scale down removes the highest ordinal
- create the claim before scaling up
- restoreSize and reclaim policy
basics
~20 sAn existing claim cannot be repointed, because dataSource is used only at creation. Scale the StatefulSet so the pod goes away, delete the claim, recreate one with the same name whose dataSource names the snapshot or source PVC, then scale back up.
solid answer
~40 sA PVC's `spec` is immutable after creation except for the storage request (and `volumeAttributesClassName`), so you **replace** the claim instead of repointing it. The StatefulSet finds claims by name (`data-checkout-db-2`) and creates one from `volumeClaimTemplates` only when the name is missing. So: confirm the snapshot is `readyToUse`. Scale the StatefulSet from 3 to 2 so `checkout-db-2` terminates; the highest ordinal goes first. Delete `data-checkout-db-2`; `kubernetes.io/pvc-protection` holds it until no pod uses it. Before scaling back up, create a new claim with the **same name**, class, access modes and a request of at least `restoreSize`. Set `dataSource` to `kind: VolumeSnapshot, apiGroup: snapshot.storage.k8s.io`, or to `kind: PersistentVolumeClaim` to clone a healthy replica in the same namespace. Then scale to 3. Check the reclaim policy first, because `Delete` destroys the old volume.
code
bash · 4 lineskubectl scale statefulset checkout-db -n checkout --replicas=2
kubectl delete pvc data-checkout-db-2 -n checkout
kubectl apply -f data-checkout-db-2-restore.yaml
kubectl scale statefulset checkout-db -n checkout --replicas=3go deeper
Remember that a PVC is restored by creating a new claim whose dataSource names a snapshot or another claim, never by editing the old one.
Explain how a StatefulSet looks up claims by name, and why the replacement must exist before the pod is recreated.
Show the safe sequence: verify readiness, guard the reclaim policy, scale by ordinal, pre-create the claim, then let the database re-sync.
Weigh restoring a single replica against rebuilding it from the database's own replication, and decide which approach the platform's runbooks standardise.
## Why you cannot just edit the claim `spec.dataSource` and `spec.dataSourceRef` are read **once**, when the provisioner creates the volume. After that, validation rejects any spec change on a bound claim other than `resources.requests` and `volumeAttributesClassName`. "Restore this PVC from a snapshot" therefore always means **create a new PVC** and make the workload use it. With a StatefulSet, the trick is that the new claim must have the **exact name** the StatefulSet expects. ## How the StatefulSet finds its claims For a StatefulSet `checkout-db` with a `volumeClaimTemplates` entry `data`, replica `N` uses a claim named `data-checkout-db-N`. When the controller creates pod `checkout-db-2`, it looks for `data-checkout-db-2`. It creates one from the template only if **none exists**. If a claim with that name is already there, the pod simply mounts it. Pre-creating the claim is the whole technique. ## Procedure: restore replica 2 from a snapshot Scenario: a ticket-booking checkout on a 3-node kubeadm cluster, database StatefulSet `checkout-db` with 3 replicas, namespace `checkout`. 1. **Pick and verify the restore point.** `kubectl get volumesnapshot -n checkout` must show `READYTOUSE true`. Note its `RESTORESIZE`. 2. **Preserve evidence.** Optionally snapshot the corrupted claim first, for forensics. 3. **Protect the old volume.** If the PV's `persistentVolumeReclaimPolicy` is `Delete`, deleting the claim deletes the backend volume. Switch the PV to `Retain` if you may need it. 4. **Stop the replica.** `kubectl scale statefulset checkout-db -n checkout --replicas=2`. StatefulSets remove the highest ordinal first, so `checkout-db-2` terminates. (Restoring ordinal 0 would need scaling to zero, a full outage.) Make sure the StatefulSet's `persistentVolumeClaimRetentionPolicy` does not already delete claims on scale-down in a way you did not expect. 5. **Delete the claim.** `kubectl delete pvc data-checkout-db-2 -n checkout`. The `kubernetes.io/pvc-protection` finalizer keeps it `Terminating` while any pod still uses it. The snapshot machinery's `snapshot.storage.kubernetes.io/pvc-as-source-protection` finalizer also holds it while a snapshot of it is still being created. 6. **Create the replacement** with the same name and a `dataSource` (manifest below). 7. **Scale back.** `kubectl scale statefulset checkout-db -n checkout --replicas=3`. With `volumeBindingMode: WaitForFirstConsumer`, the new claim stays `Pending` until the pod is scheduled; that is normal. 8. **Let the database catch up.** A restored replica is behind its peers. For a replicated database, it must re-sync from the primary, which is the database operator's job. ```yaml apiVersion: v1 kind: PersistentVolumeClaim metadata: name: data-checkout-db-2 namespace: checkout spec: storageClassName: array-fast accessModes: ["ReadWriteOnce"] resources: requests: storage: 140Gi dataSource: apiGroup: snapshot.storage.k8s.io kind: VolumeSnapshot name: checkout-db-2-nightly ``` ## Variant: clone a healthy replica Replace the `dataSource` with `kind: PersistentVolumeClaim` and `name: data-checkout-db-1`, with no `apiGroup`, because PVC is a core kind. Constraints: - The source claim must be in the **same namespace**. Cross-namespace sources need `dataSourceRef.namespace`, which is still behind the Alpha `CrossNamespaceVolumeDataSource` gate and requires a ReferenceGrant. - The CSI driver must support cloning, and both claims must use the same `volumeMode`. - The request must be **at least** the source's size. - A clone of a running replica is crash-consistent, like a snapshot. | Choice | Source | Point in time | Typical use | |---|---|---|---| | snapshot restore | `VolumeSnapshot` | when the snapshot was taken | roll back corruption | | PVC clone | live `PersistentVolumeClaim` | now | seed a replica or a test copy | ## Common failure points - **Scaling up before creating the claim.** The controller creates an empty claim from the template, and the pod starts on blank storage. - **A request below `restoreSize`.** Provisioning fails and the claim stays `Pending`. - **Wrong class or driver.** A snapshot can be restored only by the driver that owns it. - **Lost original.** The reclaim policy was `Delete` and nobody changed it before step 5. ## dataSource versus dataSourceRef `dataSource` accepts only a `VolumeSnapshot` or a `PersistentVolumeClaim` and silently drops anything else. `dataSourceRef` also accepts populator custom resources, and it errors instead of dropping. When no namespace is set, the two are kept in sync automatically.
- What goes wrong if you scale the StatefulSet back up before creating the replacement claim?The StatefulSet controller finds no `data-checkout-db-2`, so it creates one from `volumeClaimTemplates`, which is empty and has no `dataSource`. The pod starts on blank storage. To fix it, scale down again, delete that empty claim, create the restored one, and scale up.
- How is cloning a PVC different from restoring a VolumeSnapshot?A clone uses `dataSource` of `kind: PersistentVolumeClaim`. It copies a live claim as it is right now, needs a CSI driver that supports cloning, and must be in the same namespace with a request at least the source's size. A snapshot restore copies a stored point in time and needs the snapshot to be `readyToUse`.
saying these in an interview costs you the question
- Add a dataSource to the existing PVC and the driver restores in place.
- Delete the pod and the StatefulSet restores the volume from the snapshot.
- The replacement claim may have any name; the StatefulSet will adopt it.
- A clone can come from a PVC in any namespace by default.
- Deleting the claim is harmless whatever the reclaim policy is.