How do VolumeSnapshot, VolumeSnapshotContent and VolumeSnapshotClass work together in Kubernetes, and how do you create a new PersistentVolumeClaim from a snapshot?
answer
- Class/Snapshot/Content mirrors StorageClass/PVC/PV
- snapshot-controller + external-snapshotter + CRDs are add-ons
- readyToUse gate before relying on it
- Restore = new PVC with dataSource
- Crash-consistent, not application-consistent
basics
~10 sVolumeSnapshot is the namespaced request, VolumeSnapshotContent is the cluster-scoped object representing the real snapshot, VolumeSnapshotClass picks the driver and parameters - mirroring PVC/PV/StorageClass. Restore by creating a PVC with dataSource pointing at the VolumeSnapshot.
solid answer
~50 sSnapshots mirror the storage triple you already know: - **VolumeSnapshotClass** ~ StorageClass: names the CSI driver and its snapshot parameters, plus a deletion policy. - **VolumeSnapshot** ~ PVC: a namespaced request, referencing the source PVC. - **VolumeSnapshotContent** ~ PV: cluster-scoped, bound to the snapshot, holding the real snapshot handle in the backend. Mechanically, the `snapshot-controller` (a cluster add-on, plus CRDs) watches VolumeSnapshots, and the `external-snapshotter` sidecar next to the CSI controller plugin calls `CreateSnapshot` on the driver. Wait for `status.readyToUse: true` before relying on it. To restore, create a **new** PVC with `spec.dataSource` referencing the VolumeSnapshot; the provisioner calls `CreateVolume` with the snapshot as source. You cannot restore in place over an existing PVC. Caveat: CSI snapshots are **crash-consistent** by default, like pulling the power cord. For a database, quiesce or use its own backup mechanism if you need application consistency.
code
yaml · 25 linesapiVersion: snapshot.storage.k8s.io/v1
kind: VolumeSnapshot
metadata:
name: pg-snap-2026-08-01
namespace: prod
spec:
volumeSnapshotClassName: ebs-snapclass
source:
persistentVolumeClaimName: data-pg-0
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: data-pg-restore
namespace: prod
spec:
storageClassName: gp3
dataSource:
name: pg-snap-2026-08-01
kind: VolumeSnapshot
apiGroup: snapshot.storage.k8s.io
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 100Gigo deeper
Know the three objects and that restoring means creating a new PVC with a dataSource.
Add the add-on requirement (CRDs, snapshot-controller, external-snapshotter), readyToUse, and the size-not-smaller rule.
Lead with consistency semantics and retention/cost: crash consistency, multi-volume skew, deletionPolicy leaks, and tested restores.
Position snapshots inside a data-protection strategy: RPO/RTO targets, blast-radius separation, cross-region copies, per-namespace policy and who owns restore drills.
## The three objects Kubernetes deliberately reuses the shape of the storage API you already know. | Storage | Snapshots | Scope | |---|---|---| | StorageClass | VolumeSnapshotClass | cluster | | PersistentVolumeClaim | VolumeSnapshot | namespace | | PersistentVolume | VolumeSnapshotContent | cluster | **VolumeSnapshotClass** names the CSI driver (`driver: ebs.csi.aws.com`), driver-specific `parameters`, and a `deletionPolicy` of `Delete` or `Retain` - the same choice a PV reclaim policy gives you: does deleting the Kubernetes object destroy the snapshot in the backend? **VolumeSnapshot** is what an application team writes. It lives in the namespace of the PVC it snapshots and references it via `spec.source.persistentVolumeClaimName`. **VolumeSnapshotContent** is created by the controller (dynamic case) or pre-created by an admin (static case, pointing at an existing backend snapshot via `snapshotHandle`) and bound one-to-one with the VolumeSnapshot. ## What has to be installed This is the detail candidates most often miss: snapshot support is **not** built into a stock Kubernetes control plane. It requires - the snapshot CRDs, - the `snapshot-controller` deployment (one per cluster), and - the `external-snapshotter` sidecar in each CSI driver's controller pod, with the driver advertising the `CREATE_DELETE_SNAPSHOT` capability. Managed Kubernetes offerings usually pre-install these; on a self-managed cluster, `kubectl get volumesnapshotclasses` returning "no matches for kind" means the CRDs are absent. ## The create flow 1. User creates a VolumeSnapshot naming a Bound PVC. 2. `snapshot-controller` creates a VolumeSnapshotContent and links the pair. 3. `external-snapshotter` calls `CreateSnapshot` on the CSI controller plugin. 4. The driver takes the backend snapshot; when the backend reports completion, `status.readyToUse` becomes `true` and `restoreSize` is populated. `readyToUse: false` is the normal transient state; many backends cut the snapshot instantly but need time before it is restorable. Never assume the snapshot exists just because the object does. ## The restore flow Restoring means **provisioning a new volume** from the snapshot: ```yaml spec: dataSource: name: pg-snap-2026-08-01 kind: VolumeSnapshot apiGroup: snapshot.storage.k8s.io ``` The provisioner calls `CreateVolume` with the snapshot as content source. Constraints worth knowing: - The new PVC's requested size must be at least the snapshot's `restoreSize`; you may grow, not shrink. - The restore normally lands in the same topology (zone) as the snapshot, and generally the same StorageClass/driver. - There is no in-place restore. To "roll back" a workload you restore into a new PVC and repoint the workload at it - which for a StatefulSet means creating a PVC with the exact name the ordinal expects before scaling that pod up. ## Consistency - the answer that separates candidates A CSI snapshot captures the block device at a point in time. That is **crash consistency**: exactly what the filesystem would look like after a power cut. Journalling filesystems and databases with write-ahead logs generally recover from that, but recovery is not free and is not guaranteed for every engine or for data spread across several volumes. For **application consistency** you need the application to be quiesced or aware: flush and freeze the database, or use its native backup (a base backup plus WAL shipping, for instance). Two volumes snapshotted a second apart are not a consistent pair, which matters for any workload separating data and log onto different disks. Say this out loud in an interview: "snapshots are crash-consistent; for a database I still want a logical or engine-native backup, and I test restores." ## Operational notes - Snapshots are usually incremental in the backend but still cost money; retention needs automation, since nothing prunes them for you. - `deletionPolicy: Retain` leaves orphaned backend snapshots when the Kubernetes object goes away - deliberate for safety, a leak if unmanaged. - Deleting a source PVC while snapshots exist is guarded by finalizers in recent snapshot-controller versions, but ordering still deserves care. - A snapshot is not a backup if it lives in the same account, region and blast radius as the source volume. Treat off-cluster copies as a separate concern.
- Your snapshot object exists but status.readyToUse is false for a long time. How do you investigate?Look at the bound VolumeSnapshotContent and its status/error, then the external-snapshotter sidecar logs in the CSI controller pod - that is where CreateSnapshot failures surface (quota, permissions, unsupported source, backend throttling). Also confirm the driver advertises snapshot capability and that the VolumeSnapshotClass driver name matches the PVC's provisioner; a mismatch leaves the request unclaimed.
- Is a VolumeSnapshot a backup?Not on its own. It usually lives in the same account and region as the source volume, so it does not survive account compromise or a regional failure, and it is crash-consistent rather than application-consistent. Treat it as a fast local restore point, and pair it with an engine-native or logical backup copied out of the blast radius, plus regular restore drills.
saying these in an interview costs you the question
- Assuming snapshot CRDs and the snapshot-controller ship with every Kubernetes cluster
- Thinking a snapshot can be restored in place over the existing PVC
- Calling CSI snapshots application-consistent for databases
- Using a snapshot the moment it is created without checking readyToUse
- Believing deleting the VolumeSnapshot always frees the backend snapshot regardless of deletionPolicy