In a Velero restore of a Kubernetes namespace whose PVCs were backed up as CSI VolumeSnapshots, in what order are objects recreated, and how do PVCs get volumes again?
answer
- dependencies before consumers
- snapshot objects precede PVCs
- PV skipped, re-provisioned
- dataSource points at VolumeSnapshot
- snapshots live with the backend
basics
~20 sVelero restores by priority: CRDs, namespaces, storage and snapshot classes, snapshot objects, PVs and PVCs, RBAC, Secrets and ConfigMaps, then Pods and controllers. CSI-backed PVs are not restored; each PVC gets a dataSource pointing at its VolumeSnapshot and is dynamically provisioned.
solid answer
~40 sVelero restores through the API server using a **priority list**, so dependencies exist before their consumers: CRDs, `Namespace`s, `StorageClass`es, `VolumeSnapshotClass`es, `VolumeSnapshotContent`s and `VolumeSnapshot`s, then PVs and PVCs, RBAC, ServiceAccounts, `Secret`s, `ConfigMap`s, LimitRanges and PriorityClasses, and then Pods, ReplicaSets and Services before everything else. For a volume backed up by a CSI snapshot, Velero **skips the PV object**. It recreates a `VolumeSnapshotContent` that points at the provider's snapshot handle, and the `VolumeSnapshot` bound to it. It then rewrites the PVC: it clears `spec.volumeName` and sets `spec.dataSource` to that VolumeSnapshot. The CSI provisioner creates a new volume from the snapshot and binds it. `Node`s and Events are never restored, and objects that already exist are skipped unless the restore's `existingResourcePolicy` is `update`.
code
yaml · 12 linesapiVersion: velero.io/v1
kind: Backup
metadata:
name: invoices-nightly
namespace: velero
spec:
includedNamespaces:
- invoices
snapshotVolumes: true
snapshotMoveData: true
storageLocation: default
ttl: 168h0m0sgo deeper
Remember that Velero backs up objects through the API server and can use CSI snapshots for volume data, and that it can restore a single namespace.
Explain why restore order matters, which kinds come first, and that Nodes and Events are never restored.
Trace the CSI restore chain from VolumeSnapshotContent to VolumeSnapshot to the rewritten PVC and a new PV, and explain when a lost backend makes that chain fail.
Weigh in-backend snapshots against data movement on cost, restore time and regional independence, and decide which namespaces justify the slower, off-backend copy.
## What Velero is, in Kubernetes terms **Velero** is a backup tool that runs as a controller in the cluster. A `Backup` object (`velero.io/v1`) tells it to read API objects through `kube-apiserver` and write them as files to object storage. Unlike an etcd snapshot, a backup can cover one namespace or one label selector. For volume data, Velero can use **CSI VolumeSnapshots** (with its CSI support enabled), a provider snapshot plugin, or a file-system copy. A `Restore` object replays a backup into the same cluster or a different one. ## Restore ordering Objects have dependencies: a Pod that mounts a PVC or a Secret stalls in `ContainerCreating` until they exist, and custom resources are rejected until their CRD is registered. Velero's default **restore resource priorities** handle this with a high-priority list, restored in order: 1. `customresourcedefinitions`, `namespaces`, `storageclasses` 2. `volumesnapshotclass`, `volumesnapshotcontents`, `volumesnapshots` (the `snapshot.storage.k8s.io` group), plus Velero's `datauploads` 3. `persistentvolumes`, `persistentvolumeclaims` 4. `clusterroles`, `roles`, `serviceaccounts`, `clusterrolebindings`, `rolebindings` 5. `secrets`, `configmaps`, `limitranges`, `priorityclasses` 6. `pods`, `replicasets.apps`, then `endpoints` and `services` Everything else follows, and a short low-priority list comes last. Operators can override the order in Velero's server configuration. Some kinds are **never restored**: `nodes`, `events`, and Velero's own `backups` and `restores`. An object that already exists in the target is skipped with a warning when it differs, unless `existingResourcePolicy: update` asks Velero to patch it. ## How PVCs get volumes back For each backed-up PV, Velero checks how its data was protected: | Backup method | What happens to the PV object on restore | |---|---| | CSI VolumeSnapshot (or its data-mover upload) | skipped; the volume is dynamically re-provisioned | | file-system backup | skipped; re-provisioned, then data is copied into it | | provider snapshot plugin | a new volume is created from the snapshot and the PV is rewritten to point at it | | no backup, reclaim policy `Delete` | skipped; re-provisioned **empty** | On the CSI path, the chain is: - Velero recreates the **VolumeSnapshotContent** from the backup, with `deletionPolicy: Retain` and a source that references the provider's existing **snapshot handle**; - it recreates the **VolumeSnapshot** that binds to that content; - a restore action rewrites the **PVC**: `spec.volumeName` is cleared (the old PV no longer exists) and `spec.dataSource` is set to the VolumeSnapshot; - the CSI external provisioner sees a pending PVC with a snapshot data source and creates a **new volume** from it, along with a new PV that binds to the PVC. With a `WaitForFirstConsumer` StorageClass, this happens only once a Pod using the claim is scheduled. For a PDF-invoice renderer that keeps its fonts and templates on a PVC, the restored Pod therefore mounts a new volume holding snapshot-time data, under a new PV name. ```yaml apiVersion: velero.io/v1 kind: Restore metadata: name: invoices-restore namespace: velero spec: backupName: invoices-nightly includedNamespaces: - invoices restorePVs: true existingResourcePolicy: none ``` ## Where it bites in a replacement cluster - **Snapshots stay with the storage backend.** Without data movement (`snapshotMoveData` on the Backup), the snapshot exists only where the original disk lived. If that backend or region is gone, the objects restore but the PVCs cannot be provisioned. Data movement uploads snapshot contents to object storage, and the restore downloads them into a new volume. - **Classes and drivers must exist.** The target needs a StorageClass and a `VolumeSnapshotClass` for the same CSI driver. Velero selects the snapshot class at backup time, and one labelled `velero.io/csi-volumesnapshot-class` is its signal. - **Namespace mapping** can restore into a new namespace name for a side-by-side check. - **Backups expire.** The default backup TTL is 30 days (720h), after which the backup and its snapshots are garbage-collected.
- What changes if the backup was taken with snapshotMoveData enabled?At backup time, Velero's data mover copies each CSI snapshot's contents into object storage, tracked by DataUpload objects, so the data no longer depends on the original storage backend. On restore, a DataDownload writes that data into a newly provisioned volume. This is slower, but it survives the loss of a region or a storage system, and it lets you restore to a different CSI driver.
- Why do restored PVs get new names, and does anything care?On the CSI path the PV is dynamically provisioned, so the provisioner names it and gives it a new volume handle. Workloads refer to PVCs, not PVs, so they are unaffected. Anything that pinned a PV name or a disk ID, such as a backup job, a monitoring label or an external runbook, has to be updated.
saying these in an interview costs you the question
- Velero writes directly into etcd, so it restores the whole cluster at once
- A restored PVC binds back to the original PV object by its old name
- CSI snapshots are always copied into Velero's object storage
- Velero restores Node objects so Pods land on the same machines
- The API server rejects a Pod whose Secret is not restored yet