You own disaster recovery for a 140-node Kubernetes cluster shared by 22 product teams. How do you split recovery between etcd snapshots, Velero and Git re-apply, and prove it works?
answer
- failure scenario picks the tool
- namespace scope versus cluster rewind
- tiers map to schedules
- alert on newest snapshot age
- last successful restore, not backup
basics
~20 sMatch each failure to a recovery unit: rebuild plus Git re-apply for stateless workloads, Velero with CSI snapshots for namespaces with data or objects missing from Git, and etcd snapshots for rewinding the whole control plane. Prove it with timed restore drills.
solid answer
~40 sStart from failure scenarios, not tools. **Lost cluster or region**: rebuild from infrastructure code and let a GitOps controller re-apply intent, since stateless teams need nothing more. **One team's namespace destroyed or corrupted**: use Velero, because it restores one namespace, and its CSI snapshots (with data movement for critical tiers) bring back volumes. **Control plane or etcd quorum lost on otherwise healthy machines**: restore an etcd snapshot, the fastest exact recovery, but it rewinds all 22 teams at once. Snapshot etcd frequently and store it off-cluster with the PKI and encryption keys. Let teams declare an RPO tier that maps to Velero schedules. Then **prove it**: run restores regularly into a scratch cluster, time them against the RTO, and track the last *successful restore*, not the last successful backup.
go deeper
Remember that different failures need different recovery tools, and that a backup nobody has restored is unproven.
Explain why an etcd snapshot rewinds the whole cluster, while Velero can restore one namespace and Git can rebuild stateless apps.
Show how you would run a timed restore drill and which drift it typically uncovers: missing snapshot classes, Secrets absent from Git, snapshots in the wrong place.
Own the tiering and ownership model, the cost of off-backend copies and scratch clusters, and when an unacceptable cluster-wide rewind argues for splitting clusters.
## Start from failures, not tools On a **140-node cluster shared by 22 product teams**, three recovery mechanisms are available, and each one answers a different failure. A strategy that picks one tool for everything either rewinds 21 innocent teams or leaves one team's data behind. | Failure | Best recovery unit | Why | |---|---|---| | Region or whole cluster gone | new cluster from infrastructure code, plus Git re-apply, plus Velero for stateful namespaces | no stale runtime state; scales per team | | One namespace deleted or corrupted | **Velero** restore of that namespace | scoped; the other 21 teams are untouched | | etcd quorum lost or control plane corrupted, nodes healthy | **etcd snapshot** restore | exact and fast, but cluster-wide and rewinds everything | | Bad change to shared cluster-scoped objects (CRDs, webhooks) | Git revert first, snapshot as a last resort | a rewind is heavier than the fault | ## Designing each layer **etcd snapshots** belong to the platform team. - Cadence follows the RPO you would accept for a **cluster-wide** rewind. Snapshots are cheap, so take them often (every 20 minutes, say) and keep a short hourly and daily ladder. - Store them off-cluster and encrypted, together with the matching `/etc/kubernetes/pki` and the `--encryption-provider-config` file, under separate access control, because the bundle unlocks every Secret. - Alert on the **age of the newest snapshot** rather than on a job succeeding, since a silent scheduler stops alerting on failures. **Velero** is shared, with a per-team contract. - Teams label namespaces with a recovery tier, and each tier maps to a `Schedule` with a matching TTL. - Tiers that must survive a storage or region loss get snapshot **data movement**. Cheaper tiers accept in-backend CSI snapshots. - Volumes whose engine needs consistency, such as databases, need application-level backups too. That is the data teams' responsibility, not cluster state. **Git** is the default for stateless workloads such as a PDF-invoice renderer. The platform's job is to make "new cluster, point the GitOps controller at it" a routine operation, and to hunt down objects that exist only in the cluster. ## Proving it works A green backup job proves that bytes were written somewhere. It proves nothing about recovery. Evidence comes from restores: 1. **Scheduled drills** into a scratch cluster: restore the latest etcd snapshot with its PKI, and restore a rotating sample of team namespaces with Velero. 2. **Time every phase** (control plane up, Pods scheduled, PVCs bound, application checks passing) and compare the total against each tier's RTO. 3. **Run application checks**, not just `kubectl get pods`: the invoice renderer must render an invoice from restored templates. 4. **Record the drift** each drill finds: objects missing from Git, Secrets that exist nowhere else, snapshots in the wrong region. 5. **Publish the last successful restore per tier** as the metric leadership sees. A real example of what drills surface: a drill restored 19 of 22 namespaces within target. Two failed because their StorageClass had no matching `VolumeSnapshotClass` in the scratch cluster, and one because its Secrets came from an operator that was never backed up. Each finding became a fix in the platform, not a note in a runbook. ## Tradeoffs to own openly - **Rebuild versus restore.** Rebuilding avoids stale state and ghost Nodes but needs mature infrastructure code and Git hygiene. Restoring is faster to reason about but carries the past with it. - **Cost.** Off-backend data movement and a standing scratch cluster cost money. Tiering spends it where the business impact justifies it. - **Ownership.** The platform owns the tooling and the cluster-wide rewind decision. Teams own their tier choice and their data-engine backups. The incident process decides when a rewind is worth the collateral. - **Blast radius as a design input.** If a cluster-wide rewind is unacceptable for some teams, that argues for splitting them onto separate clusters.
- How do you choose the etcd snapshot interval?Work backwards from the worst acceptable rewind for the whole cluster: every change made after the last snapshot is lost or must be re-applied. Snapshots are small and cheap, so the limit is usually storage retention and the load a snapshot puts on etcd, not cost. Pair the interval with an alert on snapshot age, so a stalled job pages someone well before the RPO is breached.
- A team asks for a 5-minute RPO on its namespace. What do you tell them?Scheduled Velero backups and snapshots are poor at minute-level RPOs for data. That need is usually met by the application's own replication or continuous backup. The platform can offer frequent object backups and a tier with data movement, but a tight data RPO belongs in the data engine's design, and the team should own it.
saying these in an interview costs you the question
- A green nightly backup job proves the cluster can be recovered
- An etcd snapshot is the right tool for one team's deleted namespace
- Everything is in Git, so no drills are needed
- Snapshots can live in the same bucket and account as the cluster credentials
- One backup schedule should serve all 22 teams equally