skip to content

How do you take a backup of the Kubernetes cluster store with etcdctl, and what is the actual procedure to restore a cluster from that backup?

level: seniorimportance: must knowfreq 52%

answer

  1. snapshot save + snapshot status to verify
  2. store off-box, encrypted; cadence = RPO
  3. stop apiserver, move data dir aside, restore
  4. restore = new cluster ID, single member first
  5. re-add members one at a time; reconcile PVs after

basics

~20 s

Back up with etcdctl snapshot save (v3 API, with client TLS certs), on a schedule, stored off-box. To restore: stop the API servers, stop etcd and move the old data directory aside, run etcdctl snapshot restore into a fresh directory, start etcd as a new single-member cluster, restart the control plane, then re-add members.

solid answer

~60 s

**Backup:** `ETCDCTL_API=3 etcdctl snapshot save snap.db` against a member endpoint with the CA, client cert and key. It is a consistent point-in-time copy of the whole keyspace. Run it on a schedule, verify with `snapshot status`, and ship it off the control-plane nodes — a backup on the machine you are recovering from is not a backup. Snapshot cadence *is* your RPO. **Restore:** 1. Stop kube-apiserver on all control-plane nodes (with kubeadm, move the static-pod manifest aside) so nothing writes during recovery. 2. Stop etcd and move each data directory aside — never restore over it. 3. `etcdctl snapshot restore snap.db --data-dir /var/lib/etcd-restored`, which builds a brand-new data directory with a new cluster ID and member ID. 4. Start etcd from that directory as a single-member cluster, then start the control plane. 5. Re-add the other members one at a time (`member add`, start with empty data dirs and `--initial-cluster-state existing`). Everything written after the snapshot is lost, so afterwards reconcile: PersistentVolumes and external state did not roll back with etcd.

code

bash · 16 lines
bash
# backup
ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-$(date +%F-%H%M).db \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key

# verify
etcdctl snapshot status /backup/etcd-2026-01-01-0300.db -w table

# restore into a NEW data dir (etcd and kube-apiserver stopped)
etcdctl snapshot restore /backup/etcd-2026-01-01-0300.db \
  --data-dir /var/lib/etcd \
  --name m1 \
  --initial-cluster m1=https://10.0.0.1:2380 \
  --initial-advertise-peer-urls https://10.0.0.1:2380

go deeper

for a junior

Know the two commands — etcdctl snapshot save and etcdctl snapshot restore — and that etcd is what you back up to protect a cluster.

for a middle

Sequence the restore correctly: stop the API server, move the old data dir aside, restore to a new dir, start a single member, then re-add members.

for a senior

Own the operational envelope: cadence as RPO, off-box encrypted storage, verification, and post-restore reconciliation of PVs, orphan Pods and external resources.

for a principal

Talk about RPO/RTO targets, drill cadence, GitOps re-apply as the complement to etcd restore, snapshot secrecy (it contains all Secrets), and when rebuilding the cluster beats restoring it.

## Why this is the whole DR story etcd is the only stateful control-plane component. API servers, schedulers and controller-managers are stateless and can be reinstalled from configuration. So "restore the cluster" means "restore etcd". With no snapshot and unrecoverable quorum, the objects are gone — you rebuild the cluster and re-apply manifests from source control, which is why GitOps and etcd backups are complements, not substitutes. ## Taking the snapshot `etcdctl snapshot save` asks a member to stream a consistent copy of its backing database at a single revision. It is safe on a live member and does not stop traffic, though it reads the whole database, so it costs I/O proportional to database size. What matters in practice: - **Use the v3 API** (`ETCDCTL_API=3`; the default in etcd 3.4+) and pass `--endpoints` plus `--cacert/--cert/--key`, because etcd is mTLS-protected. - **Verify** with `etcdctl snapshot status snap.db -w table`, which prints hash, revision, key count and size. A snapshot you have never validated is a hope, not a backup. - **Store it off the node**, ideally in versioned object storage, encrypted — the snapshot contains every Secret in the cluster in whatever form etcd holds it. - **Automate the cadence.** Every 15–30 minutes is common on busy clusters; the interval is exactly how much work you are prepared to lose. - Managed control planes (EKS/GKE/AKS) do this for you and do not expose etcd; this question is about self-managed and kubeadm-style clusters. ## Restoring, step by step Restore is intrusive and must be done with the control plane quiesced, otherwise controllers write into a half-restored store. 1. **Freeze writers.** Stop kube-apiserver on every control-plane node. With kubeadm, `mv /etc/kubernetes/manifests/kube-apiserver.yaml /tmp/` makes the kubelet tear the static pod down. Do the same for controller-manager and scheduler. 2. **Stop etcd** on all members and move each data directory aside (`mv /var/lib/etcd /var/lib/etcd.bak`). Keep it — if the disks were fine, restarting the original members may still be the better recovery. 3. **Restore on one node:** `etcdctl snapshot restore snap.db --data-dir /var/lib/etcd --name m1 --initial-cluster m1=https://10.0.0.1:2380 --initial-advertise-peer-urls https://10.0.0.1:2380`. This is a local, offline operation that builds a fresh data directory and assigns a **new cluster ID** — deliberate, so a stale member of the old cluster cannot rejoin and mix histories. 4. **Start etcd** as that single-member cluster and confirm `endpoint status` shows a leader and the expected revision. 5. **Start the control plane** (restore the static-pod manifests) and sanity-check with `kubectl get nodes,pods -A`. 6. **Rebuild quorum.** For each remaining member: `etcdctl member add` (learner first is safer), then start that etcd with an empty data dir and `--initial-cluster-state existing`. One at a time, waiting for sync, so quorum is never at risk. Alternatively restore the same snapshot on every node with matching `--initial-cluster` flags and bring them up together. ## After the restore: reconcile the world Rolling etcd back in time does **not** roll back anything outside it: - Pods deleted since the snapshot may be recreated; Pods created since vanish from the API while their containers may still run on nodes until the kubelet resyncs and kills the orphans. - PersistentVolume objects return, but the underlying disks, their data and cloud-side attachments are wherever they are now — mismatches surface as stuck attach/detach. - External systems (DNS records, load balancers, cloud resources created by controllers) reconcile against older state and may thrash. - Certificates and tokens issued after the snapshot are unknown to the restored cluster. So the runbook ends not at "etcd is up" but at "current manifests re-applied from Git, nodes Ready, workloads and volumes verified". ## Practice it The common failure is not the commands — it is discovering mid-outage that the snapshot was empty, unreadable, sitting on the dead node, or taken with the wrong API version. Restore drills into a scratch cluster are the only way to know your RPO and RTO numbers are real.

  • After restoring etcd from a snapshot taken an hour ago, what inconsistencies should you expect and how do you resolve them?
    Anything written in that hour is gone: new Deployments, scaled replica counts, new Secrets and tokens; objects deleted in that window come back. Pods created after the snapshot may still be running on nodes as orphans until the kubelet resyncs and removes them. PersistentVolume objects are restored but the real disks and attachments are not, so expect stuck attach/detach. The fix is re-applying current manifests from source control, then verifying nodes, workloads and volumes.
  • Why does snapshot restore create a new cluster ID, and why must you not restore into the existing data directory?
    The new cluster ID guarantees a surviving member of the old cluster cannot silently rejoin and merge two divergent Raft histories, which would corrupt state. Restoring into the existing directory would destroy your ability to fall back to the original data if the restore turns out to be the wrong call — and healthy original members are often the better recovery path.

saying these in an interview costs you the question

  • Copying /var/lib/etcd with cp or tar on a running member and calling it a backup
  • Keeping snapshots only on the control-plane node that produced them
  • Restoring while kube-apiserver is still running
  • Assuming a restore also rolls back PersistentVolume data or cloud resources
  • Never testing a restore, so the first attempt happens during an incident

context