After losing a Kubernetes control plane, you restore last night's etcd snapshot into a replacement one. What breaks once kube-apiserver serves that state, and how do you handle it?
answer
- old identities, new machines
- CA and sa.key must match
- stop API servers, restart watchers
- ghost Nodes block rescheduling
- time travel orphans disks
basics
~20 sRestored state is old and bound to old identities. Reuse the original PKI and encryption keys, stop every API server before restoring, delete Node objects for vanished machines, restart controllers and kubelets, then reconcile drift caused by the rewind.
solid answer
~40 sThe snapshot is a point in time, and the cluster around it has moved on. First, **identity**: kubelet client certificates, kubeconfigs and ServiceAccount tokens were signed by the old CA and `sa.key`. Restore the original `/etc/kubernetes/pki` and the `--encryption-provider-config` file, or re-issue every credential. Second, **stale caches**: stop all `kube-apiserver` instances before restoring, then restart `kube-controller-manager`, `kube-scheduler`, kubelets and operators so they re-list instead of trusting newer cached state. Third, **ghost Nodes**: Node objects for machines that no longer exist go unreachable. Deleting them lets the PodGC controller clear their Pods so controllers reschedule. Fourth, **time travel**: objects created since the snapshot are gone, deleted ones reappear, PVs may point at deleted disks, and cloud resources created later are orphaned. Reconcile from Git and audit before reopening the cluster.
go deeper
Remember that a restored cluster is a snapshot of the past, and that the certificates and keys from the old control plane are needed for things to reconnect.
Be able to say why the PKI and the encryption config must be restored with the snapshot, and why Node objects for machines that no longer exist need manual cleanup.
Walk through the runbook in order: freeze change, identity, stop API servers, restore, restart watchers, ghost Nodes, forward reconcile. Name the drift classes you audit afterwards.
Discuss when a whole-cluster rewind is the wrong tool compared with rebuilding and re-applying, and how quorum-loss recovery time drives the snapshot cadence you mandate.
## Why a restore is not the end of the incident An **etcd snapshot** is the cluster store frozen at one moment. Restoring it into a replacement control plane means `kube-apiserver` now serves a past version of the world, while nodes, disks, cloud resources and people have all kept moving. On a 140-node cluster shared by 22 product teams, a snapshot 7 hours 34 minutes old can hide dozens of deploys. The etcdctl mechanics of the restore itself are a separate topic; this is about what happens after it. ## 1. Identity: the PKI must match Everything that talks to the API server proves its identity with material issued by the old cluster: - **kubelet client certificates** and every kubeconfig are signed by the old **cluster CA**; - **ServiceAccount tokens**, both projected and legacy, are signed with the old `sa.key` and verified against `sa.pub`; - the aggregation layer relies on the **front-proxy CA**, and etcd peers on the **etcd CA**; - Secrets encrypted at rest need the original **encryption configuration**. If `kubeadm init` generated fresh PKI on the new control plane, none of that is trusted: kubelets cannot re-authenticate and in-cluster clients get authentication errors. The runbook therefore puts the backed-up `/etc/kubernetes/pki` and the encryption config in place **before** the control plane starts, or it plans to re-issue every credential. ## 2. Stale caches and a rewound resourceVersion Controllers, kubelets and operators watch the API server through informers that cache objects. After a restore, `resourceVersion` values go backwards, and a surviving client may hold objects newer than anything in the store. The Kubernetes guidance is to: 1. stop **every** `kube-apiserver` instance before restoring, so nothing writes into a half-restored store; 2. restore the snapshot into all etcd members; 3. start the API servers; 4. restart `kube-controller-manager`, `kube-scheduler`, kubelets and in-cluster controllers, so each one re-lists instead of acting on cached state. ## 3. Ghost Node objects The snapshot contains `Node` objects for machines that may no longer exist. - The node lifecycle controller sees no heartbeats and taints them `node.kubernetes.io/unreachable`. - Pods tolerate that taint for **300 seconds** by default and are then evicted, but their deletion never completes, because no kubelet is left to confirm it. - **Deleting the Node objects** for vanished machines lets the **PodGC controller** delete the orphaned Pods after a short quarantine. The owning ReplicaSets and StatefulSets then create replacements on live nodes. ## 4. Time travel | Since the snapshot, someone... | After the restore | |---|---| | deployed a new PDF-invoice renderer version | the old version runs again | | deleted a namespace | it reappears, and its workloads restart | | created a PVC (a disk was provisioned) | the PV object is gone and the disk is orphaned | | deleted a PVC (the disk was deleted) | the PV object points at a missing disk, so attach fails | | created a LoadBalancer Service | the cloud load balancer is orphaned | The fix is **forward reconciliation**. A GitOps controller re-applies what was merged after the snapshot. You audit PVs against the storage backend, and you clean up orphaned cloud resources. Stateful workloads whose data moved on need their own decision: a database restored from its own backup may no longer match the rewound objects. ## Quorum loss is the common trigger Often the control plane is not destroyed at all; etcd has simply lost quorum for good, for example when two of three members' disks are gone. The survivor cannot accept writes on its own. Recovery means building a **new** etcd cluster from the latest snapshot (or from the survivor's data) as a single member, pointing the API servers at it, and then adding members back one at a time. After that, every step above still applies. ## A runbook skeleton 1. Freeze change: pause deploy pipelines and tell the 22 teams. 2. Restore PKI and encryption config, stop API servers, restore etcd, start the control plane. 3. Restart controllers, schedulers, kubelets and operators. 4. Delete ghost Nodes and watch Pods reschedule. 5. Reconcile forward from Git and audit PVs and cloud resources. 6. Reopen, and record what drifted for the next drill.
- Two of three etcd members are permanently gone and one survives. Do you still need a snapshot?The survivor cannot make progress without a majority, so you rebuild. The documented path is to start a new cluster from the latest snapshot, or from the survivor's data, as a single member. Then point the API servers at it and add fresh members one at a time. A snapshot is the safer input when the survivor's data is suspect, and taking one regularly is what makes this path fast.
- Why stop every kube-apiserver before restoring, rather than restoring under a running control plane?A running API server keeps writing: leases, status updates, and events from controllers acting on their caches. Those writes can land in a store that is being replaced, or conflict with the restored data, leaving the cluster in a mixed state. Stopping all instances makes the restore atomic from the clients' point of view.
- After the restore, does a GitOps controller help or hurt?Mostly it helps: it re-applies everything merged after the snapshot, which undoes most of the rewind. It can hurt if its own configuration was rewound, or if pruning is on and the restored state contains objects Git no longer lists. Those get deleted, which may be correct but should be expected. Check its sync scope before letting it run.
saying these in an interview costs you the question
- Let kubeadm generate fresh certificates; the restored cluster will just work
- Once the API server is up, the restore is finished
- Old Node objects clean themselves up, so leave them
- Controllers need no restart because they read directly from etcd
- Restoring a snapshot only adds missing objects and never removes newer ones