skip to content

After editing kube-apiserver.yaml in /etc/kubernetes/manifests on a kubeadm control-plane node, kubectl stops responding. How do you diagnose and recover?

level: seniorimportance: nice to knowfreq 34%

answer

  1. kubectl needs the thing you broke
  2. kubelet watches a directory
  3. container runtime CLI on the node
  4. hidden files are skipped
  5. mirror pods only reflect

basics

~20 s

The kubelet restarts the API server from the edited static pod manifest, so kubectl is useless; debug on the node with crictl ps -a, crictl logs and the kubelet journal, fix the file or restore a backup kept outside the manifests directory.

solid answer

~40 s

Control-plane components on a kubeadm node are **static pods**: the kubelet watches `/etc/kubernetes/manifests` and recreates a component whenever its file changes. A bad flag or YAML error means the API server crash-loops or never starts, and every `kubectl` call fails because kubectl talks to that very API server. So I work on the node: `crictl ps -a --name kube-apiserver` to see restarts, `crictl logs <id>` for the flag or certificate error, and `journalctl -u kubelet` for manifest parse errors. I fix the file, or restore a copy I saved **outside** the directory, because the kubelet loads every non-hidden file there, `.bak` included. Editing the mirror pod through the API never changes the running component.

code

bash · 3 lines
bash
sudo crictl ps -a --name kube-apiserver
sudo crictl logs <container-id> 2>&1 | tail -n 20
sudo journalctl -u kubelet --since '10 min ago' | grep -i apiserver

go deeper

for a junior

Recall that control-plane components on kubeadm nodes are files in /etc/kubernetes/manifests that the kubelet runs.

for a middle

Explain how the kubelet reacts to manifest changes and why a mirror pod cannot be used to change the component.

for a senior

Demonstrate node-level debugging with crictl and the kubelet journal, safe backups, and rolling one control-plane node at a time.

for a principal

Push control-plane changes through reviewed kubeadm configuration or node automation, so a single typo cannot take down every API server.

## Why kubectl goes silent On a kubeadm-built node, the API server, controller manager, scheduler and (with stacked etcd) etcd run as **static pods**. Their definitions are plain files in `/etc/kubernetes/manifests`, the kubelet's `staticPodPath`. The kubelet watches that directory and also re-reads it periodically (`fileCheckFrequency`, 20 seconds by default). When a file changes, the kubelet stops the old container and starts one from the new spec. `kubectl` is only a client of kube-apiserver. If the edit breaks the API server, there is nothing for kubectl to talk to: expect `connection refused` on port 6443 through the load balancer if every control-plane node was edited, or intermittent errors if only one was. Picture the operators of **a multiplayer game-session backend** on **an 80-node cluster that autoscales between 20 and 80 nodes**. An engineer adds an audit flag to the API server on all three control-plane nodes at once, mistypes a file path, and the cluster's API disappears. Running game sessions keep going, because kubelets keep existing containers alive, but nothing new can be scheduled, and the node autoscaler cannot add capacity. ## Diagnosis on the node Work from the machine, not the API: 1. **Is the container running at all?** `crictl ps -a --name kube-apiserver` shows the container, its state and attempt count. A rising attempt count means a crash loop. 2. **Why did it exit?** `crictl logs <container-id>` usually names the problem: an unknown flag, a missing certificate file, an unreachable etcd endpoint. 3. **Did the kubelet even accept the file?** If no container exists, read `journalctl -u kubelet`. A YAML syntax error or an invalid pod spec is reported there and the pod is not started. 4. **Is a mounted path missing?** New flags that reference files need matching `hostPath` volumes and `volumeMounts` in the same manifest; a flag pointing at a path that is not mounted fails inside the container even though the file exists on the host. | Symptom | Where to look | Typical cause | |---|---|---| | Container restarting | `crictl logs` | bad flag value, missing cert, etcd unreachable | | No container at all | kubelet journal | YAML error, invalid spec | | Two API server containers or odd behaviour | `ls -a` of the manifests dir | a backup file left in the directory | | Change seems ignored | manifest file vs mirror pod | edit made through `kubectl edit` | ## Recovery - **Fix the file in place** and wait for the kubelet to pick it up, or restore the previous version. - **Keep backups outside the directory.** The kubelet reads every file in `staticPodPath` whose name does not start with a dot, regardless of extension, so `kube-apiserver.yaml.bak` is parsed as another static pod definition and can clash with the edited one. - **Edit one control-plane node at a time**, confirm its API server is healthy behind the load balancer, then move on. With several API servers, a mistake on one node leaves the others serving. - To force a restart without changing the spec, move the file out of the directory, wait for the container to stop, and move it back. ## Why the API route does not work For every static pod the kubelet creates a **mirror pod** in the API, marked with the `kubernetes.io/config.mirror` annotation, named after the component plus the node name, such as `kube-apiserver-cp-2`. It exists so you can see the component with `kubectl get pods -n kube-system`. It is a read-only reflection: - changing it through the API does not change the file, so the running container is unaffected; - deleting it only makes the kubelet recreate it; - the source of truth stays the file on disk. ## Making the change durable A hand edit on a kubeadm node can be overwritten later, because `kubeadm upgrade` regenerates control-plane manifests from the stored cluster configuration. Put the flag into `ClusterConfiguration` (for example `apiServer.extraArgs`) as well, so the next regeneration keeps it. That is also the safer place to review the change before it reaches any node. ## What interviewers listen for The strong answer names the circular dependency (kubectl needs the component you broke), switches to node-level tools without hesitation, and mentions both the backup-file trap and the mirror pod. The weak answer keeps retrying `kubectl` or tries to delete the kube-apiserver pod.

  • Why does a flag change you made by hand disappear after a kubeadm upgrade?
    `kubeadm upgrade` regenerates the control-plane static pod manifests from the cluster configuration kubeadm stored at init. Flags added only to the file are not in that configuration, so they are lost. Adding them to `ClusterConfiguration`, for example under `apiServer.extraArgs`, makes them survive regeneration.
  • Your new API server flag points at a file that exists on the host, yet the container reports it missing. Why?
    The container sees only paths mounted into it. kubeadm's manifest mounts specific host directories; a file elsewhere needs a matching `hostPath` volume and `volumeMount` added to the same static pod manifest.

saying these in an interview costs you the question

  • Use kubectl edit on the kube-apiserver pod to fix the flag
  • Deleting the kube-apiserver mirror pod restarts it with the fixed spec
  • Saving kube-apiserver.yaml.bak next to the original is a safe backup
  • Running workloads stop the moment the API server goes down
  • Editing all control-plane nodes at once is fine because the kubelet validates flags