skip to content

How do you bootstrap a highly available Kubernetes control plane with kubeadm, and what must be decided before the first kubeadm init?

level: seniorimportance: should knowfreq 54%

answer

  1. decide the address first
  2. load balancer in the certificate
  3. encrypted Secret, short lifetime
  4. local versus external etcd config
  5. one control-plane node at a time

basics

~20 s

Put a load balancer or DNS name in front of the future API servers, pass it as --control-plane-endpoint to the first kubeadm init, share certificates with --upload-certs, then run kubeadm join --control-plane on the others; choose stacked or external etcd up front.

solid answer

~40 s

The decisions that are hard to change later come first: a stable **`controlPlaneEndpoint`** (load balancer or DNS name, which also goes into the API server certificate), the **pod and service CIDRs**, and whether etcd is **stacked** (kubeadm runs an etcd static pod on every control-plane node) or **external** (`etcd.external` with endpoints and client certs). Then run `kubeadm init --control-plane-endpoint ... --upload-certs` on the first node; kubeadm encrypts the shared CA, service-account and front-proxy keys into the `kubeadm-certs` Secret, readable with a certificate key for **two hours**. Install the CNI, then add the other control-plane nodes one at a time with `kubeadm join ... --control-plane --certificate-key <key>`, then workers. Without an endpoint, kubeadm refuses the second control-plane join.

code

yaml · 19 lines
yaml
apiVersion: kubeadm.k8s.io/v1beta4
kind: ClusterConfiguration
kubernetesVersion: v1.37.0
controlPlaneEndpoint: k8s-api.games.internal:6443
apiServer:
  certSANs:
    - k8s-api.games.internal
networking:
  podSubnet: 10.244.0.0/16
  serviceSubnet: 10.96.0.0/12
etcd:
  external:
    endpoints:
      - https://10.40.7.11:2379
      - https://10.40.7.12:2379
      - https://10.40.7.13:2379
    caFile: /etc/kubernetes/pki/etcd/ca.crt
    certFile: /etc/kubernetes/pki/apiserver-etcd-client.crt
    keyFile: /etc/kubernetes/pki/apiserver-etcd-client.key

go deeper

for a junior

Recall that HA needs several control-plane nodes behind one address, and that kubeadm join has a --control-plane flag for adding them.

for a middle

Explain how --upload-certs and the certificate key move the shared CA and service-account keys to new control-plane nodes, and why they expire.

for a senior

Show which settings are locked in at first init, how you sequence joins with stacked etcd, and how you refresh an expired key safely.

for a principal

Argue stacked versus external etcd in terms of machine count, failure domains and who operates etcd, and decide when self-managed HA is worth running at all.

## What makes kubeadm HA different A single-node control plane is a single point of failure. A highly available (HA) control plane runs several API servers, each with its own controller manager and scheduler, plus an etcd cluster sized for quorum. The quorum maths and leader election are architecture topics; here the question is **how kubeadm builds it**, and which choices are locked in by the very first `kubeadm init`. Take **an 80-node cluster that autoscales between 20 and 80 nodes** running **a multiplayer game-session backend**. Every one of those workers talks to the API server; if the address they were given points at one machine, that machine can never be retired. ## Decisions to make before the first init 1. **A stable control-plane endpoint.** A load balancer or DNS name, for example `k8s-api.games.internal:6443`, passed as `--control-plane-endpoint` or `controlPlaneEndpoint` in `ClusterConfiguration`. It is written into every kubeconfig and into the API server certificate. kubeadm's join preflight fails with *"unable to add a new control plane instance to a cluster that doesn't have a stable controlPlaneEndpoint address"* when it was never set. 2. **Extra certificate names.** Any other names clients will use go into `apiServer.certSANs` (flag `--apiserver-cert-extra-sans`). 3. **Network ranges.** `networking.podSubnet` (`--pod-network-cidr`) must match what the CNI plugin expects, and `networking.serviceSubnet` (`--service-cidr`) must not overlap your VPC. Both are painful to change later. 4. **etcd topology.** Stacked or external; see the table below. 5. **The load balancer health check.** It should probe the API server port and pull a node out while it is down, including during the first init when only one backend exists. ## Stacked versus external etcd, from kubeadm's side | Aspect | Stacked etcd | External etcd | |---|---|---| | Where etcd runs | static pod on each control-plane node | separate hosts you run | | kubeadm config | `etcd.local` (the default) | `etcd.external` with `endpoints`, `caFile`, `certFile`, `keyFile` | | Who creates members | kubeadm, during `init` and the `etcd-join` phase of `join` | you, before `kubeadm init` | | Certificates shared on join | cluster CA, SA keys, front-proxy CA, etcd CA | cluster CA, SA keys, front-proxy CA, etcd client cert and key | | Machines needed | fewer | more | With external etcd, kubeadm writes the endpoints into the API server's `--etcd-servers` flag and creates no etcd static pod at all. Why you would pay for extra machines is a topology judgment for the HA design itself. ## The bootstrap sequence 1. Create the load balancer with all future control-plane machines as backends. 2. On the first control-plane node, run `kubeadm init --config cluster.yaml --upload-certs`. 3. Install the CNI plugin, so nodes can become Ready. 4. On each additional control-plane node, run the printed `kubeadm join <endpoint> --token ... --discovery-token-ca-cert-hash ... --control-plane --certificate-key <key>`, **one node at a time**, waiting for each to be healthy. 5. Join workers with the plain `kubeadm join` command. ## How certificate sharing works Every control-plane node must hold the **same** cluster CA, service-account signing key pair and front-proxy CA, or tokens and certificates signed on one node would be rejected on another. - `--upload-certs` encrypts those files with a random 32-byte **certificate key** and stores them in the `kubeadm-certs` Secret in `kube-system`. - `kubeadm join --control-plane --certificate-key <key>` downloads and decrypts them in its `control-plane-prepare` phase, then generates the node's own leaf certificates locally. - The Secret and key are short-lived: **two hours** by default. To add a control-plane node later, run `kubeadm init phase upload-certs --upload-certs` for a new key, then `kubeadm token create --print-join-command --certificate-key <key>`. - The alternative is copying the files by hand to `/etc/kubernetes/pki` on each new node before joining without `--certificate-key`. The certificate key decrypts the cluster CA private key, so treat it with at least the care you give the bootstrap token. ## Common mistakes - Initialising with a node IP and trying to add HA later. - Joining two control-plane nodes simultaneously with stacked etcd, which grows the etcd member list while a new member is still syncing. - Reusing a join command from yesterday: the token and the certificate key have both expired. - Letting the load balancer send traffic to a node whose API server is not ready yet.

  • You set up a single control-plane cluster on a node IP months ago. What does converting it to HA involve?
    kubeadm will not add control-plane nodes without `controlPlaneEndpoint`. You must introduce a load balancer or DNS name, add it to the stored cluster configuration, reissue the API server certificate with the new SAN, and update every kubeconfig, including kubelets, to use it. It is doable but invasive, which is why the endpoint is decided before init.
  • Why join stacked-etcd control-plane nodes one at a time?
    Each `kubeadm join --control-plane` adds an etcd member. A new member counts toward quorum as soon as it is added, before it has synced. Adding two at once from a single-member cluster can leave quorum depending on members that are not ready, so you wait for each node's etcd and API server to be healthy first.

saying these in an interview costs you the question

  • controlPlaneEndpoint can be added later with no certificate or kubeconfig changes
  • The kubeadm-certs Secret stays readable forever after kubeadm init
  • Each control-plane node should generate its own cluster CA
  • With external etcd, kubeadm still runs an etcd static pod
  • The certificate key is harmless to paste into shared chat
  • All control-plane nodes can be joined in parallel with stacked etcd