In an HA Kubernetes control plane, why do kubelets and clients reach kube-apiserver through a load balancer, and how should it check backends?
answer
- stateless servers, one address
- controlPlaneEndpoint chosen up front
- layer 4, no TLS termination
- readyz includes etcd readiness
- the default kubernetes Service
basics
~20 skube-apiserver replicas are stateless and all active, so clients need one stable address that survives any single server's loss. Put a TCP pass-through load balancer or virtual IP in front of them, and have it health-check each server's /readyz endpoint.
solid answer
~50 sUnlike kube-scheduler and kube-controller-manager, `kube-apiserver` runs **active-active**: all state lives in etcd, so any replica can serve any request. Kubelets, kube-proxy, CI and humans each have one server URL in their kubeconfig, so that URL must stay valid when a control-plane node dies. kubeadm calls this address `controlPlaneEndpoint`. On a bare-metal cluster with no cloud load balancer, that usually means a floating virtual IP plus a TCP proxy on the control-plane nodes, or a hardware load balancer. The load balancer should pass TLS through at layer 4, because client-certificate authentication needs the TLS handshake to reach the API server. It should health-check `/readyz`, which fails when the server's etcd connection is unhealthy and, with `--shutdown-delay-duration`, as soon as a graceful shutdown starts. Pods inside the cluster use the `kubernetes` Service instead, whose endpoints list every API server.
code
bash · 7 lines# Probe each API server the way the load balancer should
for ip in 10.40.0.21 10.40.0.22 10.40.0.23; do
curl -sk "https://${ip}:6443/readyz?verbose" | tail -n 2
done
# In-cluster view: one endpoint per live API server
kubectl get endpointslices -n default -l kubernetes.io/service-name=kubernetesgo deeper
Remember that API servers all serve at once, and clients reach them through one stable address that survives any single node failure.
Explain the load balancer's contract: layer-4 pass-through for client-certificate authentication, /readyz health checks, and the kubernetes Service path for Pods.
Cover the operational details: graceful drain with shutdown delay, skewed load after a failover, and why a TCP-only check sends traffic to servers that cannot reach etcd.
Judge the endpoint as part of the HA design: it must not become the new single point of failure, and on bare metal that choice shapes the whole control-plane build.
## Active-active, not active-passive An HA Kubernetes control plane runs its components in two different modes: - **kube-apiserver: active-active.** It keeps no authoritative state. Everything lives in etcd, so three replicas can all serve reads, writes and watches at the same time. - **kube-controller-manager and kube-scheduler: active-passive.** They elect a single leader through a Lease object, because two active copies would fight over the same objects. Active-active only helps if clients can reach whichever replica is alive. That is the load balancer's job. ## Why clients need one stable endpoint Every client finds the API server through a kubeconfig that holds **one** server URL. That includes kubelets, kube-proxy, a GitOps controller, CI jobs and engineers' laptops. If that URL names one control-plane node, its failure disconnects every client even though two healthy API servers remain. The fix is a stable address in front of all replicas. In kubeadm this is `ClusterConfiguration.controlPlaneEndpoint`, a DNS name or IP. It should be chosen before the cluster is built, because it is written into kubeconfigs and the API server's serving certificate. ## Building it without a cloud load balancer The shipment-tracking platform runs on a 9-node bare-metal cluster with no cloud load balancer. Common patterns: 1. **Floating virtual IP plus a local TCP proxy.** A VRRP-style daemon moves one IP between the three control-plane nodes. A TCP proxy on each node forwards port 6443 to all three API servers. 2. **A dedicated hardware or appliance load balancer** in front of the three nodes, if the data centre has one. 3. **A virtual IP managed from inside the cluster** by a static Pod or DaemonSet that announces the address. This is convenient, but it must work while the control plane is starting up. 4. **DNS round-robin** is the weak option. Clients cache answers and keep dialing a dead address until the record expires. Whichever you choose, the endpoint itself must not become a new single point of failure. A load balancer on one box moves the outage somewhere else. ## How the load balancer should behave - **Pass TLS through at layer 4.** kube-apiserver authenticates many clients, including kubelets, by client certificate. Terminating TLS at the load balancer would drop that certificate before it reaches the API server. - **Health-check `/readyz`, not just the TCP port.** A server can accept connections while it cannot serve. `/readyz` includes an `etcd-readiness` check, and it is readable without credentials by default through the `system:public-info-viewer` ClusterRole, as long as anonymous authentication stays enabled. - **Drain on shutdown.** With `--shutdown-delay-duration`, a terminating API server keeps serving normally while `/readyz` fails immediately. The load balancer notices and stops routing new connections before the process exits. - **Expect sticky connections.** Clients hold long-lived HTTP/2 connections for watches. After a failed server comes back, load stays skewed toward the survivors. The `--goaway-chance` flag makes kube-apiserver randomly close a small fraction of HTTP/2 connections so clients reconnect through the load balancer and spread out again. ## In-cluster clients take a different path Pods do not use the external endpoint by default. They use the `kubernetes` Service in the `default` namespace. Each API server records itself through the **lease endpoint reconciler**, which is the default for `--endpoint-reconciler-type`. The older `master-count` mode and its `--apiserver-count` flag are deprecated. The Service's endpoints therefore list exactly the live API servers, and kube-proxy spreads in-cluster traffic across them. | Client | Path to an API server | |---|---| | kubelet, kube-proxy, kubectl, CI | kubeconfig URL -> VIP or load balancer -> any healthy replica | | Pods using in-cluster config | `kubernetes` Service -> kube-proxy -> any live replica | ## Common failure pattern A load balancer that checks only TCP port 6443 keeps sending traffic to an API server whose local etcd member is broken. Every request that lands there fails, and the errors look random. Switching the check to `/readyz` fixes that class of problem.
- Why not point each kubelet at the API server on its nearest control-plane node instead of a load balancer?Then the kubelet loses the control plane whenever that one node fails, even though other replicas are healthy. Its Lease renewals stop, and the node lifecycle controller eventually marks a healthy worker unreachable. One stable, health-checked endpoint lets any node survive the loss of any single API server.
- After a failed API server rejoins, it gets almost no traffic for hours. Why, and what helps?Clients hold long-lived HTTP/2 connections, mostly watches, and the load balancer only balances new connections. The survivors keep the existing ones. Setting `--goaway-chance` makes kube-apiserver randomly send GOAWAY on a small fraction of requests, so clients reconnect through the load balancer and spread out again. It should not be used with a single API server or without a load balancer.
saying these in an interview costs you the question
- Terminate TLS at the load balancer to inspect API traffic
- A TCP connect check on 6443 proves the API server is healthy
- Only the leader kube-apiserver can accept writes
- Set --apiserver-count so the kubernetes Service lists all servers
- DNS round-robin is as good as a health-checked virtual IP