skip to content

The control-plane node running the active kube-controller-manager loses power, and a shipment-tracking rollout stalls for 47 seconds. What governs Kubernetes leader-election failover, and would you tune it?

level: seniorimportance: nice to knowfreq 33%

answer

  1. a Lease in kube-system
  2. 15, 10, 2
  3. stop before anyone can take over
  4. acquireTime versus power-off time
  5. endpoint failover and cache sync

basics

~20 s

A standby takes over only after the leader's Lease looks expired: 15 seconds by default, plus a jittered 2-second retry. The rest of a 47-second stall usually comes from endpoint failover, stuck connections and cache sync.

solid answer

~50 s

kube-controller-manager and kube-scheduler each hold a `Lease` in `kube-system`, named after the component. The leader renews it. Standbys poll it and may take it over only after they have seen no change for `--leader-elect-lease-duration` (default 15s). They retry every `--leader-elect-retry-period` (default 2s, with jitter). A leader that cannot renew within `--leader-elect-renew-deadline` (default 10s) stops. kube-controller-manager exits so it quits before anyone else can take over. A power loss therefore costs roughly 15-20 seconds before the new leader starts. The new leader then syncs its informer caches, and if the scheduler leader was on the same node, it fails over as well. Standbys and kubelets that used the dead node's endpoint also wait on failover and connection timeouts. To diagnose, compare the Lease's `acquireTime` with the power-off time. Shortening the lease rarely helps and raises the risk of losing leadership during a brief API server hiccup.

code

bash · 2 lines
bash
kubectl -n kube-system get lease kube-controller-manager kube-scheduler \
  -o custom-columns=NAME:.metadata.name,HOLDER:.spec.holderIdentity,ACQUIRED:.spec.acquireTime,RENEWED:.spec.renewTime,TRANSITIONS:.spec.leaseTransitions

go deeper

for a junior

Remember that only one controller-manager and one scheduler act at a time, and a standby takes over when the leader's Lease expires.

for a middle

Explain the three timers and their defaults, and why the renew deadline must be shorter than the lease duration.

for a senior

Break a real stall into lease expiry, endpoint failover, stuck connections and cache sync, using the Lease timestamps and logs, and fix the largest part first.

for a principal

Weigh faster failover against API write load and self-inflicted leadership loss, and decide whether a stall of this size even violates an SLO.

## How the Lease lock works In an HA control plane, three copies of `kube-controller-manager` and three copies of `kube-scheduler` run, but only one of each acts. They coordinate through a **Lease** object from `coordination.k8s.io/v1`: - The lock objects are the Leases `kube-controller-manager` and `kube-scheduler` in the `kube-system` namespace. - `spec.holderIdentity` names the current leader, `spec.leaseDurationSeconds` states the lease length, and `spec.acquireTime` and `spec.renewTime` record when leadership was taken and last renewed. `spec.leaseTransitions` counts how many times leadership has changed hands. - The lock's safety comes from **etcd**, not from the replica count. Updates go through kube-apiserver with optimistic concurrency, so two candidates can never both succeed in taking the same Lease. Two replicas are as safe as three. The third only adds spare capacity. ## The three timers | Flag | Default | Meaning | |---|---|---| | `--leader-elect-lease-duration` | 15s | How long a standby must see no change before it may take the lease | | `--leader-elect-renew-deadline` | 10s | How long the leader keeps retrying renewal before it stops leading | | `--leader-elect-retry-period` | 2s | Interval between attempts, with jitter (up to 2.2 times as long) | client-go enforces the ordering: the lease duration must be greater than the renew deadline, and the renew deadline must be greater than 1.2 times the retry period. The gap between 10s and 15s matters. A leader that cannot renew stops acting **before** a standby is allowed to take over, so two leaders never act at the same time. In kube-controller-manager, losing leadership logs `leaderelection lost/stopped` and the process exits. The kubelet then restarts it as a standby. ## Timeline of the 47-second stall The shipment-tracking team's 9-node bare-metal cluster loses the control-plane node that holds both Leases. One plausible breakdown: 1. **0-15 s:** standbys see an unchanged Lease. Expiry is measured from when *they* last saw the record change, not from `renewTime`. 2. **15-~20 s:** the next jittered retry succeeds, and the log shows `Successfully acquired lease`. `leaseTransitions` goes up. 3. **Additional delay: the endpoint.** If the virtual IP lived on the dead node, or the load balancer's health check is slow, standbys and kubelets wait for failover. Requests already sent to the dead server hang until the client times out. 4. **Additional delay: startup.** The new kube-controller-manager starts its controllers and waits for informer caches to sync before the Deployment controller acts. If the kube-scheduler leader was also on that node, new Pods also wait for its separate failover. On a small cluster, steps 1-2 explain under 20 seconds. The rest of the 47 seconds is usually steps 3-4. ## Diagnosing it - Read the Lease objects and compare `acquireTime` with the time the node lost power. - Search the new leader's log for `Attempting to acquire leader lease...` and `Successfully acquired lease`, and measure the gap between them. - Compare the VIP or load balancer's failover time with the gap between the lease takeover and the rollout resuming. - Check whether kube-scheduler moved at the same time. Both leaders on one node double the stall. ## Should you tune it? Usually not. - **Shorter leases mean more writes and more fragility.** Every renewal is an API write. A lease that is too short turns a brief API or etcd slowdown into a lost leadership, a process exit and a restart, which causes a bigger stall than the one you tried to fix. - **Fix the endpoint first.** Faster VIP failover and `/readyz` checks remove more of the stall than shaving seconds off the lease. - **Planned restarts behave differently.** kube-scheduler releases its lease on graceful shutdown, so a standby can take over at once. kube-controller-manager does not by default. Releasing on exit is behind the `ControllerManagerReleaseLeaderElectionLockOnExit` feature gate, alpha and off by default since 1.36. A rolling restart of the kube-controller-manager leader therefore still waits for the lease to expire. - If you do tune, change all three flags together, keep the required ordering, and test under API server load.

  • Why do controller-manager replicas not need an odd count the way etcd members do?
    They do not vote. The Lease is a single object stored in etcd and updated through the API server with optimistic concurrency, so etcd's quorum decides who wins. Two replicas give one spare, and three give two. Split-brain protection comes from the renew deadline being shorter than the lease duration, not from replica count.
  • The leader is healthy, but its path to the API servers breaks for 12 seconds. What happens?
    It cannot renew within the 10-second renew deadline, so kube-controller-manager stops leading and exits. At about 15 seconds without a visible change, a standby that can still reach the API takes the Lease. The kubelet restarts the old process, which comes back as a standby. The price is a failover and cache resync caused by a transient network fault.

saying these in an interview costs you the question

  • Standbys take over the instant the leader's node loses power
  • A deposed leader keeps running controllers until someone kills it
  • Controller-manager replicas must be odd so they can vote
  • Setting the lease duration to 2 seconds is a safe fix for stalls
  • kube-controller-manager always releases its lease on graceful shutdown