skip to content

kube-controller-manager runs on all three control-plane nodes of a highly available cluster. Why don't three copies of the ReplicaSet controller each create their own set of Pods?

level: seniorimportance: should knowfreq 50%

answer

  1. active/passive, not active/active
  2. Lease in kube-system, holderIdentity + renewTime
  3. renewDeadline 10s < leaseDuration 15s
  4. failover = lease expiry + rebuild informers
  5. apiserver resourceVersion conflicts fence stale writers

basics

~20 s

Because of leader election: all three instances start, but each competes to hold a Lease object in the API server. Only the lease holder runs its control loops; the others idle as hot standbys and take over if the lease expires.

solid answer

~50 s

kube-controller-manager (and kube-scheduler) are **active/passive**, unlike kube-apiserver which is active/active. Every replica starts and immediately contends for a **Lease** object (`kube-system/kube-controller-manager`) using optimistic concurrency: acquire or renew by updating the Lease with a resourceVersion precondition, so exactly one writer wins. Only the winner starts its controllers; the losers loop, watching the lease and doing no reconciliation. The leader renews continuously (`--leader-elect-renew-deadline`, default 10s, against a `--leader-elect-lease-duration` of 15s and a 2s retry period). If it crashes or is partitioned, the lease goes unrenewed, expires, and another replica acquires it — a failover of roughly the lease duration during which no reconciliation happens. A well-behaved leader that loses the lease stops its loops (typically by exiting) rather than continuing to write. The API server is still the safety net: writes are versioned, so even a brief two-leader overlap cannot silently corrupt state.

go deeper

for a junior

Know that only one instance is active at a time and the rest are standbys that take over automatically.

for a middle

Name the Lease object in kube-system, the renew/expire timing, and that the standby starts by listing and watching to rebuild its cache.

for a senior

Explain the resourceVersion compare-and-swap as the exclusion mechanism, the failover window and its operational impact, and how to spot lease flapping in metrics.

for a principal

Reason about the deliberate choice of time-based leases over fencing tokens, the availability-versus-failover-latency tradeoff in tuning lease duration, and the requirement that operators with external side effects elect a leader.

## Why singleton control loops Controllers are convergent but not commutative in the short term. Two ReplicaSet controllers observing `replicas: 3` with zero Pods would each create three Pods; the next sync would delete the surplus, but you would have churned six Pods and disrupted workloads. Rather than make every loop safe for concurrent actors, Kubernetes makes the loops **singleton**: one active instance cluster-wide per controller-manager (and per scheduler). ## The lease mechanism Leader election is built on ordinary API objects, not a separate consensus service — etcd's Raft already provides the linearizable store, so election reduces to a compare-and-swap on one object. - The lock is a **coordination.k8s.io/v1 Lease** in `kube-system`, named after the component. Historically Endpoints or ConfigMap annotations were used; Lease is the modern default because it is cheap to write and does not pollute Endpoints. - Its fields: `holderIdentity` (usually hostname plus a random suffix), `leaseDurationSeconds`, `acquireTime`, `renewTime`, `leaseTransitions`. - **Acquire**: if the lease is absent or `renewTime` is older than the lease duration, a candidate writes itself as holder. The write carries the observed `resourceVersion`, so the API server rejects all but one concurrent attempt with a conflict — that is the mutual exclusion. - **Renew**: the holder rewrites `renewTime` every retry period. If it cannot renew before the renew deadline (API server unreachable, GC pause, node freeze), it must assume it lost leadership and stop. Relevant flags: `--leader-elect` (on by default), `--leader-elect-lease-duration` (15s), `--leader-elect-renew-deadline` (10s), `--leader-elect-retry-period` (2s). The invariant is renewDeadline < leaseDuration, with enough margin that a slow but healthy leader is not displaced. ## What failover looks like Kill the leader and nothing visible happens for a few seconds: no reconciliation, so a Pod that dies is not replaced, a rollout pauses, a NotReady node is not processed. Once the lease expires a standby acquires it, builds its informer caches from a **list-then-watch** against the API server, and resumes. Because reconciliation is level-triggered, the new leader needs no handoff state — it reads the world and computes the delta. Nothing is lost; work is only delayed. `leaseTransitions` increments, and you can watch the transition with `kubectl get lease -n kube-system kube-controller-manager -o yaml` or the component's `leader_election_master_status` metric. ## Split brain and why it is survivable Leases are time-based, and clocks and pauses are imperfect: a leader stalled by a long GC pause may believe it still holds a lease that has already been taken. Brief overlap is therefore possible. The system tolerates it because: - Every write goes through the API server with **optimistic concurrency** — a stale actor's update fails on resourceVersion conflict. - Controllers are **idempotent and level-triggered** — the worst outcome is extra churn (a surplus Pod created then deleted), not corrupted state. - The stale leader's renewal is failing, so it exits within the renew deadline. This is deliberately weaker than a fencing-token design; Kubernetes trades strict exclusion for simplicity, relying on convergence to clean up. ## Contrast with other components - **kube-apiserver**: stateless and horizontally scalable — all replicas serve traffic behind a load balancer, no election. - **etcd**: has its own Raft leader, unrelated to Kubernetes leases; do not conflate them. - **kube-scheduler**: same lease pattern as controller-manager, its own Lease object. - **kubelet**: no election — one per node, each owning only its own Pods. - **Custom controllers/operators**: use the same client-go leader-election library, and should, for exactly the same reason. Running two operator replicas without it is a real production bug. ## Operating notes Symptoms of election trouble: rapidly incrementing `leaseTransitions` (flapping, usually from an overloaded or high-latency API server) — reconciliation then stalls repeatedly. Extending lease duration reduces flapping but lengthens failover. On single-node control planes leader election still runs; it is a no-op contest with one candidate, and disabling it saves a trivial amount of write traffic but removes the safety net if a second instance ever appears.

  • During a controller-manager failover, what is unavailable, and for how long?
    Reconciliation stops for roughly the unrenewed portion of the lease plus cache warm-up — commonly on the order of ten to twenty seconds. In that window failed Pods are not replaced, rollouts do not progress and unreachable nodes are not tainted. Nothing running is affected, and no state is lost, because the new leader recomputes desired versus actual from the API server rather than resuming a queue.
  • Is a brief period with two leaders dangerous?
    It is possible under clock skew or long pauses and is tolerated by design. Every controller write is an optimistic-concurrency update, so the stale leader's writes lose on resourceVersion conflicts, and level-triggered reconciliation converges away any duplicate objects that slipped through. The stale instance also fails its renewal and exits within the renew deadline, so the overlap is bounded.
  • Should a custom operator you write use leader election?
    Yes, if you run more than one replica for availability. Use client-go's leader election with a Lease in your operator's namespace, run the reconcilers only while leading, and terminate on lost leadership. Without it, two replicas will duplicate side effects such as creating external cloud resources, which reconciliation cannot always undo.

A talking stick: only whoever holds it acts, and they must keep announcing they still hold it or someone else picks it up.

saying these in an interview costs you the question

  • Saying all controller-manager replicas run loops concurrently and 'it's fine because reconciliation is idempotent'.
  • Confusing the Kubernetes leader-election Lease with the etcd Raft leader.
  • Claiming leader election guarantees strictly one leader at all times; overlap is possible and merely tolerated.
  • Thinking a failover loses queued work — the new leader recomputes from current state.
  • Assuming kube-apiserver also elects a leader; it is active/active behind a load balancer.

context