What is leader imbalance, and how do preferred-leader election and auto.leader.rebalance.enable address it?
answer
- preferred leader = first in replica list
- leadership doesn't return after broker recovers
- auto.leader.rebalance.enable default true
- leader.imbalance.per.broker.percentage = 10%
- election needs preferred replica in ISR
basics
~20 sThe first replica in a partition's replica list is its 'preferred leader'. After failures, leadership drifts off preferred replicas, concentrating load on a few brokers. auto.leader.rebalance.enable=true periodically moves leadership back to preferred replicas when imbalance exceeds a threshold.
solid answer
~40 sEach partition has an ordered replica list; the first entry is the 'preferred leader', chosen so leadership is evenly spread when the cluster is healthy. When a broker restarts or fails, its partitions' leadership fails over to other replicas — and stays there even after the broker recovers, so leadership becomes concentrated on a subset of brokers (leader imbalance). That makes some brokers hot. The controller can run a preferred-leader election to move leadership back to the preferred replicas (provided they're in the ISR). With auto.leader.rebalance.enable=true (default), the controller checks every leader.imbalance.check.interval.seconds and triggers a rebalance when any broker's imbalance ratio exceeds leader.imbalance.per.broker.percentage (default 10%). You can also trigger it manually with kafka-leader-election.sh --election-type preferred (or the older kafka-preferred-replica-election).
go deeper
Know the preferred leader is the first replica and imbalance means leadership is uneven.
Explain why imbalance arises after restarts and name auto.leader.rebalance.enable + the percentage threshold.
Discuss ISR requirement, the churn cost of auto-rebalance, and when to run preferred election manually.
Weigh auto vs manual rebalancing policy at scale and integrate it into rolling-restart automation.
## Preferred leader: the concept A partition's replicas are stored as an **ordered list**, e.g. `[3, 1, 2]`. The **preferred leader** is the *first* broker in that list (broker 3 here). When Kafka first assigns partitions, it arranges these lists so that, if every preferred leader is the actual leader, leadership is spread evenly across brokers — balancing the read/write load (leaders do all client I/O). ## How imbalance arises Leadership only sits on the preferred replica while the cluster is healthy. When a broker goes down or restarts: 1. Its partitions lose their leaders, and leadership **fails over** to another in-sync replica. 2. When the broker comes back, leadership does **not** automatically return — the new leaders keep serving. After a few rolling restarts, leadership piles up on a handful of brokers. Those brokers handle disproportionate client traffic → hotspots, higher latency, uneven disk/network use. This is **leader imbalance**. ## Preferred-leader election The fix is a **preferred-leader election**: the controller moves each partition's leadership back to its preferred (first-listed) replica — but only if that replica is currently in the **ISR** (you can't make a stale replica the leader without risking data loss). Two ways to trigger it: - **Manual:** `kafka-leader-election.sh --bootstrap-server ... --election-type PREFERRED --all-topic-partitions` (older name: `kafka-preferred-replica-election.sh`). - **Automatic:** controlled by three configs: - `auto.leader.rebalance.enable` (default **true**) — turns on the periodic check. - `leader.imbalance.check.interval.seconds` (default 300) — how often the controller checks. - `leader.imbalance.per.broker.percentage` (default 10) — if the fraction of a broker's partitions that are NOT led by their preferred leader exceeds this percent, a rebalance fires. ## Edge cases - Auto-rebalance triggers a brief leadership change for affected partitions, causing momentary unavailability and metadata churn — some large-cluster operators **disable** it and run preferred election deliberately (e.g. at the end of a rolling restart) to control timing. - Preferred election does NOT move data; it only changes which existing replica is leader. To change *where replicas live*, you need a reassignment. - A preferred leader that is out of the ISR is skipped until it catches up.
- Does preferred-leader election move partition data between brokers?No. It only changes which existing replica is the leader. Moving replicas to different brokers requires a partition reassignment (kafka-reassign-partitions).
- Why might a large cluster disable auto.leader.rebalance.enable?Auto-rebalance can fire at unpredictable times, causing leadership churn and brief unavailability across many partitions at once. Operators disable it and run preferred election deliberately (e.g. after a rolling restart) to control timing and blast radius.
saying these in an interview costs you the question
- Saying leadership automatically returns to the preferred broker after it recovers — it does not, hence imbalance.
- Confusing preferred-leader election (changes leader among existing replicas) with reassignment (moves replicas).
- Thinking election can promote an out-of-sync replica — it requires the preferred replica be in the ISR.