How does auto.leader.rebalance.enable work, and what do leader.imbalance.check.interval.seconds and leader.imbalance.per.broker.percentage control?
answer
- enable=true default, controller-driven
- check.interval.seconds default 300
- per.broker.percentage default 10
- ratio of non-preferred leaders per broker
- election = brief unavailability + metadata refresh
basics
~20 sauto.leader.rebalance.enable (default true) lets the controller automatically run preferred leader election. The check interval sets how often it inspects imbalance, and the per-broker percentage is the imbalance threshold that triggers a rebalance for a broker.
solid answer
~40 sThese are broker-level configs read by the controller. auto.leader.rebalance.enable=true makes the controller periodically scan for leader imbalance and trigger preferred leader election automatically. leader.imbalance.check.interval.seconds (default 300) is how often that scan runs. leader.imbalance.per.broker.percentage (default 10) is the threshold: for each broker the controller computes the ratio of partitions whose current leader is NOT the preferred leader, relative to the partitions for which that broker is the preferred leader; if that ratio exceeds the percentage for any broker, a preferred election is triggered. So with the defaults, every 5 minutes the controller rebalances if any broker is more than 10% imbalanced. Operators sometimes disable this to control exactly when leadership moves, since elections cause brief per-partition unavailability and a client metadata refresh.
go deeper
Know the three config names and that they auto-trigger preferred election with a default 5-minute check and 10% threshold.
Explain how the imbalance ratio is computed and why a team might disable auto-rebalance.
Weigh election churn vs imbalance tolerance and tune the percentage/interval for the workload.
Define an org policy: auto vs scheduled elections, interaction with throughput SLOs and controller load.
## The three configs All three are **broker configs** evaluated by the active **controller** (the broker that manages cluster metadata). ### `auto.leader.rebalance.enable` - Default **true**. - When true, the controller runs a background task that periodically looks for leader imbalance and, if found, triggers a **preferred leader election** automatically — no operator action needed. - When false, leadership only returns to preferred replicas when you run it manually (e.g. `kafka-leader-election.sh`). This is common in large or latency-sensitive clusters where operators want to schedule elections during low-traffic windows. ### `leader.imbalance.check.interval.seconds` - Default **300** (5 minutes). - How often the controller's background task wakes up to measure imbalance. Only meaningful when auto-rebalance is enabled. ### `leader.imbalance.per.broker.percentage` - Default **10**. - The trigger threshold, expressed as a percentage. For each broker the controller computes an **imbalance ratio**: roughly imbalance% = (partitions led by this broker where it is NOT the preferred leader) / (partitions for which this broker is the preferred leader) * 100 Intuitively: of all the partitions this broker *should* be leading, how many are currently being led elsewhere (or it's leading partitions it shouldn't). If **any** broker's ratio exceeds the configured percentage, the controller triggers a preferred election across the cluster. ## Putting it together (defaults) Every 300 seconds the controller checks all brokers; if even one is >10% imbalanced, it runs preferred leader election to restore leadership to the preferred replicas (those currently in ISR). ## Cost and edge cases - A leader election makes each affected partition **briefly unavailable** (a few hundred ms) while leadership transfers, and clients must refresh metadata to find the new leader. That's why high-throughput shops often set `auto.leader.rebalance.enable=false` and trigger elections manually on a schedule. - Auto-rebalance only fixes imbalance that preferred election can fix — it cannot move replicas or fix imbalance caused by a permanently down preferred broker. - Lowering the percentage makes rebalancing more aggressive (more frequent elections, more churn); raising it tolerates more imbalance before acting. - The check interval should not be set extremely low; frequent scans plus frequent elections add controller load and client metadata churn.
- Why might a team set auto.leader.rebalance.enable=false in production?To control exactly when leadership moves. Each election briefly interrupts affected partitions and forces client metadata refreshes, so they prefer to trigger preferred election manually during low-traffic windows.
- What happens if you set leader.imbalance.per.broker.percentage very low, like 1?Almost any drift triggers a rebalance, causing frequent elections and continuous metadata churn and partition blips — usually undesirable.
saying these in an interview costs you the question
- Saying the percentage is an absolute partition count rather than a ratio.
- Claiming auto-rebalance can move replica data between brokers.
- Asserting leader elections are free / zero-impact on clients.