skip to content

Preferred Leader Election and Partition Balancing

Keeping leadership evenly spread with preferred-leader election and partition reassignment. Interviewers ask because leader skew quietly overloads one broker while the rest idle.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is the 'preferred leader' for a Kafka partition, and why does it matter?

level: juniorimportance: must knowfreq 70%

answer

  1. First broker in AR list
  2. Only the leader serves traffic
  3. Restarts drift leadership away
  4. Re-spreads load, doesn't move data
  5. Must be in ISR to be elected

basics

~20 s

The preferred leader is the first broker listed in a partition's replica assignment (the assigned-replica list). Kafka tries to make it the leader so leadership is spread evenly across brokers and no single broker is overloaded.

solid answer

~40 s

Every partition has an ordered list of replicas called the assigned replicas (AR). The preferred leader is the first broker in that list. When a topic is created, Kafka assigns these lists round-robin so leadership is balanced across the cluster. Over time, broker restarts and failures cause leadership to drift onto whichever replicas happened to be in-sync, concentrating leaders on fewer brokers. Because only the leader handles produce and consume traffic for a partition, an imbalanced set of leaders means uneven CPU, network, and disk load. Preferred leader election restores leadership to the first replica in each AR list (when it is in-sync), re-spreading the load. This is why operators monitor leader imbalance and either rely on auto.leader.rebalance.enable or run kafka-leader-election.sh PREFERRED.

go deeper

for a junior

Know it's the first replica in the list and that Kafka prefers it to balance load.

for a middle

Explain why leadership drifts after restarts and that only ISR members are eligible.

for a senior

Connect leader placement to broker resource utilization and the auto/manual rebalance mechanisms.

for a principal

Reason about cluster-wide load modeling and when leader balancing alone is insufficient (needs replica reassignment).

## First principles A Kafka **topic** is split into **partitions**. Each partition is replicated onto several brokers for fault tolerance. The list of brokers holding a partition's replicas is called the **assigned replicas (AR)** — and crucially it is *ordered*. For example partition 0 might have AR = [3, 1, 2], meaning broker 3 holds the first replica, broker 1 the second, broker 2 the third. At any moment exactly one replica is the **leader**: it is the only replica that serves producer writes and consumer reads for that partition. The others are **followers** that just copy the leader's log. The subset of replicas that are fully caught up is the **in-sync replica set (ISR)**. ## What 'preferred leader' means The **preferred leader** is simply **the first broker in the AR list** (broker 3 in the example). It is 'preferred' because Kafka assigns AR lists round-robin at topic-creation time, so if every partition's leader were its first replica, leadership would be spread evenly across all brokers. ## Why it matters Because only the leader does I/O for a partition, where the leaders sit determines load distribution. Two things push leadership away from the preferred replica: 1. **Broker restart / failure**: when the preferred leader broker goes down, leadership moves to another in-sync replica. When the broker comes back it rejoins as a *follower*, not automatically as leader. 2. **Repeated churn**: over many restarts, leaders pile up on the brokers that stayed alive. The result is **leader imbalance**: some brokers host far more leaders than others, so they saturate on CPU/network while others idle. Restoring leadership to the preferred replicas re-balances the load. A replica is only eligible to become leader if it is currently in the ISR — an out-of-sync preferred replica is skipped until it catches up. ## How it's triggered - Automatically, if `auto.leader.rebalance.enable=true` (the default), the controller periodically checks imbalance and runs preferred election. - Manually, via `kafka-leader-election.sh --election-type PREFERRED`. ## Edge cases - If the preferred replica is not in the ISR, the election does nothing for that partition (it won't elect an out-of-sync replica unless unclean election is allowed, which is a separate, dangerous setting). - Preferred election only *moves leadership*; it never moves data or changes the replica assignment. Moving replicas requires reassignment.

  • Why does leadership drift away from the preferred replica over time?
    When a broker restarts or fails, its leaderships move to surviving in-sync replicas. On rejoin the broker comes back as a follower, so leaders accumulate on the brokers that stayed up.
  • Does preferred leader election move partition data between brokers?
    No. It only changes which existing replica is the leader. Moving the actual data requires a replica reassignment with kafka-reassign-partitions.sh.

saying these in an interview costs you the question

  • Saying the preferred leader is whichever broker has the most free capacity — it is statically the first replica in the AR list.
  • Claiming preferred election relocates partition data or changes the replica set.
  • Thinking an out-of-sync preferred replica gets elected anyway.

context

open as a page

How does auto.leader.rebalance.enable work, and what do leader.imbalance.check.interval.seconds and leader.imbalance.per.broker.percentage control?

level: middleimportance: must knowfreq 60%

basics

~20 s

auto.leader.rebalance.enable (default true) lets the controller automatically run preferred leader election. The check interval sets how often it inspects imbalance, and the per-broker percentage is the imbalance threshold that triggers a rebalance for a broker.

open as a page

How do you manually trigger preferred leader election with kafka-leader-election.sh, and how does PREFERRED differ from UNCLEAN election?

level: middleimportance: should knowfreq 45%

basics

~10 s

Run kafka-leader-election.sh with --election-type PREFERRED to move leadership back to preferred replicas. PREFERRED only elects in-sync replicas (safe). UNCLEAN can elect an out-of-sync replica to restore availability, risking data loss.

open as a page

When is kafka-reassign-partitions.sh needed instead of preferred leader election, and how does reassignment differ from a leader election?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Preferred leader election only changes which existing replica leads. kafka-reassign-partitions.sh actually moves or changes replicas across brokers — needed for adding/removing brokers, fixing skewed data placement, or changing replication factor. It copies data, so it is heavyweight.

open as a page

Design a leadership-balancing strategy for a large, latency-sensitive cluster during rolling restarts. What do you tune, and what are the failure modes?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Often disable auto.leader.rebalance.enable and trigger preferred election deliberately after each restart batch, scoped via JSON and staggered, so leadership churn is controlled. Watch for client metadata storms, imbalance from out-of-sync preferred replicas, and unclean-election durability risks.

open as a page