When is kafka-reassign-partitions.sh needed instead of preferred leader election, and how does reassignment differ from a leader election?
answer
- Election = leader only, no data move
- Reassign = move replicas/data across brokers
- Needed for add/remove broker, RF change, skew
- --generate / --execute / --verify
- Throttle replication; AR order sets preferred leader
basics
~20 sPreferred leader election only changes which existing replica leads. kafka-reassign-partitions.sh actually moves or changes replicas across brokers — needed for adding/removing brokers, fixing skewed data placement, or changing replication factor. It copies data, so it is heavyweight.
solid answer
~50 sLeader election (auto or kafka-leader-election.sh) only re-elects leaders among the replicas a partition already has — it moves no data. kafka-reassign-partitions.sh changes the replica assignment itself: it can relocate replicas onto different brokers, change replication factor, and as a side effect set new preferred leaders by reordering AR. You need it when leadership balancing isn't enough: adding new brokers (which start empty), decommissioning a broker, or correcting data/leader skew that elections can't fix because the preferred replicas live on overloaded brokers. The workflow is --generate a candidate plan, --execute it (data is replicated to the new brokers, which can saturate network — throttle with --throttle / replication-quota), then --verify until complete. Because it physically copies log data, reassignment is far heavier and slower than an election, and you order the new AR list to also fix leader balance.
go deeper
Know that reassignment moves replicas/data while election only changes the leader.
List when reassignment is required (add/remove broker, RF change, skew) and the generate/execute/verify flow.
Operate it safely: throttling, batching, editing AR order to set preferred leaders, verifying and removing throttles.
Design cluster scale-out/decommission plans balancing data movement cost, durability, and leadership distribution.
## Two different operations **Leader election** (preferred): chooses a new *leader* from a partition's *existing* replicas. No data moves. Cheap and fast. Fixes *leadership* imbalance. **Partition reassignment**: changes the *set and order of replicas* (the AR list) for partitions, potentially placing them on *different brokers*. Data must be copied to the new brokers. Expensive and slow. Fixes *data placement* and *replica* imbalance. ## When election is NOT enough Preferred leader election can only spread leadership across the brokers that *already hold replicas*. It cannot help when: - **You add brokers** — new brokers hold no replicas, so no partition's preferred leader is on them. You must reassign replicas onto them. - **You remove/decommission a broker** — you must move its replicas elsewhere first. - **Data/storage skew** — some brokers hold far more partition data; rebalancing leadership doesn't move bytes. - **Change replication factor** — add or drop replicas per partition. - **Preferred replicas all sit on hot brokers** — you must relocate replicas, then the new AR ordering also sets better preferred leaders. ## kafka-reassign-partitions.sh workflow 1. `--generate`: given `--topics-to-move-json-file` and `--broker-list`, it proposes a `reassignment-json-file` (and prints the current assignment so you can roll back). The proposal is a starting point, not optimal — you typically hand-edit AR ordering to control preferred leaders. 2. `--execute --reassignment-json-file plan.json`: the controller starts moving replicas. New replicas catch up from the leader; once in ISR, the old replicas are removed. 3. `--verify`: reports per-partition status (in progress / completed) and removes throttles when done. ## Throttling — critical at scale Replica movement saturates inter-broker network and can starve normal replication and client traffic. Use `--throttle <bytes/sec>` (sets `leader.replication.throttled.rate` / `follower.replication.throttled.rate`) to cap it. Forgetting to throttle (or to remove the throttle after --verify) is a classic incident cause. ## Interaction with preferred leaders The **order** of brokers in each partition's new AR list sets the new preferred leaders. So a well-designed reassignment both relocates replicas *and* rebalances leadership in one step — then preferred election (auto or manual) settles leadership onto the new first replicas. ## Edge cases - Reassignment is durable/idempotent through the controller; it survives controller failover. - It is heavy: schedule during low traffic, throttle, and reassign in batches rather than the whole cluster at once. - It does not by itself trigger leadership change to the new preferred replica immediately for already-led partitions; a preferred election finalizes leader placement.
- Why must you throttle a large partition reassignment?Replica movement copies log data over the inter-broker network and can saturate it, starving normal replication and client traffic. --throttle caps the replication rate; remove it after --verify completes.
- After a reassignment, how do you ensure leadership lands on the intended brokers?Order the new AR list so the intended broker is first (the preferred leader), then run or wait for a preferred leader election to move leadership onto it.
saying these in an interview costs you the question
- Claiming preferred leader election can rebalance a freshly added empty broker — it can't; reassignment is required.
- Saying reassignment is lightweight like an election — it copies data and can saturate the network.
- Forgetting that AR ordering determines the new preferred leader.