You start a brand-new Kafka broker with a fresh broker.id and join it to the cluster. Why does it not start serving traffic, and what must you do to put data on it?
answer
- new broker.id = idle, no auto-rebalance
- kafka-reassign-partitions --generate/--execute/--verify
- follower catches up -> joins ISR
- --throttle replication rate
- Cruise Control / Self-Balancing automates it
basics
~10 sA new broker joins the cluster but Kafka never auto-moves existing partitions onto it. You must explicitly reassign replicas to the new broker using kafka-reassign-partitions to scale out.
solid answer
~40 sAdding a broker is just registering a new broker.id with the cluster (KRaft controller or, historically, ZooKeeper). Kafka does NOT automatically rebalance existing partitions onto it, so the broker sits idle except for partitions of newly created topics. To actually use the new capacity you generate a reassignment plan with kafka-reassign-partitions.sh (--generate to propose, --execute to run, --verify to track progress) that lists the new broker.id in the replica set of chosen partitions. Replicas then bootstrap as followers, fetch from the leader until in-sync (ISR), and only after catching up can leadership be moved to them. Throttle the move with --throttle to avoid saturating the network. Confluent/commercial offerings add auto-balancers (Self-Balancing Clusters, Cruise Control) that automate this.
go deeper
Know that adding a broker requires explicit reassignment; Kafka does not auto-balance.
Drive kafka-reassign-partitions end to end (--generate/--execute/--verify) and understand ISR catch-up.
Add throttling, preferred-leader election, and explain why open-source needs Cruise Control for automation.
Reason about capacity planning, throttle budgets vs client SLA, and when to adopt an auto-balancer fleet-wide.
**What a Kafka broker is:** A broker is one server process in a Kafka cluster, identified by a unique integer `broker.id` (or `node.id` in KRaft mode). Topics are split into **partitions**; each partition has a **leader** replica (handles reads/writes) and zero or more **follower** replicas on other brokers for redundancy. The set of replicas that are fully caught up with the leader is the **ISR (in-sync replicas)**. **Why a new broker is idle:** When you boot a process with a new `broker.id`, it registers itself with the cluster metadata (the **KRaft controller quorum** in modern Kafka, or **ZooKeeper** in pre-3.x/legacy deployments). Registration only makes the broker *available*; Kafka has **no automatic data rebalancing** built into open-source Apache Kafka. Existing partitions keep their current replica assignments, so the new broker holds no data and serves no leaders. It will only naturally receive partitions for *newly created topics* (the partition assigner spreads new topics across all live brokers). **How to actually scale out:** Use the CLI tool **`kafka-reassign-partitions.sh`**: 1. Write a `topics-to-move.json` listing topics, then run with `--generate --broker-list "1,2,3,4"` (including the new id) to get a *proposed* reassignment JSON. 2. Review/edit it, then run `--execute --reassignment-json-file plan.json`. This rewrites each partition's replica set to include the new broker. 3. The new replica starts as a **follower**, fetches records from the current leader, and grows its log until it joins the ISR. Only an in-sync replica is eligible to become leader. 4. Run `--verify` to confirm completion and **remove throttles**. **Throttling:** Moving terabytes can saturate NICs and starve client traffic, so pass `--throttle <bytes/sec>` (this sets `leader.replication.throttled.rate` / `follower.replication.throttled.rate` dynamically). **Edge cases & notes:** New brokers do not auto-acquire leadership even after catch-up unless you also run **preferred leader election** (or have `auto.leader.rebalance.enable=true`, which only rebalances toward the *preferred* (first) replica). Tools like **Cruise Control** (open source) and **Confluent Self-Balancing Clusters** automate generation + throttling + goal-based balancing so you don't hand-craft JSON.
- After reassignment finishes, the new broker holds replicas but still serves no leaders. Why?Reassignment only places replicas; leadership stays where it was. You need preferred leader election (kafka-leader-election.sh --election-type PREFERRED) or auto.leader.rebalance.enable to move leaders onto the new broker's replicas.
- What does the --verify step do besides reporting status?Besides reporting per-partition completion, --verify removes the replication throttles that --execute set, so you don't leave the cluster permanently rate-limited.
saying these in an interview costs you the question
- Claiming Kafka auto-rebalances partitions onto a new broker (open-source Apache Kafka does not).
- Thinking a new broker immediately starts taking write/read traffic for existing topics.
- Forgetting that leadership must be moved separately via preferred leader election.