skip to content

Walk me through safely decommissioning a broker so you can take it out of the cluster permanently without data loss.

level: middleimportance: must knowfreq 65%

answer

  1. drain replicas off broker.id first
  2. reassign to exclude target, --throttle
  3. --verify until zero partitions + zero URP
  4. controlled.shutdown migrates leaders gracefully
  5. never drop below min.insync.replicas

basics

~10 s

Move every replica off that broker.id to other brokers using kafka-reassign-partitions, wait until it hosts zero partitions, then shut it down and deregister it. Never just kill the process while it holds replicas.

solid answer

~40 s

Decommissioning = draining all data off the node before removing it. Steps: (1) Build a reassignment plan that rewrites every partition currently on the target broker.id to a replica set excluding it (often using a tool/script or Cruise Control's remove-broker). (2) --execute with a --throttle so the drain doesn't starve clients. (3) --verify until the broker hosts zero replicas and no leaderships; remove throttles. (4) Trigger controlled shutdown (controlled.shutdown.enable=true, the default) so leaderships migrate gracefully, then stop the process and deregister the node. Critical: do NOT remove the broker if doing so would drop any partition below its min.insync.replicas or leave a replica without enough copies — you'd risk under-replication or, with unclean leader election, data loss. Verify URP (under-replicated partitions) is zero before and after.

go deeper

for a junior

Know the order: drain replicas off, verify empty, then shut down — never kill first.

for a middle

Execute the full reassign-drain-verify flow and check URP and controlled shutdown.

for a senior

Reason about min.insync.replicas, throttle budgets, and the dead-broker variant.

for a principal

Set fleet policy: RF/min.insync defaults, automation (Cruise Control remove_broker), and durability invariants enforced in tooling.

## Goal Permanently remove a broker without losing data or violating durability guarantees. A broker holds **leader** and **follower** replicas; if you simply kill it while it owns the *only* in-sync copy of some partition, that data becomes unavailable (or lost if unclean leader election is enabled). ## Step 1 — Drain the replicas Every partition whose replica list contains the target `broker.id` must be reassigned to a replica set that *excludes* it and includes a different broker. You produce a reassignment JSON (by hand, via a helper script, or via Cruise Control's `remove_broker` / Confluent Self-Balancing). Run `kafka-reassign-partitions.sh --execute --reassignment-json-file drain.json`. New replicas bootstrap as followers and catch up into the **ISR (in-sync replica set)**. ## Step 2 — Throttle A full drain copies the broker's entire dataset across the network. Pass `--throttle <bytes/sec>` so replication doesn't saturate the NIC and degrade producer/consumer latency. Throttles are stored as dynamic configs (`leader/follower.replication.throttled.rate`). ## Step 3 — Verify Run `--verify` repeatedly until all moves report *completed*; this also clears the throttle. Then: - Confirm the broker now hosts **zero partitions** and **zero leaderships**. - Check that **under-replicated partitions (URP) = 0** across the cluster (JMX `UnderReplicatedPartitions` or `kafka-topics.sh --describe --under-replicated-partitions`). ## Step 4 — Controlled shutdown With `controlled.shutdown.enable=true` (default), when you stop the broker it asks the controller to **move any leaderships it still holds to other in-sync replicas first**, avoiding a hard leader-election storm. Then the process exits cleanly. ## Step 5 — Deregister - In **KRaft**, the node is unregistered from the controller's metadata (it ages out / you remove it from the configured voters or observers as appropriate). - In legacy **ZooKeeper** mode the ephemeral broker znode disappears on disconnect. ## Durability invariants you must not break - Never drain a partition in a way that leaves fewer than `min.insync.replicas` in-sync copies for an `acks=all` topic — producers would start failing with `NOT_ENOUGH_REPLICAS`. - And never rely on **unclean leader election** to recover a partition whose last good replica was the broker you removed: that elects an out-of-sync replica as leader and **truncates/loses** committed messages. ## Edge cases If the broker is already dead (disk failure) you can't drain it gracefully — you reassign *away from a dead broker*, and the surviving replicas re-replicate; if a partition's only surviving replica is the dead one, you face the unclean-leader-election dilemma. Always size replication factor (e.g., RF=3) so losing one broker never strands a partition.

  • How is decommissioning a healthy broker different from removing a broker that already died from disk failure?
    A healthy broker can be drained gracefully (replicas copied off, controlled shutdown). A dead broker can't serve data, so you reassign partitions away from it and surviving replicas re-replicate; if a partition's only good copy was on the dead node you face the unclean-leader-election / data-loss tradeoff.
  • What metric tells you it's safe to proceed and finish?
    UnderReplicatedPartitions (URP) should be 0, and the target broker should report zero partitions and zero leaderships. Non-zero URP means replicas are still catching up and removing the node could break durability.

saying these in an interview costs you the question

  • Just stopping/killing the broker process while it still holds replicas.
  • Ignoring min.insync.replicas — draining can push a partition below it and break acks=all producers.
  • Skipping throttling on a large drain and saturating the network.
  • Relying on unclean leader election to 'recover' after removal (that loses data).

context