What does controlled.shutdown.enable do during a broker restart, and what governs whether it succeeds?
answer
- default true; ControlledShutdown RPC to controller
- migrates leaderships to in-sync replicas before exit
- max.retries + retry.backoff.ms
- needs RF>=2 and healthy ISR to work
- moves leaders not data (still drain for decommission)
basics
~10 sOn shutdown, the broker asks the controller to move its partition leaderships to other in-sync replicas before it exits, so leader failover is graceful instead of a sudden election. It's on by default.
solid answer
~40 scontrolled.shutdown.enable (default true) makes a stopping broker send a ControlledShutdown request to the controller, which proactively migrates leadership of every partition the broker leads to another in-sync replica, then lets the broker exit. This avoids a burst of unclean/abrupt leader elections, minimizes producer/consumer disruption, and prevents the broker from being elected leader of partitions whose data it would soon abandon. It retries up to controlled.shutdown.max.retries with controlled.shutdown.retry.backoff.ms between attempts. It can only move a leadership to a replica that is currently in the ISR — if a partition led by this broker has no other in-sync replica, that leadership can't migrate cleanly and the broker may shut down anyway, leaving the partition leaderless until an election occurs. So replication factor >= 2 and healthy ISRs are prerequisites for controlled shutdown to actually be 'controlled'.
go deeper
Know it gracefully hands off leadership before a broker stops, and it's on by default.
Explain the controller RPC, ISR requirement, and the retry settings.
Tie it to rolling restarts and explain failure modes (RF=1, lagging ISR).
Design restart automation that waits for ISR health and verifies leadership drain between steps.
**The problem it solves:** When a broker stops, every partition it was **leader** for suddenly has no leader. Without coordination, the **controller** detects the broker is gone and triggers leader elections reactively — a brief window where those partitions are unavailable, and a risk of electing a less-ideal (or, if misconfigured, out-of-sync) replica. **What controlled shutdown does:** With `controlled.shutdown.enable=true` (the default since long ago), a broker that is asked to stop first sends a **ControlledShutdown** RPC to the cluster **controller**. The controller, *before* the broker dies, **moves leadership of each partition the broker leads to another replica that is in the ISR (in-sync replica set)**. Only once leaderships are migrated (or retries exhausted) does the broker complete its exit. The net effect: leadership transitions are planned and fast, clients reconnect to the new leaders with minimal interruption, and there's no risk of the departing broker being (re)elected leader. **Tunables:** - `controlled.shutdown.enable` — master switch (default `true`). - `controlled.shutdown.max.retries` — how many times the broker retries the request if the controller is busy/unreachable. - `controlled.shutdown.retry.backoff.ms` — wait between retries. **What governs success:** The controller can only hand a leadership to a replica that is **currently in the ISR**. Therefore: - Partitions with **replication factor 1** (no follower) cannot migrate leadership at all — they go offline when the broker stops. - Partitions whose followers have fallen out of the ISR (lagging) can't receive leadership until they catch up. - If retries are exhausted, the broker may shut down **anyway**, leaving affected partitions to be handled by a normal (reactive) controller election. **Interplay with rolling restarts:** Controlled shutdown is what makes **rolling upgrades** safe — you stop one broker at a time, leadership drains off it gracefully, you upgrade/restart it, it rejoins and re-enters ISRs, then you move to the next. Doing this without controlled shutdown (or with unhealthy ISRs) causes repeated availability blips. **Relationship to decommissioning:** Controlled shutdown handles **leaderships at stop time**; it does NOT move the broker's **data/replicas** off the node. For a permanent decommission you still must reassign replicas away first (drain) — controlled shutdown alone leaves the broker's replicas behind and under-replicated once it's gone. **Edge case — last in-sync replica:** If the shutting-down broker is the *only* in-sync replica for a partition (others lagging), controlled shutdown cannot cleanly transfer it; whether the partition recovers depends on whether `unclean.leader.election.enable` permits electing a lagging replica (with data-loss risk) or the partition stays offline until a good replica returns.
- Does controlled shutdown move the broker's replicas (data) off the node?No. It only migrates partition leaderships at stop time. The broker's follower/leader replicas (the data) stay assigned to it; for a permanent removal you must reassign replicas away first, or the cluster becomes under-replicated once the broker is gone.
- Why might controlled shutdown fail to migrate a partition's leadership?Because the controller can only hand leadership to a replica in the ISR. If the partition is RF=1 or all followers are lagging out of the ISR, there's no eligible target, so leadership can't migrate cleanly.
saying these in an interview costs you the question
- Saying controlled shutdown moves the broker's data/replicas (it only moves leaderships).
- Assuming it always succeeds regardless of ISR health or replication factor.
- Thinking it must be manually enabled — it's on by default.
- Confusing it with a decommission/drain step.