skip to content

Walk through how controlled shutdown works for a broker in KRaft and why it matters for availability.

level: seniorimportance: should knowfreq 30%

answer

  1. SIGTERM -> wantShutDown in heartbeat
  2. controller moves leaders off first
  3. vs abrupt: wait full session.timeout, partitions leaderless
  4. controlled.shutdown.enable default true
  5. rolling restart: wait URP=0 between brokers

basics

~10 s

On graceful stop, the broker signals shutdown via its heartbeat. The controller moves leadership off it to in-sync replicas before it exits, avoiding abrupt leader elections and minimizing produce/consume disruption.

solid answer

~40 s

Controlled shutdown is the graceful-stop path that avoids the availability hit of an abrupt failure. When the broker is asked to stop (SIGTERM), instead of just dying, it sets wantShutDown in its BrokerHeartbeat. The active controller responds by proactively reassigning leadership for every partition the broker leads to other in-sync replicas and removing it from ISRs, all through metadata records. The controller signals completion via the heartbeat response (shouldShutDown), and only then does the broker exit. This differs from an unclean stop where partitions stay leaderless until the session.timeout fences the broker (default 9000 ms), during which producers/consumers see leadership errors. controlled.shutdown.enable (default true) governs the behavior. Doing rolling restarts one broker at a time, waiting for under-replicated partitions to return to zero between brokers, keeps availability high.

go deeper

for a junior

Know that graceful shutdown moves leadership away first and is better than killing the process.

for a middle

Explain the wantShutDown heartbeat flag and the contrast with abrupt failure waiting on session timeout.

for a senior

Detail the full flow, controlled.shutdown.enable, and the rolling-restart URP=0 procedure.

for a principal

Reason about availability math across rolling upgrades, min.insync.replicas interactions, and the limits when replication is insufficient.

## The two ways a broker can stop - **Unclean / abrupt**: the process dies (kill -9, crash, power loss). The controller does not find out until heartbeats stop arriving and `broker.session.timeout.ms` (default 9000 ms) elapses. Throughout that window, every partition the broker *led* has no leader, so producers and consumers for those partitions get errors (e.g. `NOT_LEADER_OR_FOLLOWER`) and must wait for new leaders to be elected. - **Controlled / graceful**: the broker coordinates with the controller *before* exiting so leadership is moved away first, shrinking the unavailability window to near zero. ## The controlled-shutdown flow in KRaft In ZooKeeper mode there was a dedicated `ControlledShutdown` RPC. In **KRaft**, controlled shutdown is folded into the **heartbeat protocol**: 1. The broker receives a shutdown signal (typically **SIGTERM** from the service manager or `kafka-server-stop.sh`). 2. On its next `BrokerHeartbeat`, the broker sets **`wantShutDown = true`**, telling the controller it intends to leave. 3. The active controller begins moving work off it: for each partition the broker **leads**, it elects a new leader from the remaining **in-sync replicas** and writes the corresponding metadata records. It also removes the broker from ISRs. 4. The controller indicates progress/completion in the heartbeat response (`shouldShutDown`). The broker keeps heartbeating until the controller confirms there is nothing left to move. 5. The broker then **exits cleanly**. Because leadership already moved, clients experience only brief, targeted metadata refreshes rather than a timeout-length outage. ## Config knobs - **`controlled.shutdown.enable`** (default **true**): turns the graceful path on. If false, every stop behaves like an abrupt failure. - In KRaft the legacy `controlled.shutdown.max.retries` / `controlled.shutdown.retry.backoff.ms` are less central because coordination happens through ongoing heartbeats rather than discrete retried RPCs, but the operational goal is the same. ## Why it matters for availability Without controlled shutdown, a routine **rolling restart** (config change, version upgrade) would cause every restart to incur up to a session-timeout of partial unavailability. With it, the window is the time to elect and propagate new leaders — typically sub-second. The standard operational recipe: 1. Stop one broker gracefully. 2. Wait until the cluster's **under-replicated partitions** count returns to 0 and the restarted broker is **unfenced** and caught up. 3. Only then move to the next broker. This preserves `min.insync.replicas` guarantees and avoids ever taking two replicas of the same partition down at once. ## Edge cases - If a partition led by the shutting-down broker has **no other in-sync replica**, the controller cannot cleanly move leadership; that partition will go offline (or require unclean leader election) regardless of controlled shutdown. Controlled shutdown reduces but cannot eliminate availability loss when replication is insufficient. - After exit, the broker is fenced; on restart it re-registers (new epoch), catches up, and is unfenced before taking leadership again.

  • How does controlled shutdown differ mechanically in KRaft versus ZooKeeper mode?
    ZK mode used a dedicated ControlledShutdown RPC; KRaft folds the intent into the BrokerHeartbeat via a wantShutDown flag, with the controller confirming completion in the heartbeat response before the broker exits.
  • During a rolling restart, what metric should you watch before stopping the next broker?
    Under-replicated partitions (URP) should be back to 0, and the just-restarted broker should be unfenced and caught up, ensuring no partition is left with insufficient in-sync replicas.

saying these in an interview costs you the question

  • Claiming KRaft still uses a separate ControlledShutdown RPC instead of the heartbeat flag
  • Saying controlled shutdown eliminates all unavailability even when a partition has only one in-sync replica
  • Confusing controlled shutdown (graceful) with the session-timeout fencing path (abrupt)
  • Restarting multiple brokers at once during a rolling restart

context