skip to content

Broker Registration and Heartbeats

How a broker joins the cluster and stays unfenced by heartbeating the controller, and what fencing means for its traffic. Comes up in incident-style questions about brokers that look up but serve nothing.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

In a KRaft Kafka cluster, how does a broker tell the controller it is still alive, and what RPC is involved?

level: juniorimportance: must knowfreq 55%

answer

  1. BrokerRegistration then BrokerHeartbeat RPCs
  2. interval 2000ms, session timeout 9000ms
  3. no ZooKeeper ephemeral znodes
  4. miss heartbeats -> fenced
  5. response carries fence/shutdown flags

basics

~10 s

Each broker periodically sends a BrokerHeartbeat request to the active controller. The interval is set by broker.heartbeat.interval.ms (default 2000 ms). If heartbeats stop, the controller eventually fences the broker.

solid answer

~40 s

In KRaft, brokers are no longer tracked via ZooKeeper ephemeral znodes. Instead each broker first sends a BrokerRegistration RPC to the active controller, then keeps itself alive by periodically sending BrokerHeartbeat RPCs. The cadence is governed by broker.heartbeat.interval.ms (default 2000 ms). The controller tracks the last-heartbeat time; if a broker misses heartbeats for longer than broker.session.timeout.ms (default 9000 ms), the controller fences the broker — it is excluded from ISRs and stops being chosen as a leader. The heartbeat response also tells the broker whether it should remain fenced, begin controlled shutdown, etc. Heartbeats carry the broker epoch so the controller can detect stale or restarted brokers.

go deeper

for a junior

Know the two RPCs (register, heartbeat), the two configs and their defaults, and that missing heartbeats leads to fencing.

for a middle

Explain the registration->heartbeat->unfence flow and that the response drives the broker state machine.

for a senior

Discuss broker epochs/incarnation ids, why timeout must be a multiple of the interval, and partition-vs-death ambiguity.

for a principal

Reason about quorum-side liveness tracking replacing ZK, failure-detector tradeoffs, and tuning for cloud network jitter.

## Background: why heartbeats exist Kafka clusters must continuously know which brokers are alive so the controller can keep partition leadership and the in-sync replica (ISR) sets correct. In the old **ZooKeeper** mode, each broker held an *ephemeral znode* in ZooKeeper; if the broker's ZK session expired, the znode vanished and the controller learned the broker was gone. ZooKeeper effectively did the liveness tracking. **KRaft** (Kafka Raft, KIP-500) removes ZooKeeper. Liveness is now tracked by the **controller quorum** itself using an explicit RPC protocol. ## The lifecycle 1. **Registration** — On startup a broker sends a `BrokerRegistration` RPC to the active controller. It includes the broker id, the listeners/endpoints, supported feature ranges, the cluster id, and an **incarnation id** (a fresh UUID per process start). The controller assigns a **broker epoch** (the log offset at which the registration record was committed) and records the broker, initially **fenced**. 2. **Heartbeats** — The broker then loops, sending `BrokerHeartbeat` RPCs every `broker.heartbeat.interval.ms` (default **2000 ms**). Each heartbeat carries the broker id, the **broker epoch**, the current metadata offset the broker has caught up to, and flags such as `wantFence` and `wantShutDown`. 3. **Staying unfenced** — Once the broker has caught up to the metadata log and signals readiness, the controller **unfences** it, making it eligible to host leaders and join ISRs. 4. **Fencing on timeout** — The controller records the time of the last heartbeat. If no heartbeat arrives within `broker.session.timeout.ms` (default **9000 ms**), the controller **fences** the broker: it is removed from ISRs and is no longer eligible for leadership. The broker is not deleted — it can recover by resuming heartbeats. ## The heartbeat response The controller's `BrokerHeartbeat` response is the back-channel that drives the broker's state machine. It returns booleans like `isCaughtUp`, `isFenced`, and `shouldShutDown`, telling the broker whether it may unfence, must stay fenced, or may finish a controlled shutdown. ## Key configs - `broker.heartbeat.interval.ms` (default 2000) — how often the broker pings. - `broker.session.timeout.ms` (default 9000) — how long the controller waits before fencing. Must be a comfortable multiple of the interval so a single dropped heartbeat doesn't fence a healthy broker. ## Edge cases - A network partition between broker and controller looks identical to a dead broker — the broker gets fenced even if it is otherwise healthy. - A restarted broker sends a new incarnation id; the controller issues a new epoch, so stale heartbeats from a previous incarnation are rejected.

  • What replaced the ZooKeeper ephemeral znode mechanism for broker liveness in KRaft?
    Explicit BrokerRegistration + periodic BrokerHeartbeat RPCs to the active controller, which tracks last-heartbeat time and fences brokers that miss the session timeout.
  • What happens if a broker misses a single heartbeat?
    Nothing immediately — the controller only fences after broker.session.timeout.ms (9000 ms) elapses with no heartbeat, which is several intervals, so one dropped heartbeat is tolerated.

saying these in an interview costs you the question

  • Saying brokers still use ZooKeeper ephemeral znodes in KRaft mode
  • Claiming a single missed heartbeat immediately fences the broker
  • Confusing broker.heartbeat.interval.ms with broker.session.timeout.ms
  • Saying the heartbeat goes to ZooKeeper rather than the active controller

context

open as a page

Explain fencing vs unfencing of a broker in KRaft: what triggers each transition and what are the operational effects?

level: middleimportance: must knowfreq 50%

basics

~10 s

A fenced broker is registered but excluded from leadership and ISRs. Brokers start fenced; they become unfenced once caught up to the metadata log and heartbeating normally. Missing heartbeats or controlled shutdown re-fences them.

open as a page

What is a broker epoch in KRaft, how is it assigned, and what problem does it solve during broker restarts?

level: seniorimportance: should knowfreq 35%

basics

~20 s

A broker epoch is a monotonically increasing id the controller assigns at registration (the metadata log offset of the registration record). Every heartbeat carries it, so the controller can reject stale requests from a previous broker incarnation.

open as a page

Walk through how controlled shutdown works for a broker in KRaft and why it matters for availability.

level: seniorimportance: should knowfreq 30%

basics

~10 s

On graceful stop, the broker signals shutdown via its heartbeat. The controller moves leadership off it to in-sync replicas before it exits, avoiding abrupt leader elections and minimizing produce/consume disruption.

open as a page

You operate a KRaft cluster in a cloud with occasional network jitter. How do you reason about tuning broker.heartbeat.interval.ms and broker.session.timeout.ms, and what are the failure modes at each extreme?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

Keep the session timeout a comfortable multiple of the heartbeat interval (defaults: 2000 ms / 9000 ms ~= 4.5x). Too short a timeout causes false fencing on jitter; too long delays detection of real failures.

open as a page