Walk through what happens, in terms of metadata-log records, when a broker starts up and joins a KRaft cluster.
answer
- BrokerRegistration RPC + incarnation id
- RegisterBrokerRecord appended
- starts fenced, not leader-eligible
- BrokerHeartbeat -> session timeout liveness
- unfence record after catch-up; FenceBrokerRecord on miss
basics
~20 sThe broker sends a registration request to the active controller, which appends a RegisterBrokerRecord (broker id, endpoints, incarnation) to the log. The broker starts fenced, sends periodic heartbeats, and once caught up the controller appends an unfence record so it can host partition leaders.
solid answer
~40 sOn startup a broker connects to the controller quorum and issues a `BrokerRegistration` request with its node id, listener endpoints, supported features, and a fresh *incarnation id*. The active controller validates it (e.g. no conflicting live incarnation of the same id) and appends a `RegisterBrokerRecord` to `__cluster_metadata`. The broker enters the cluster *fenced* — present in metadata but not eligible to lead partitions. It then sends periodic `BrokerHeartbeat` requests; the controller tracks liveness via `broker.session.timeout.ms`. The broker must also catch up by replaying the metadata log to the current offset. Once it is caught up and heartbeating, the controller appends an unfence change (`BrokerRegistrationChangeRecord`/`UnfenceBrokerRecord`), making it eligible to be assigned leadership. If heartbeats stop, the controller fences it again (`FenceBrokerRecord`) and moves leadership elsewhere.
go deeper
Know a starting broker registers with the controller and only later becomes able to host leaders.
Describe RegisterBrokerRecord, the fenced state, and heartbeats.
Explain incarnation id, catch-up-before-unfence, session timeout, and re-fencing on heartbeat loss.
Discuss how registration/fencing as ordered records yields a reconstructable membership history and interacts with leader-election and controlled shutdown.
## Registration request When a Kafka broker process starts in KRaft mode, it does not write to ZooKeeper (there is none). Instead it contacts the controller quorum and sends a `BrokerRegistration` RPC containing: - its `node.id`, - its advertised listener endpoints, - the broker features/versions it supports, - and an **incarnation id** — a UUID generated for *this* process start. The incarnation id lets the controller distinguish a brand-new start from a stale previous registration of the same broker id (important after crashes/restarts). ## The record The active controller — the sole writer — validates the request (rejecting, for example, a duplicate live incarnation or an incompatible feature set) and appends a **RegisterBrokerRecord** to the `__cluster_metadata` log. Once committed (quorum majority), every node replaying the log learns this broker exists. ## Fenced on arrival A freshly registered broker is **fenced**: it appears in metadata but is *not* eligible to host partition leaders. Why? It may not yet have replayed the metadata log up to the current offset, so it doesn't yet know about all the partitions it should serve. Letting it lead immediately could serve stale or missing state. ## Heartbeats and catch-up The broker periodically sends **BrokerHeartbeat** requests to the controller. These: - (a) prove the broker is alive (the controller uses `broker.session.timeout.ms` to decide liveness), and - (b) report the broker's current metadata offset so the controller knows how caught-up it is. The broker concurrently fetches and replays the metadata log to reach the current offset. ## Unfencing When the broker is sufficiently caught up and heartbeating normally, the active controller appends a change record (a `BrokerRegistrationChangeRecord`, conceptually an *unfence*) flipping it to active. Now the controller may assign it partition leadership, and subsequent `PartitionChangeRecord`s can name it as a leader. ## Failure / re-fencing If the controller stops receiving heartbeats within the session timeout, it appends a **FenceBrokerRecord**, marking the broker fenced again. Fencing triggers leader re-election for partitions that broker led — those `PartitionChangeRecord`s move leadership to in-sync replicas on other brokers. If the broker comes back, it re-registers with a new incarnation id and repeats the cycle. **Controlled shutdown** is similar but graceful: the broker asks to be fenced cleanly so leaders are migrated before it stops. ## Edge cases - a broker restarting fast may briefly have two incarnations known to the controller — the new incarnation supersedes the old; - an unclean broker that never catches up stays fenced and never gets leadership; - and because all of this is recorded as ordered log records, the join/leave history is fully reconstructable by replaying the log.
- What is the purpose of the incarnation id in broker registration?It uniquely identifies a specific process start, letting the controller distinguish a fresh registration from a stale prior one for the same broker id, e.g. after a crash and quick restart.
- Why is a newly registered broker fenced until it catches up?Until it has replayed the metadata log to the current offset, it doesn't know all partitions/state it should host, so leading would risk serving stale or missing metadata.
- What config governs how long the controller waits before fencing a silent broker?broker.session.timeout.ms — if no heartbeat arrives within it, the controller fences the broker and re-elects leaders for its partitions.
saying these in an interview costs you the question
- Saying the broker writes its own registration to the log
- Claiming a broker can lead partitions the instant it registers
- Saying registration goes through ZooKeeper
- Forgetting heartbeats / the session-timeout fencing mechanism