skip to content

Walk through what happens during first-boot quorum formation in a KRaft cluster: how do controllers find each other, elect a leader, and reach the point where brokers can register?

level: seniorimportance: should knowfreq 28%

answer

  1. Validate meta.properties → bind controller listener
  2. Raft Vote RPC; persist vote in quorum-state
  3. Majority ceil((n+1)/2) elects active controller
  4. Leader writes genesis metadata.version to log
  5. Brokers send BrokerRegistration + heartbeats
  6. No quorum → no leader → brokers can't register

basics

~20 s

On first boot each controller reads its formatted meta.properties (cluster.id, node.id) and the voter set (static voters config or bootstrap.checkpoint). Controllers contact each other on the controller listener, run a Raft election to pick an active controller (leader), and once a quorum agrees, brokers can register and the cluster is live.

solid answer

~50 s

First boot: each controller starts, validates its meta.properties (matching cluster.id, its node.id, and directory.id), and learns the voter set — from static controller.quorum.voters or from the bootstrap.checkpoint genesis snapshot in a dynamic cluster. Controllers connect on the controller.listener.names endpoint and run KRaft's Raft election: candidates request votes (Vote RPC), each node persists its vote in quorum-state, and a candidate winning a majority of the voter set becomes the active controller (Raft leader) for an epoch. The leader appends the initial metadata records (bootstrap metadata.version, etc.) to __cluster_metadata-0 and replicates them. Brokers, started with controller.quorum.bootstrap.servers / voters, send a BrokerRegistration request to the active controller, get a broker epoch, and begin sending periodic heartbeats. Only once a majority of controllers is up can a leader be elected — a quorum needs ceil((n+1)/2) of n voters; below that, the cluster will not bootstrap and brokers cannot register.

go deeper

for a junior

Know that controllers must start and elect a leader before the cluster is usable.

for a middle

Describe the format→elect→register sequence and the majority requirement.

for a senior

Detail Raft Vote/Fetch RPCs, quorum-state persistence, epochs, and BrokerRegistration/heartbeat.

for a principal

Reason about quorum sizing, availability under partial outages, split-brain prevention, and combined-mode election ordering.

## Setup recap Before first boot every node has been formatted: its `log.dirs`/`metadata.log.dir` contains `meta.properties` with the shared **cluster.id**, this node's **node.id**, and (for dynamic quorums) a **directory.id**. Controllers also know the **voter set** — either statically via `controller.quorum.voters=id@host:port,...` or, for KIP-853 dynamic clusters, from the **`bootstrap.checkpoint`** genesis snapshot written at format time. ## Step 1 — node starts and validates identity Each controller process boots, reads `meta.properties`, and checks the cluster.id is consistent across all its directories and with the quorum it's joining. A mismatch → `InconsistentClusterIdException` and the node exits. It binds its **controller listener** (named by `controller.listener.names`, e.g. `CONTROLLER` on port 9093) — controllers talk to each other and to brokers over this listener. ## Step 2 — Raft leader election KRaft uses a **Raft** consensus protocol over the controller quorum. Raft works in **terms/epochs**: - A controller with no leader becomes a **candidate**, increments the epoch, votes for itself, and sends **Vote** RPCs to the other voters. - Each voter grants at most one vote per epoch and **persists** its choice in the **`quorum-state`** file so a restart can't make it vote twice (which would violate safety). - A candidate that collects votes from a **majority** of the voter set becomes the **active controller** (Raft leader) for that epoch. Majority = `ceil((n+1)/2)` of `n` voters: 2-of-3, 3-of-5, etc. Election requires a **quorum to be reachable**. With 3 controllers you need at least 2 up; with only 1 of 3 up, no leader is elected and the cluster does not come alive. ## Step 3 — leader writes genesis metadata The newly elected active controller appends the initial metadata records to **`__cluster_metadata-0`**: the bootstrap **`metadata.version`** (feature level set at format time / from bootstrap metadata), and any seed records. Followers replicate via **Fetch** RPCs (in KRaft followers pull from the leader). High watermark advances once a majority has the records, making them committed. ## Step 4 — brokers register Brokers (nodes with `broker` in `process.roles`) start with `controller.quorum.bootstrap.servers` (or `controller.quorum.voters`) so they know whom to contact. Each broker sends a **BrokerRegistration** request to the active controller. The controller assigns the broker a **broker epoch** and records the registration in the metadata log. The broker then sends periodic **BrokerHeartbeat** requests; missing heartbeats (beyond `broker.session.timeout.ms`) cause the controller to fence the broker. Once registered and unfenced, the broker is part of the cluster and can host partitions. ## Combined-mode nuance With `process.roles=broker,controller`, a node is both a voter and a broker in the same JVM. Election among the controller portions happens first; the broker portion registers with whoever is the active controller (possibly itself). ## Why ordering matters Brokers cannot meaningfully register until an active controller exists, because registration is a metadata write that must be committed by the quorum. So the dependency chain on first boot is: **format → controllers reachable → leader elected → genesis metadata committed → brokers register/heartbeat → cluster ready.** ## Edge cases - If fewer than a majority of controllers start, you'll see repeated election attempts and brokers stuck retrying registration. - A split or wrong `controller.quorum.voters` (mismatched ids/endpoints across nodes) prevents convergence. - `quorum-state` corruption can make a node unable to participate safely; it must recover its Raft state.

  • Why is the vote persisted in quorum-state across restarts?
    Raft safety requires a node to never vote twice in the same epoch. Persisting the vote means a crash-and-restart can't 'forget' it and grant a second vote, which could otherwise elect two leaders and split-brain the quorum.
  • How many controllers must be up for a 5-node controller quorum to elect a leader?
    A majority: ceil((5+1)/2) = 3. With only 2 reachable, no leader can be elected and the cluster won't bootstrap or accept metadata writes.

saying these in an interview costs you the question

  • Saying brokers can register before any active controller exists.
  • Claiming a single reachable controller out of three is enough to form the quorum.
  • Describing controller replication as ISR/leader-push instead of Raft follower-fetch.
  • Forgetting that the vote is durably persisted (quorum-state) for safety.
  • Confusing broker epoch with controller/Raft epoch.

context