skip to content

Raft Protocol and Leader Election

Kafka's pull-based Raft variant: epochs, votes, majority commit, and how the metadata leader is elected. Interviewers ask to see whether you understand the consensus, not just that Kafka uses Raft now.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is KRaft and what role does the Raft consensus protocol play in it?

level: juniorimportance: must knowfreq 70%

answer

  1. Kafka Raft = no ZooKeeper
  2. KIP-500
  3. __cluster_metadata topic
  4. controllers = voters, brokers = observers
  5. ordered metadata log with offsets+epochs

basics

~20 s

KRaft (Kafka Raft) is Kafka's built-in consensus mechanism that replaces ZooKeeper for storing cluster metadata. A small group of controllers uses a Raft-style protocol to elect a leader and agree on an ordered log of metadata changes.

solid answer

~40 s

KRaft (KRaft = Kafka Raft Metadata mode) is Kafka's self-managed metadata quorum, introduced by KIP-500, that removes the external ZooKeeper dependency. Cluster metadata (topics, partitions, configs, ACLs, broker registrations) lives in an internal replicated log, the __cluster_metadata topic. A small set of nodes with process.roles=controller form a quorum of voters. They run KRaft, a tailored variant of the Raft consensus algorithm, to elect a single active controller (the leader) and replicate an ordered, append-only log of metadata records. Other nodes (brokers) are observers that fetch and replay this log to build their in-memory metadata image. Because metadata is now an event log with offsets and epochs, recovery and propagation are faster and simpler than the ZooKeeper-based design.

go deeper

for a junior

Know that KRaft replaces ZooKeeper and uses Raft consensus to keep cluster metadata in an internal log.

for a middle

Explain voters vs observers, the __cluster_metadata topic, and that the active controller is the Raft leader.

for a senior

Discuss KIP-500 motivation, pull-based replication difference from vanilla Raft, and how offsets/epochs speed failover.

for a principal

Reason about quorum sizing, controller/broker role separation, migration from ZooKeeper, and operational trade-offs at scale.

## The problem KRaft solves Before KRaft, Kafka used **ZooKeeper** (a separate distributed coordination service) to store cluster metadata: which topics and partitions exist, which broker leads each partition, configs, ACLs, and broker liveness. This meant operators ran and tuned two systems, and metadata propagation had scaling limits. ## What KRaft is **KRaft** stands for **Kafka Raft (metadata mode)**. Proposed in **KIP-500** and made production-ready over several releases, it moves metadata into Kafka itself. - **Consensus**: a way for multiple machines to agree on a single, ordered sequence of values even if some machines fail. - **Raft**: a well-known consensus algorithm (from the paper *In Search of an Understandable Consensus Algorithm*) built around a single elected leader and a replicated log. - **Quorum**: the majority of voters needed to make progress. With 3 voters, a quorum is 2; with 5 voters, 3. ## How it works at a high level 1. Nodes configured with `process.roles=controller` are **voters**. They participate in elections and replicate the metadata log. 2. The voters elect one **active controller** (the Raft **leader**). Only the leader appends new metadata records. 3. Metadata records form an ordered, append-only log stored as the internal topic **`__cluster_metadata`** (a single partition). Each record has an **offset** and is stamped with a **leader epoch**. 4. **Brokers** (`process.roles=broker`) are **observers**: they don't vote but fetch the metadata log and replay it to build an in-memory **metadata image**. 5. A record is **committed** once a majority of voters have it; only committed records advance the **high watermark (HWM)** and become visible. ## Why it matters - One system to operate instead of two. - Faster failover and metadata propagation because metadata is an event log with offsets/epochs. - Supports far more partitions per cluster. ## Edge cases / notes - A node can be both a controller and a broker (`process.roles=broker,controller`) in small/dev setups, but production typically separates them. - KRaft is a **variant** of Raft, not vanilla Raft: replication is **pull-based** (followers Fetch from the leader, mirroring how Kafka consumers fetch), whereas classic Raft pushes via AppendEntries. - ZooKeeper mode was deprecated and removed in later Kafka 4.x; KRaft is now the only mode.

  • Where is the metadata actually stored in KRaft?
    In a single-partition internal log topic named __cluster_metadata, replicated across the controller voters; each record has an offset and a leader epoch.
  • What is the difference between a voter and an observer in KRaft?
    Voters (controllers) participate in leader elections and count toward quorum/commit; observers (brokers) only fetch and replay the committed log, they neither vote nor count toward the majority.

saying these in an interview costs you the question

  • Saying KRaft still needs ZooKeeper as a fallback
  • Claiming every broker votes in elections
  • Confusing the metadata log with a normal user topic that consumers read

context

open as a page

How does leader election work in KRaft, and what is a leader epoch?

level: middleimportance: must knowfreq 65%

basics

~20 s

When voters detect no leader, a candidate increases the epoch (a term counter), votes for itself, and sends Vote requests to peers. If a majority grants votes, it becomes leader for that epoch. The epoch is a monotonically increasing number that totally orders leadership periods.

open as a page

How does KRaft decide a metadata record is committed, and how does the high watermark advance?

level: seniorimportance: must knowfreq 50%

basics

~20 s

A record is committed once it has been replicated (fetched) by a majority of the voters. The high watermark is the highest offset known to be on a majority; the leader advances it as Fetch offsets arrive, and only records at or below the HWM are visible/applied.

open as a page

How does KRaft replicate the metadata log, and why is it described as pull-based rather than push-based Raft?

level: seniorimportance: must knowfreq 55%

basics

~20 s

In KRaft, followers and observers pull records from the leader using Fetch requests, just like Kafka consumers pull from a partition leader. Classic Raft instead has the leader push records via AppendEntries. KRaft reuses Kafka's existing fetch/replication machinery.

open as a page

How do you size a KRaft controller quorum, and what fault tolerance does each size give?

level: principalimportance: should knowfreq 45%

basics

~20 s

Use an odd number of controllers, typically 3 or 5. A quorum of N tolerates floor((N-1)/2) failures: 3 tolerate 1, 5 tolerate 2. Odd counts give the best fault tolerance per node since a majority must still be reachable to elect a leader and commit.

open as a page