skip to content

KRaft Architecture and Metadata Quorum

The KRaft design: a controller quorum keeping cluster metadata in a Raft-replicated internal topic, with process.roles deciding what each node does. Core modern-Kafka knowledge for any operations or architecture interview.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is KRaft in Apache Kafka, and what problem was it introduced to solve?

level: juniorimportance: must knowfreq 80%

answer

  1. KIP-500 removes ZooKeeper
  2. Kafka Raft = self-managed metadata
  3. __cluster_metadata internal topic
  4. faster failover, millions of partitions
  5. ZK removed in Kafka 4.0

basics

~10 s

KRaft (Kafka Raft) is Kafka's built-in consensus protocol that replaces Apache ZooKeeper for storing cluster metadata. It lets Kafka manage its own metadata internally, so you no longer run a separate ZooKeeper cluster.

solid answer

~40 s

KRaft (short for Kafka Raft Metadata mode) was introduced in KIP-500 to remove Kafka's dependency on Apache ZooKeeper. Before KRaft, Kafka stored cluster metadata (topics, partitions, broker registrations, ACLs, configs) in an external ZooKeeper ensemble, which added an extra system to operate, limited the number of partitions a cluster could hold, and made metadata propagation slow. KRaft moves this metadata into Kafka itself: a quorum of controller nodes runs a Raft-based consensus protocol (KIP-595) and stores metadata as an event log in an internal topic called __cluster_metadata. This simplifies deployment to a single system, scales to millions of partitions, and makes controller failover much faster because the new controller already has the metadata log in memory rather than reloading it from ZooKeeper.

go deeper

for a junior

Know KRaft = Kafka's replacement for ZooKeeper, stores metadata inside Kafka itself.

for a middle

Tie KRaft to KIP-500/595, the __cluster_metadata topic, and the operational wins (one system, faster failover).

for a senior

Explain the Raft quorum mechanics, why failover is faster (in-memory log vs ZK reload), and the partition-scaling argument.

for a principal

Reason about migration strategy (KIP-866), the metadata-as-event-log design, and trade-offs vs the old controller model when designing or upgrading clusters.

## Background: why metadata needs a home A Kafka cluster is a set of **brokers** (servers that store and serve message data). To coordinate, the cluster needs shared **metadata**: which topics exist, how many partitions each has, which broker is the leader for each partition, broker registrations (who is alive), configuration, and ACLs (access control lists). This metadata must be consistent across the whole cluster and survive restarts. ## The old way: ZooKeeper Historically Kafka stored that metadata in **Apache ZooKeeper**, a separate distributed coordination service you had to install, secure, and operate alongside Kafka. One broker was elected the **controller**; it watched ZooKeeper and pushed metadata changes to the other brokers. Problems: (1) you operate two systems; (2) the controller had to load all metadata from ZooKeeper on failover, which could take tens of seconds in large clusters; (3) the design capped practical partition counts (roughly the low hundreds of thousands). ## KRaft: Kafka manages its own metadata **KRaft** = **K**afka **Raft** metadata mode, introduced by **KIP-500** (the umbrella proposal) and implemented with **KIP-595** (the Raft protocol). It eliminates ZooKeeper. Metadata is stored as an **event log** — an ordered, append-only sequence of records describing every metadata change — inside a special internal topic named **`__cluster_metadata`**. This topic has a **single partition** that is replicated across a small group of **controller** nodes using the **Raft consensus algorithm**. **Raft** is a well-known consensus protocol: one node is the **leader** (the active controller), the others are **followers**; the leader appends records and replicates them, and a record is **committed** once a **majority (quorum)** of voters has it. KRaft uses a *pull*-based variant where followers fetch from the leader. ## Why it is better - **One system to run** instead of Kafka + ZooKeeper. - **Faster failover**: every controller already has the metadata log; the new leader is up to date in memory, so promotion is near-instant rather than a multi-second reload. - **Scales further**: millions of partitions become feasible because metadata propagates as incremental log records (a delta) rather than full snapshots. - **Simpler mental model**: metadata is just another Kafka log, consumed the way brokers already consume logs. ## Timeline (worth knowing) KRaft was previewed in 2.8, marked production-ready in Kafka 3.3, became the default for new clusters in 3.x, and ZooKeeper support was **removed entirely in Kafka 4.0** — so modern Kafka is KRaft-only. ## Edge cases / caveats - KRaft does not change how *data* (your records) is replicated between brokers — that still uses the normal ISR/replication mechanism. KRaft is only about **metadata**. - Migration from a ZooKeeper cluster to KRaft is a defined, staged process (KIP-866), not a flip of one flag.

  • Does KRaft change how the actual message data is replicated between brokers?
    No. KRaft only governs cluster metadata. Topic/partition record data is still replicated between brokers via the usual leader/follower ISR replication; KRaft replaces only the ZooKeeper-backed metadata layer.
  • Which Kafka version removed ZooKeeper support entirely?
    Kafka 4.0. KRaft was production-ready in 3.3 and default for new clusters thereafter; 4.0 dropped ZooKeeper mode, making KRaft the only option.

saying these in an interview costs you the question

  • Saying KRaft replaces broker-to-broker data replication (it only handles metadata)
  • Claiming KRaft still needs ZooKeeper internally
  • Confusing KRaft (consensus for metadata) with Kafka's normal ISR replication for topic data
  • Thinking KRaft is an optional plugin rather than the built-in, now-default mode

context

open as a page

Describe the __cluster_metadata topic in KRaft: how it is structured, who reads and writes it, and why it is special.

level: middleimportance: must knowfreq 65%

basics

~20 s

__cluster_metadata is the internal Kafka topic that holds all cluster metadata as an event log. It has a single partition replicated across the controller quorum via Raft. The active controller writes to it; controllers and brokers read it to stay in sync.

open as a page

Explain the process.roles, node.id, and controller.quorum.voters configuration properties in KRaft. What does each control?

level: middleimportance: must knowfreq 70%

basics

~20 s

process.roles sets whether a node is a broker, controller, or both (combined). node.id is the node's unique integer ID in the cluster. controller.quorum.voters lists the controller nodes (id@host:port) that form the metadata Raft quorum, so every node knows who the controllers are.

open as a page

How does metadata propagate from the active controller to brokers in KRaft, and how does this differ from the ZooKeeper-era controller model?

level: seniorimportance: should knowfreq 45%

basics

~20 s

In KRaft, brokers pull metadata changes by fetching the __cluster_metadata log from the active controller and applying records incrementally to a local metadata cache. In the ZooKeeper era, the controller pushed full LeaderAndIsr/UpdateMetadata RPCs to brokers, which was slower and harder to scale.

open as a page

In a KRaft controller quorum, how is a metadata write committed, how many controller failures can the cluster tolerate, and what happens when quorum is lost?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A metadata record is committed once a majority of controller voters have persisted it. With N voters the cluster tolerates floor((N-1)/2) failures (so 3 tolerate 1, 5 tolerate 2). If a majority is lost, no new metadata can be committed and the controller becomes read-only until quorum returns.

open as a page