skip to content

In KRaft mode, what is the active controller and what role does it play in managing cluster metadata?

level: juniorimportance: must knowfreq 70%

answer

  1. Single metadata writer
  2. Raft leader of __cluster_metadata
  3. TopicRecord / PartitionRecord / BrokerRegistration
  4. brokers replay = event sourcing
  5. no ZooKeeper (KIP-500)

basics

~20 s

The active controller is the single elected controller node that is the only one allowed to write cluster metadata (topics, partitions, broker registrations) by appending records to a replicated log. Other nodes follow and replay that log.

solid answer

~40 s

In KRaft (Kafka Raft) mode, controllers form a quorum, and exactly one of them is elected the active controller (the Raft leader of the __cluster_metadata topic). The active controller is the sole metadata writer: all cluster-state changes (create topic, reassign partition, register broker) become records it appends to the replicated metadata log. The other controllers are followers that replicate the log; brokers are observers that fetch and replay it. This single-writer design serializes all metadata changes through one ordered log, giving every node a consistent, eventually-identical view of cluster state without ZooKeeper. If the active controller fails, the Raft quorum elects a new leader from the up-to-date followers, and it resumes writing.

go deeper

for a junior

Know that one elected controller writes metadata and others follow; ZooKeeper is gone.

for a middle

Explain the metadata log, record types, and brokers replaying it as an event-sourced state machine.

for a senior

Discuss the Raft quorum, single-writer rationale, commit-on-majority, and failover semantics.

for a principal

Reason about availability vs consistency tradeoffs, quorum sizing, and how this design removes a whole class of ZooKeeper coordination problems.

## Metadata, and where it lives Apache Kafka stores not just user data but also *metadata* — which topics exist, how many partitions each has, where partition replicas live, which brokers are alive, ACLs, configs, and so on. Historically this metadata lived in ZooKeeper, an external coordination service. **KRaft** (KRaft = Kafka Raft Metadata mode, introduced by KIP-500) removes ZooKeeper and stores metadata inside Kafka itself. ## The controllers In KRaft you designate certain nodes with `process.roles=controller` (or `broker,controller` for combined nodes). These controller nodes form a *quorum* — typically 3 or 5 — that runs the Raft consensus protocol over an internal topic named `__cluster_metadata` (a single-partition replicated log). ## The active controller Raft elects exactly one leader of that log; in Kafka terminology this leader is the **active controller**. It is the *only* node permitted to *write* metadata. Every metadata change is expressed as a record — for example: - a `TopicRecord` (a new topic), - a `PartitionRecord` (a partition's replica assignment and leader), - or a `BrokerRegistration`/`RegisterBrokerRecord` (a broker joining). The active controller appends these records, in order, to the metadata log. ## Followers and brokers - The other controllers are Raft *followers*; they replicate the log so they are ready to take over. - Brokers (the nodes that actually serve produce/fetch traffic) act as *observers*: they continuously fetch the metadata log and **replay** it to rebuild their in-memory picture of the cluster. This is an ***event-sourced state machine***: state = result of applying the ordered stream of records. ## Why a single writer Funneling all writes through one node and one ordered log means there is one authoritative order of events. Every node that replays the same prefix of the log reaches the same state, so the cluster converges to a consistent view without distributed locks or a separate coordination system. ## Failover If the active controller crashes, the Raft quorum holds an election and promotes an up-to-date follower to active controller. Because followers already have the replicated log, the new leader can resume appending almost immediately. A **majority (quorum)** of controllers must be available for writes to proceed. ## Edge cases to know - a record is only durable/committed once a majority of the quorum has it; - brokers that lag are eventually consistent (they catch up by replaying); - and the active controller is a logical role distinct from a partition *leader* for user data — they are different leadership concepts.

  • How is a new active controller chosen if the current one fails?
    The Raft quorum runs a leader election; an up-to-date follower with the latest committed log wins and becomes the new active controller, then resumes appending records.
  • What is the difference between the active controller and a partition leader?
    The active controller is the leader of the metadata log and writes cluster metadata. A partition leader handles produce/fetch for one user-data partition. They are independent roles, often on different nodes.

saying these in an interview costs you the question

  • Saying KRaft still needs ZooKeeper
  • Claiming multiple controllers can write metadata simultaneously
  • Confusing the active controller with a topic-partition leader
  • Saying every broker is also a controller

context