What is the Kafka controller, and how does its role differ between ZooKeeper mode and KRaft mode?
answer
- controller = cluster brain
- leader election from ISR
- ZK: 1 broker via /controller znode
- KRaft: Raft quorum + __cluster_metadata log
- process.roles + controller.quorum.voters
basics
~20 sThe controller is the broker/node responsible for cluster-wide management: partition leader election, tracking broker liveness, and propagating metadata. In ZooKeeper mode one elected broker is the controller; in KRaft, dedicated controller nodes run a Raft quorum that stores metadata in a log.
solid answer
~40 sThe controller is Kafka's cluster coordinator. It detects broker failures, elects new partition leaders from the in-sync replicas, manages replica reassignment, and pushes metadata updates to all brokers. In the legacy ZooKeeper architecture, exactly one broker is elected controller (via a ZooKeeper ephemeral znode); it reads/writes cluster state in ZooKeeper and on its loss another broker takes over. KRaft (KIP-500) removes ZooKeeper: a small set of **controller** nodes form a **Raft** quorum and store the entire cluster metadata as a replicated **__cluster_metadata** log. One controller is the active leader; the others are hot standbys that already have the log, so failover is near-instant and metadata scales to millions of partitions. Brokers become metadata followers that replay the log. Roles are set by process.roles=broker, controller, or both; the quorum is configured via controller.quorum.voters.
go deeper
Know there is one controller coordinating the cluster and electing partition leaders.
Distinguish ZooKeeper single-controller from KRaft quorum at a high level.
Explain Raft metadata log, failover speed, quorum sizing, and process.roles.
Reason about KRaft migration, quorum topology/fault tolerance, and metadata scaling limits.
## What the controller does A Kafka cluster needs one component making cluster-wide decisions so brokers stay consistent. That component is the **controller**. Its core duties: - **Broker liveness**: detect when a broker dies or rejoins. - **Leader election**: for each partition, pick a leader replica from the **ISR** (in-sync replicas) when the current leader fails. - **Replica/partition management**: handle reassignments, topic creation/deletion, and partition expansion. - **Metadata propagation**: push the resulting state (who leads what, ISR membership) to all brokers so clients can be routed correctly. ## ZooKeeper mode (legacy) Historically Kafka stored cluster metadata in **ZooKeeper**, an external coordination service. Brokers race to create an ephemeral **/controller** znode; the winner is the single active controller. It maintains state in ZooKeeper and broadcasts updates via control requests (LeaderAndIsr, UpdateMetadata). Drawbacks: controller failover requires the new controller to **reload all metadata from ZooKeeper**, which is slow and bounds the cluster's partition count; ZooKeeper is a separate system to operate and secure. ## KRaft mode (KIP-500) **KRaft** (Kafka Raft) removes ZooKeeper. Metadata lives inside Kafka itself as a special replicated log topic **__cluster_metadata**, managed by the **Raft** consensus protocol. Key pieces: - **Controller nodes**: a dedicated quorum (commonly 3 or 5) of nodes whose job is metadata. Set via `process.roles=controller`. - **Quorum**: configured by `controller.quorum.voters` (id@host:port list). One controller is the Raft **leader** (active controller); the rest are **followers**/hot standbys that already hold the metadata log. - **Brokers**: `process.roles=broker` nodes act as metadata **followers**, replaying the __cluster_metadata log to learn cluster state. A node can be `broker,controller` (combined mode) for small/dev clusters. Benefits: because standby controllers already have the log, failover is near-instant (no full reload); metadata is an event log enabling far larger partition counts and simpler operations (no ZooKeeper). KRaft is the default/required architecture in current Kafka (ZooKeeper removed in Kafka 4.0). ## node.id and identity In KRaft, `node.id` uniquely identifies each node across broker and controller roles, replacing broker.id's role. Controller voters reference these ids. ## Edge cases / nuances - **Active vs standby**: only one controller is active; clients/brokers never talk to controllers for data, only metadata flows. - **Unclean leader election** (`unclean.leader.election.enable`) is still a controller decision: when no ISR replica survives, allowing an out-of-sync replica to become leader trades data loss for availability. - **Combined mode** is fine for dev but discouraged for large prod clusters (isolation of the metadata quorum). - A KRaft quorum tolerates floor((N-1)/2) failures, so 3 voters tolerate 1, 5 tolerate 2.
- Why is controller failover faster in KRaft than in ZooKeeper mode?Standby controllers already replicate the __cluster_metadata log, so a new active controller is already caught up; ZooKeeper-mode failover requires reloading all metadata from ZooKeeper.
- How many controller nodes can fail in a 5-node KRaft quorum while staying available?Two. A Raft quorum of N tolerates floor((N-1)/2) failures, so 5 voters need 3 to maintain a majority and tolerate 2 losses.
- What config decides whether a node is a broker, controller, or both?process.roles (e.g. broker, controller, or broker,controller), with the quorum members listed in controller.quorum.voters.
saying these in an interview costs you the question
- Saying every broker is a controller simultaneously
- Claiming KRaft still uses ZooKeeper
- Thinking clients send produce/fetch to the controller
- Confusing the cluster controller with a partition leader
- Saying the controller stores the actual message data