What controller-health signals should you alert on in a Kafka cluster, and how does this differ between ZooKeeper-based and KRaft clusters?
answer
- ActiveControllerCount summed = exactly 1
- 0 = no controller, >1 = split brain
- Controller event queue growing = metadata stalling
- ZK mode: also quorum + session expirations
- KRaft: __cluster_metadata Raft quorum + metadata lag
basics
~20 sAlert if the active controller count isn't exactly 1 (0 = no controller, >1 = split brain), and on rising controller queue size or slow leader elections. In KRaft, also watch the metadata quorum: leader presence, follower lag, and unfetched metadata.
solid answer
~50 sThe controller manages cluster metadata — leader elections, partition reassignments, ISR changes. Core alert: ActiveControllerCount across the cluster must sum to exactly 1. Zero means no controller (no elections happen, failures don't get healed); more than one indicates a split-brain/transition bug. Also alert on a growing controller event-queue (event processing falling behind), high ControllerEventQueueTimeMs, and slow/failed leader elections (LeaderElectionRateAndTimeMs spiking). In a ZooKeeper deployment the controller's state lives in ZK, so you also monitor ZK health: ensemble has a quorum, session expirations, and request latency — ZK problems cascade into controller instability. In KRaft (KIP-500), the controller quorum is a Raft group of controller nodes storing metadata in the __cluster_metadata log; you alert on the Raft quorum instead: an active leader exists, follower/observer lag (MetadataLag / lastApplied offset), and that brokers are caught up to the metadata log. KRaft removes ZK as a separate failure domain but adds the metadata-log quorum as the thing to watch.
go deeper
Know the controller coordinates the cluster and that exactly one active controller should exist.
Add the ActiveControllerCount alert and the idea that controller queue growth signals trouble.
Distinguish ZK quorum/session monitoring from KRaft metadata-quorum/lag monitoring and choose alerts accordingly.
Own the migration story (ZK → KRaft), redefine runbooks and alert sets for the metadata quorum, and teach the failure-domain shift.
## What the controller does Exactly one broker (ZooKeeper mode) or one controller node (KRaft mode) acts as the **active controller**. It owns cluster-wide metadata operations: electing partition leaders, shrinking/expanding ISRs, processing broker join/leave, and driving partition reassignments. If the controller is unhealthy, the cluster stops self-healing: a failed broker's partitions don't get new leaders (they go offline), and ISR changes stall. ## The universal alert: ActiveControllerCount JMX: `kafka.controller:type=KafkaController,name=ActiveControllerCount`, reported by each broker as 0 or 1. **Summed across the cluster it must equal exactly 1.** - **Sum = 0**: no active controller. Elections aren't happening; the cluster is leaderless at the metadata level. Critical page. - **Sum > 1**: two brokers think they're controller (split brain or a stuck failover). Dangerous — page. This is one of the highest-signal Kafka alerts because it gates the whole self-healing machinery. ## Other controller signals - **Controller event queue size / time**: `ControllerEventQueueSize` and `ControllerEventQueueTimeMs`. A growing backlog means metadata changes are processed slowly — leader elections and ISR updates lag, prolonging outages. - **Leader election rate/latency**: `LeaderElectionRateAndTimeMs`. Frequent or slow elections indicate churn (flapping brokers) or a struggling controller. - **Unclean leader elections**: `UncleanLeaderElectionsPerSec > 0` means an out-of-sync replica was promoted — possible data loss; alert immediately if `unclean.leader.election.enable` is ever true. ## ZooKeeper mode The controller persists state in **ZooKeeper**, a separate quorum service. ZK problems directly destabilize the controller, so you also alert on: - **Quorum/ensemble health**: a majority of ZK nodes must be up; loss of quorum freezes all metadata writes. - **Session expirations**: a broker's ZK session expiring can trigger spurious controller failover and leader elections. - **ZK request latency / outstanding requests**: high latency slows every controller operation. ZK is a distinct failure domain — many historical Kafka incidents were really ZK incidents. ## KRaft mode (KIP-500) KRaft removes ZooKeeper. A small set of **controller nodes** form a **Raft quorum** and store all metadata in an internal log topic, `__cluster_metadata`. One controller is the Raft **leader**; brokers are **observers** that replay the metadata log. New things to alert on: - **Quorum leader present**: the Raft group must have an elected leader; no leader = no metadata progress. - **Metadata log lag**: each broker/controller's `lastApplied`/`MetadataLag` — how far behind the metadata log it is. A broker far behind has a stale view of leadership and ISRs. - **Quorum follower health**: enough voters caught up to maintain majority; a follower that falls behind can't be counted toward commits. `kafka-metadata-quorum.sh --describe` exposes leader, voters, observers, and their offsets/lag. ## Why the distinction matters operationally In ZK mode you have two systems to keep healthy and two sets of runbooks. KRaft collapses them: fewer moving parts, faster failover and recovery, and no ZK split-brain class of bug — but your monitoring must now track the metadata quorum itself. Mis-applying ZK-era runbooks (e.g. 'restart ZooKeeper') to a KRaft cluster is a common on-call mistake. ## Edge cases - During a controlled controller failover, ActiveControllerCount can momentarily read 0 across brokers; require a short sustained window before paging. - A network partition can transiently produce sum > 1 readings as failover settles; correlate with election metrics before declaring split brain. - In KRaft, a broker can be healthy for client traffic but metadata-lagging — leaders it reports may be stale; the lag metric catches this.
- Why must ActiveControllerCount summed across the cluster equal exactly 1?Each broker reports 0 or 1. Sum 0 means no active controller, so no leader elections or ISR healing happen — the cluster can't recover from failures. Sum >1 means split brain: two brokers acting as controller, which corrupts metadata coordination. Only exactly one is correct.
- In KRaft, ZooKeeper is gone — what replaces the ZK health checks in your alerting?You monitor the KRaft metadata quorum: that the Raft group has an elected leader, that voters/followers are caught up enough to hold majority, and each node's metadata-log lag (lastApplied/MetadataLag). kafka-metadata-quorum.sh --describe surfaces leader, voters, observers, and offsets.
saying these in an interview costs you the question
- Saying you want more than one active controller for redundancy
- Believing KRaft still needs ZooKeeper health monitoring
- Ignoring controller event-queue growth as a leading indicator
- Treating any momentary ActiveControllerCount=0 during failover as an instant page without a sustained window