skip to content

What controller-health signals should you alert on in a Kafka cluster, and how does this differ between ZooKeeper-based and KRaft clusters?

level: seniorimportance: should knowfreq 55%

answer

  1. ActiveControllerCount summed = exactly 1
  2. 0 = no controller, >1 = split brain
  3. Controller event queue growing = metadata stalling
  4. ZK mode: also quorum + session expirations
  5. KRaft: __cluster_metadata Raft quorum + metadata lag

basics

~20 s

Alert if the active controller count isn't exactly 1 (0 = no controller, >1 = split brain), and on rising controller queue size or slow leader elections. In KRaft, also watch the metadata quorum: leader presence, follower lag, and unfetched metadata.

solid answer

~50 s

The controller manages cluster metadata — leader elections, partition reassignments, ISR changes. Core alert: ActiveControllerCount across the cluster must sum to exactly 1. Zero means no controller (no elections happen, failures don't get healed); more than one indicates a split-brain/transition bug. Also alert on a growing controller event-queue (event processing falling behind), high ControllerEventQueueTimeMs, and slow/failed leader elections (LeaderElectionRateAndTimeMs spiking). In a ZooKeeper deployment the controller's state lives in ZK, so you also monitor ZK health: ensemble has a quorum, session expirations, and request latency — ZK problems cascade into controller instability. In KRaft (KIP-500), the controller quorum is a Raft group of controller nodes storing metadata in the __cluster_metadata log; you alert on the Raft quorum instead: an active leader exists, follower/observer lag (MetadataLag / lastApplied offset), and that brokers are caught up to the metadata log. KRaft removes ZK as a separate failure domain but adds the metadata-log quorum as the thing to watch.

go deeper

for a junior

Know the controller coordinates the cluster and that exactly one active controller should exist.

for a middle

Add the ActiveControllerCount alert and the idea that controller queue growth signals trouble.

for a senior

Distinguish ZK quorum/session monitoring from KRaft metadata-quorum/lag monitoring and choose alerts accordingly.

for a principal

Own the migration story (ZK → KRaft), redefine runbooks and alert sets for the metadata quorum, and teach the failure-domain shift.

## What the controller does Exactly one broker (ZooKeeper mode) or one controller node (KRaft mode) acts as the **active controller**. It owns cluster-wide metadata operations: electing partition leaders, shrinking/expanding ISRs, processing broker join/leave, and driving partition reassignments. If the controller is unhealthy, the cluster stops self-healing: a failed broker's partitions don't get new leaders (they go offline), and ISR changes stall. ## The universal alert: ActiveControllerCount JMX: `kafka.controller:type=KafkaController,name=ActiveControllerCount`, reported by each broker as 0 or 1. **Summed across the cluster it must equal exactly 1.** - **Sum = 0**: no active controller. Elections aren't happening; the cluster is leaderless at the metadata level. Critical page. - **Sum > 1**: two brokers think they're controller (split brain or a stuck failover). Dangerous — page. This is one of the highest-signal Kafka alerts because it gates the whole self-healing machinery. ## Other controller signals - **Controller event queue size / time**: `ControllerEventQueueSize` and `ControllerEventQueueTimeMs`. A growing backlog means metadata changes are processed slowly — leader elections and ISR updates lag, prolonging outages. - **Leader election rate/latency**: `LeaderElectionRateAndTimeMs`. Frequent or slow elections indicate churn (flapping brokers) or a struggling controller. - **Unclean leader elections**: `UncleanLeaderElectionsPerSec > 0` means an out-of-sync replica was promoted — possible data loss; alert immediately if `unclean.leader.election.enable` is ever true. ## ZooKeeper mode The controller persists state in **ZooKeeper**, a separate quorum service. ZK problems directly destabilize the controller, so you also alert on: - **Quorum/ensemble health**: a majority of ZK nodes must be up; loss of quorum freezes all metadata writes. - **Session expirations**: a broker's ZK session expiring can trigger spurious controller failover and leader elections. - **ZK request latency / outstanding requests**: high latency slows every controller operation. ZK is a distinct failure domain — many historical Kafka incidents were really ZK incidents. ## KRaft mode (KIP-500) KRaft removes ZooKeeper. A small set of **controller nodes** form a **Raft quorum** and store all metadata in an internal log topic, `__cluster_metadata`. One controller is the Raft **leader**; brokers are **observers** that replay the metadata log. New things to alert on: - **Quorum leader present**: the Raft group must have an elected leader; no leader = no metadata progress. - **Metadata log lag**: each broker/controller's `lastApplied`/`MetadataLag` — how far behind the metadata log it is. A broker far behind has a stale view of leadership and ISRs. - **Quorum follower health**: enough voters caught up to maintain majority; a follower that falls behind can't be counted toward commits. `kafka-metadata-quorum.sh --describe` exposes leader, voters, observers, and their offsets/lag. ## Why the distinction matters operationally In ZK mode you have two systems to keep healthy and two sets of runbooks. KRaft collapses them: fewer moving parts, faster failover and recovery, and no ZK split-brain class of bug — but your monitoring must now track the metadata quorum itself. Mis-applying ZK-era runbooks (e.g. 'restart ZooKeeper') to a KRaft cluster is a common on-call mistake. ## Edge cases - During a controlled controller failover, ActiveControllerCount can momentarily read 0 across brokers; require a short sustained window before paging. - A network partition can transiently produce sum > 1 readings as failover settles; correlate with election metrics before declaring split brain. - In KRaft, a broker can be healthy for client traffic but metadata-lagging — leaders it reports may be stale; the lag metric catches this.

  • Why must ActiveControllerCount summed across the cluster equal exactly 1?
    Each broker reports 0 or 1. Sum 0 means no active controller, so no leader elections or ISR healing happen — the cluster can't recover from failures. Sum >1 means split brain: two brokers acting as controller, which corrupts metadata coordination. Only exactly one is correct.
  • In KRaft, ZooKeeper is gone — what replaces the ZK health checks in your alerting?
    You monitor the KRaft metadata quorum: that the Raft group has an elected leader, that voters/followers are caught up enough to hold majority, and each node's metadata-log lag (lastApplied/MetadataLag). kafka-metadata-quorum.sh --describe surfaces leader, voters, observers, and offsets.

saying these in an interview costs you the question

  • Saying you want more than one active controller for redundancy
  • Believing KRaft still needs ZooKeeper health monitoring
  • Ignoring controller event-queue growth as a leading indicator
  • Treating any momentary ActiveControllerCount=0 during failover as an instant page without a sustained window

context