What does the ActiveControllerCount metric mean in a KRaft cluster, and what value should it have?
answer
- kafka.controller:...ActiveControllerCount
- 1 on leader, 0 elsewhere
- cluster sum must == 1
- sum 0 = no controller, writes stall
- alert on != 1
basics
~10 sActiveControllerCount is a JMX gauge that is 1 on the controller node that is the current active (leader) controller and 0 on the others. Summed across the cluster it should always equal exactly 1.
solid answer
~40 sActiveControllerCount is exposed via JMX under `kafka.controller:type=KafkaController,name=ActiveControllerCount`. In KRaft, each controller node reports 1 if it is currently the active/leader controller and 0 otherwise. Summed across all controller nodes the value must be exactly 1: a healthy cluster has one and only one active controller. A cluster-wide sum of 0 means there is no active controller — the quorum can't elect a leader (lost majority, network partition, or all controllers down), and metadata writes stall. A sum of 2 (transiently possible during a split-brain-like window or misreading per-node values) is a serious anomaly. Operationally you alert on `sum(ActiveControllerCount) != 1`. It pairs naturally with kafka-metadata-quorum.sh describe, which shows the same leader from the Raft layer's perspective.
go deeper
Remember: exactly one node has ActiveControllerCount=1; cluster sum should be 1.
Know the JMX object name and that sum=0 means no controller and stalled metadata writes.
Set the alert with a short 'for' window to tolerate normal elections; correlate with quorum describe and voter majority.
Fold it into control-plane SLOs and capacity planning (voter count vs fault tolerance), and design runbooks for the sum=0 recovery path.
## What it measures `ActiveControllerCount` is a JMX **gauge** published by every controller node: ``` kafka.controller:type=KafkaController,name=ActiveControllerCount ``` - On the node that is currently the **active controller** (the Raft leader of the `__cluster_metadata` log that is also serving as the cluster's controller), the gauge reads **1**. - On every other controller node it reads **0**. ## The invariant: cluster-wide sum == 1 Because Kafka must have exactly one authority appending metadata records, the **sum across all controller nodes** must be exactly **1** at all times in a healthy cluster. | Sum | Meaning | |---|---| | **1** | Healthy: one active controller. | | **0** | No active controller — quorum cannot elect a leader. Metadata is read-only/stalled; topic creation, config changes, partition reassignment, broker registration all block. Causes: majority of voters down, network partition isolating the would-be leader, all controllers down. | | **2+** | Anomalous. Should not happen in steady state; if observed it usually reflects a measurement/aggregation artifact, an in-flight election window being sampled, or a serious bug. Investigate immediately. | ## Why 0 is the dangerous case Raft requires a **majority of voters** to elect a leader. With 3 controllers you need 2 online; with 5 you need 3. If you drop below majority, no leader can be elected, ActiveControllerCount sums to 0, and the control plane freezes. Importantly, **already-produced data on brokers keeps flowing** for a while (brokers use cached metadata), but anything requiring metadata changes is blocked, and the situation degrades. ## How to use it 1. **Alerting**: alert when `sum(ActiveControllerCount) != 1` over a short window (allow a few seconds for normal elections). 2. **Correlate** with `kafka-metadata-quorum.sh describe --status` — the LeaderId there should be the same node reporting 1. 3. **Per-node dashboards**: graph each controller separately so you can see leadership move during a controlled failover (one node drops to 0, another rises to 1). ## Edge cases - Brief elections (e.g. during a rolling restart of controllers) can momentarily show 0; that's why alerts use a short delay/`for` window rather than firing instantly. - This metric is about the **controller leadership**, not partition leadership of regular topics (that's a different concern). Don't confuse it with `OfflinePartitionsCount` or partition leader metrics. - In a combined-mode node (process.roles=broker,controller) the same process can host both the broker and an active controller; the metric still reflects controller leadership only.
- Your monitoring shows sum(ActiveControllerCount)=0 for 30 seconds. What's happening and what still works?No controller could be elected — likely the voter quorum lost majority (controllers down or partitioned). Metadata changes (topic create, config, reassignment, broker registration) stall, but existing producers/consumers can keep working off cached metadata for a while. Restore voter majority to recover.
- How does this metric relate to kafka-metadata-quorum.sh describe?Both report controller leadership. The node reporting ActiveControllerCount=1 should match LeaderId from `describe --status`. JMX gives continuous time-series for alerting; the CLI gives an on-demand authoritative snapshot including voter lag.
saying these in an interview costs you the question
- Saying it counts the number of controller nodes — it's a leadership flag, not a node count.
- Thinking a healthy cluster can have a sum of 2 or more.
- Confusing it with partition-leader or OfflinePartitionsCount metrics.
- Believing sum=0 immediately stops all produce/consume — existing traffic survives on cached metadata for a time.