skip to content

What do ActiveControllerCount and OfflinePartitionsCount tell you, and what is healthy for each?

level: middleimportance: must knowfreq 60%

answer

  1. sum(ActiveControllerCount) must == 1
  2. 0 = no brain, 2+ = split-brain
  3. OfflinePartitions = no leader = unavailable
  4. both = availability alerts
  5. KRaft: controller quorum, same intent

basics

~20 s

ActiveControllerCount is per-broker and is 1 on the single controller broker and 0 on the rest, so cluster-wide it must sum to exactly 1. OfflinePartitionsCount is the number of partitions with no leader — healthy is 0; any positive value means those partitions are unavailable.

solid answer

~40 s

ActiveControllerCount (kafka.controller:type=KafkaController,name=ActiveControllerCount) is a per-broker gauge that reads 1 on whichever broker is the elected controller and 0 everywhere else. Summed across the cluster it must equal exactly 1: a sum of 0 means no controller (no leadership management happening) and a sum of 2+ means a split-brain. OfflinePartitionsCount (kafka.controller:type=KafkaController,name=OfflinePartitionsCount) is reported by the controller and counts partitions that currently have no leader — usually because all in-sync replicas are down. Healthy is 0; any positive value means producers and consumers for those partitions get errors until a leader is restored. Both are top-tier availability alerts. In KRaft mode the controller quorum changes the wording but you still alert that exactly one active controller exists and that offline partitions stay at 0.

go deeper

for a junior

Knows one controller exists and offline partitions = unavailable, healthy is 0.

for a middle

Can explain the per-broker sum rule, split-brain, and how all-ISR-down + no unclean election causes offline partitions.

for a senior

Builds correct aggregate alerts, distinguishes offline vs under-replicated, and weighs unclean-leader trade-offs.

for a principal

Defines availability SLOs and runbooks for controller failover/split-brain across ZK and KRaft topologies.

## ActiveControllerCount **The controller** is the single broker responsible for cluster-management duties: electing partition leaders, tracking ISR changes, and propagating metadata. In a healthy cluster exactly **one** broker is the controller. `kafka.controller:type=KafkaController,name=ActiveControllerCount` is a **per-broker gauge**: it reads **1** on the broker that is currently the active controller and **0** on all others. The monitoring rule is therefore about the **sum across the cluster**: - **Sum = 1** -> healthy. - **Sum = 0** -> *no* active controller. Leader elections and ISR updates stall; the cluster can't react to broker failures. This often happens during a controller failover, ZK/quorum connectivity loss, or a controller crash. - **Sum >= 2** -> **split-brain**: two brokers each think they're the controller, which can corrupt metadata. This is a severe alert. ## OfflinePartitionsCount `kafka.controller:type=KafkaController,name=OfflinePartitionsCount` is reported by the **controller** and counts partitions that have **no leader** at all. A partition goes offline when every replica that could be leader is unavailable — typically all replicas in the ISR are down, and unclean leader election is disabled (the default), so no out-of-sync replica is promoted. - **Healthy = 0.** - **> 0** means those partitions are completely **unavailable**: producers get `NotLeaderOrFollowerException`/timeouts and consumers can't fetch. This is a direct customer-facing outage for the affected partitions. ## Why they pair These are the two classic *availability* (as opposed to durability) broker alerts. UnderReplicatedPartitions is a warning that durability margin is shrinking; OfflinePartitionsCount is the realized outage. ActiveControllerCount tells you whether the brain that *fixes* leadership is even functioning. ## KRaft note In **KRaft** mode (ZooKeeper-free), control plane responsibility moves to a **controller quorum** of nodes running a Raft metadata log. There is still an active controller (the quorum leader), and you still alert that exactly one active controller exists and OfflinePartitionsCount is 0. Additional KRaft metrics track the metadata log (e.g. metadata lag), but the broker-health intent is unchanged. ## Edge cases - Because ActiveControllerCount is per-broker, alerting on a single broker's value is wrong — you must aggregate. A common alert is `sum(ActiveControllerCount) != 1`. - A momentary sum=0 during a deliberate controller move is normal; sustained 0 is not. - OfflinePartitionsCount only meaningfully reports from the active controller; on non-controller brokers it's 0, so aggregate with max or read it from the controller.

  • Your alerting shows sum(ActiveControllerCount) = 0 for several minutes. What's happening and why does it matter?
    No broker is acting as controller, so leader elections, ISR updates, and failover handling are stalled. The cluster can't recover from broker failures during this window, so it's a critical alert — usually a controller failover that didn't complete or quorum/ZK connectivity loss.
  • How can OfflinePartitionsCount go positive even when brokers are running?
    If all in-sync replicas for a partition are unavailable and unclean leader election is disabled, the controller refuses to promote an out-of-sync replica, so the partition has no leader and is offline despite other brokers being up.

saying these in an interview costs you the question

  • Expecting ActiveControllerCount to be 1 on every broker (it's 1 only on the controller, 0 elsewhere)
  • Alerting on a single broker's ActiveControllerCount instead of the cluster sum
  • Thinking offline partitions are the same as under-replicated (offline = no leader = down; under-replicated = leader exists but fewer in-sync replicas)
  • Assuming enabling unclean leader election is a free fix — it risks data loss

context