skip to content

Broker JMX Metrics and Health Signals

The handful of broker MBeans that actually matter: under-replicated partitions, offline partitions, active controller count, and handler idle percentages. Interviewers ask which few metrics you would put on a dashboard first.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What does the UnderReplicatedPartitions broker metric mean, and what should it normally read?

level: juniorimportance: must knowfreq 70%

answer

  1. |ISR| < replication factor
  2. kafka.server:type=ReplicaManager
  3. healthy = 0
  4. sum across brokers (per-leader)
  5. transient OK on rolling restart, sustained = page

basics

~20 s

UnderReplicatedPartitions counts how many partitions led by this broker have fewer in-sync replicas than the configured replication factor. In a healthy cluster it should be 0. A sustained non-zero value means replicas are falling behind or down.

solid answer

~40 s

Exposed at kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions, it's a per-broker gauge of how many partitions for which this broker is leader have |ISR| < replication factor — i.e. at least one replica is not caught up. Steady state should be 0. A sustained positive value signals a follower broker that is down, slow (disk/network/GC), or repeatedly dropping out of the ISR. Because each leader reports only its own partitions, you sum the gauge across brokers or watch the cluster max. It's one of the highest-priority broker alerts: under-replication erodes durability headroom and, if replicas keep dropping, can lead to offline partitions or unclean-leader risk. Briefly non-zero during a rolling restart or reassignment is expected; persistently non-zero is not.

go deeper

for a junior

Knows it should be 0 and that non-zero means replicas are behind or a broker is down.

for a middle

Understands ISR, replica.lag.time.max.ms, and why to sum across brokers.

for a senior

Distinguishes it from UnderMinIsr/OfflinePartitions and correlates with IsrShrinks and GC/disk causes.

for a principal

Sets cluster-wide durability SLOs, alert thresholds and dwell times, and ties it to min.insync.replicas and unclean-leader policy.

## Background terms - **Replica:** a copy of a partition's log on a broker. Each partition has one **leader** and N-1 **followers**, where N is the **replication factor**. - **ISR (In-Sync Replica set):** the subset of replicas that are caught up to the leader within `replica.lag.time.max.ms` (default 30s). Followers that fall behind are removed from the ISR. - A partition is **under-replicated** when `|ISR| < replication factor` — at least one replica that *should* be in sync is not. ## The metric `kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions` is a **gauge** reporting the number of partitions *led by this broker* that are currently under-replicated. Each broker only reports the partitions for which it is leader, so the cluster-wide picture is the **sum across all brokers** (or you alert on any broker's gauge > 0). ## What healthy looks like Steady state is **0**. Transient non-zero values are normal during: - rolling broker restarts (followers briefly catch up), - partition reassignments / cluster expansion, - a leader election storm. ## What a sustained non-zero value means - A **follower broker is down** — its replicas can't stay in sync. - A follower is **slow**: disk saturation, network saturation, long GC pauses, or an overloaded fetcher thread, causing it to fall outside `replica.lag.time.max.ms`. - Misconfiguration (e.g. replication factor higher than the number of live brokers). ## Why it's high-priority Under-replication shrinks your durability margin. If `min.insync.replicas` is set and acks=all producers can no longer satisfy it, **writes start failing**. If replicas keep dropping until only the leader is left and that leader then fails, you face an offline partition or (if unclean leader election is enabled) potential data loss. So `UnderReplicatedPartitions > 0` for more than a few minutes is a classic page-worthy alert. ## Related signals to look at together - `UnderMinIsrPartitionCount` — partitions below `min.insync.replicas` (more severe; acks=all writes already failing). - `IsrShrinksPerSec` / `IsrExpandsPerSec` — replicas flapping in and out of the ISR. - `OfflinePartitionsCount` — the worst case, no leader at all. ## Edge cases - A broker that is leader for **nothing** reports 0 even if the cluster is unhealthy — always aggregate across brokers. - During a clean shutdown the controller moves leadership first, so under-replication should stay low; a spike during shutdown hints at a problem.

  • How is UnderReplicatedPartitions different from UnderMinIsrPartitionCount?
    UnderReplicated counts partitions with |ISR| below the replication factor (durability margin shrinking). UnderMinIsr counts partitions with |ISR| below min.insync.replicas — at that point acks=all producers are already being rejected, so it's more urgent.
  • Why must you aggregate this gauge across brokers?
    Each broker reports only the under-replicated partitions for which it is the leader. A single broker's value undercounts the cluster, and a broker leading no partitions reports 0 regardless of cluster health.

saying these in an interview costs you the question

  • Saying it counts all replicas in the cluster rather than partitions led by this broker
  • Claiming any non-zero value is always an emergency (transient spikes on restart are normal)
  • Confusing it with consumer lag
  • Thinking a single broker's gauge represents the whole cluster

context

open as a page

What is JMX in the context of a Kafka broker, and how do you expose and read broker metrics through it?

level: juniorimportance: must knowfreq 55%

basics

~20 s

JMX (Java Management Extensions) is the Java standard for exposing runtime metrics as MBeans. A Kafka broker publishes its internal metrics as JMX MBeans. You enable it by setting JMX_PORT (or jmxremote system properties) and read it with tools like JConsole, jmxterm, or a Prometheus JMX exporter.

open as a page

What do ActiveControllerCount and OfflinePartitionsCount tell you, and what is healthy for each?

level: middleimportance: must knowfreq 60%

basics

~20 s

ActiveControllerCount is per-broker and is 1 on the single controller broker and 0 on the rest, so cluster-wide it must sum to exactly 1. OfflinePartitionsCount is the number of partitions with no leader — healthy is 0; any positive value means those partitions are unavailable.

open as a page

What do RequestHandlerAvgIdlePercent and NetworkProcessorAvgIdlePercent measure, and how do you interpret low values?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Both are saturation gauges from 0.0 to 1.0. RequestHandlerAvgIdlePercent is the fraction of time the I/O (request handler) threads are idle; NetworkProcessorAvgIdlePercent is the fraction of time the network threads are idle. Low values (near 0) mean those thread pools are saturated and the broker is a bottleneck.

open as a page

What do IsrShrinksPerSec and IsrExpandsPerSec measure, and what does a high or flapping rate tell you?

level: seniorimportance: should knowfreq 40%

basics

~20 s

They are rate meters counting how often replicas leave the in-sync replica set (shrink) or rejoin it (expand) per second. Occasional events are normal; a high or oscillating rate means replicas are repeatedly falling behind and catching up — a sign of an overloaded or unstable broker.

open as a page

What does LeaderElectionRateAndTimeMs measure, and how would you design alerting around broker health JMX metrics as a whole?

level: principalimportance: should knowfreq 30%

basics

~20 s

LeaderElectionRateAndTimeMs is a timer the controller exposes: it tracks how often partition leader elections happen and how long they take (in ms). Frequent or slow elections signal instability. For overall broker health, alert on a small curated set — offline/under-replicated partitions, controller count, thread idle, ISR churn, election rate — with thresholds and dwell times tuned to severity.

open as a page