skip to content

What does the UnderReplicatedPartitions broker metric mean, and what should it normally read?

level: juniorimportance: must knowfreq 70%

answer

  1. |ISR| < replication factor
  2. kafka.server:type=ReplicaManager
  3. healthy = 0
  4. sum across brokers (per-leader)
  5. transient OK on rolling restart, sustained = page

basics

~20 s

UnderReplicatedPartitions counts how many partitions led by this broker have fewer in-sync replicas than the configured replication factor. In a healthy cluster it should be 0. A sustained non-zero value means replicas are falling behind or down.

solid answer

~40 s

Exposed at kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions, it's a per-broker gauge of how many partitions for which this broker is leader have |ISR| < replication factor — i.e. at least one replica is not caught up. Steady state should be 0. A sustained positive value signals a follower broker that is down, slow (disk/network/GC), or repeatedly dropping out of the ISR. Because each leader reports only its own partitions, you sum the gauge across brokers or watch the cluster max. It's one of the highest-priority broker alerts: under-replication erodes durability headroom and, if replicas keep dropping, can lead to offline partitions or unclean-leader risk. Briefly non-zero during a rolling restart or reassignment is expected; persistently non-zero is not.

go deeper

for a junior

Knows it should be 0 and that non-zero means replicas are behind or a broker is down.

for a middle

Understands ISR, replica.lag.time.max.ms, and why to sum across brokers.

for a senior

Distinguishes it from UnderMinIsr/OfflinePartitions and correlates with IsrShrinks and GC/disk causes.

for a principal

Sets cluster-wide durability SLOs, alert thresholds and dwell times, and ties it to min.insync.replicas and unclean-leader policy.

## Background terms - **Replica:** a copy of a partition's log on a broker. Each partition has one **leader** and N-1 **followers**, where N is the **replication factor**. - **ISR (In-Sync Replica set):** the subset of replicas that are caught up to the leader within `replica.lag.time.max.ms` (default 30s). Followers that fall behind are removed from the ISR. - A partition is **under-replicated** when `|ISR| < replication factor` — at least one replica that *should* be in sync is not. ## The metric `kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions` is a **gauge** reporting the number of partitions *led by this broker* that are currently under-replicated. Each broker only reports the partitions for which it is leader, so the cluster-wide picture is the **sum across all brokers** (or you alert on any broker's gauge > 0). ## What healthy looks like Steady state is **0**. Transient non-zero values are normal during: - rolling broker restarts (followers briefly catch up), - partition reassignments / cluster expansion, - a leader election storm. ## What a sustained non-zero value means - A **follower broker is down** — its replicas can't stay in sync. - A follower is **slow**: disk saturation, network saturation, long GC pauses, or an overloaded fetcher thread, causing it to fall outside `replica.lag.time.max.ms`. - Misconfiguration (e.g. replication factor higher than the number of live brokers). ## Why it's high-priority Under-replication shrinks your durability margin. If `min.insync.replicas` is set and acks=all producers can no longer satisfy it, **writes start failing**. If replicas keep dropping until only the leader is left and that leader then fails, you face an offline partition or (if unclean leader election is enabled) potential data loss. So `UnderReplicatedPartitions > 0` for more than a few minutes is a classic page-worthy alert. ## Related signals to look at together - `UnderMinIsrPartitionCount` — partitions below `min.insync.replicas` (more severe; acks=all writes already failing). - `IsrShrinksPerSec` / `IsrExpandsPerSec` — replicas flapping in and out of the ISR. - `OfflinePartitionsCount` — the worst case, no leader at all. ## Edge cases - A broker that is leader for **nothing** reports 0 even if the cluster is unhealthy — always aggregate across brokers. - During a clean shutdown the controller moves leadership first, so under-replication should stay low; a spike during shutdown hints at a problem.

  • How is UnderReplicatedPartitions different from UnderMinIsrPartitionCount?
    UnderReplicated counts partitions with |ISR| below the replication factor (durability margin shrinking). UnderMinIsr counts partitions with |ISR| below min.insync.replicas — at that point acks=all producers are already being rejected, so it's more urgent.
  • Why must you aggregate this gauge across brokers?
    Each broker reports only the under-replicated partitions for which it is the leader. A single broker's value undercounts the cluster, and a broker leading no partitions reports 0 regardless of cluster health.

saying these in an interview costs you the question

  • Saying it counts all replicas in the cluster rather than partitions led by this broker
  • Claiming any non-zero value is always an emergency (transient spikes on restart are normal)
  • Confusing it with consumer lag
  • Thinking a single broker's gauge represents the whole cluster

context