Why are under-replicated partitions (URP) and offline partitions the two most important things to alert on in a Kafka cluster, and how should the alert severities differ?
answer
- Offline = no leader = page now
- URP = ISR < RF = durability lost
- OfflinePartitionsCount (controller)
- UnderReplicatedPartitions (ReplicaManager)
- Sustained URP, not transient spike
basics
~10 sOffline partitions mean some partitions have no leader, so reads/writes fail — page immediately (critical). Under-replicated partitions mean replicas are behind the leader, reducing durability but still serving traffic — warn, page if sustained.
solid answer
~40 sAlert on these because they directly map to availability and durability. Offline partitions (broker metric kafka.controller:type=KafkaController,name=OfflinePartitionsCount > 0) mean partitions have no leader: producers and consumers for those partitions get errors. That is a hard outage — page immediately, critical. Under-replicated partitions (kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions > 0) mean the in-sync replica set (ISR) is smaller than the configured replication factor: data is still served by the leader, but you've lost redundancy, so a single further failure could cause data loss or offline partitions. A brief blip during a broker restart or rolling deploy is normal; alert when URP > 0 is sustained (e.g. 5-15 minutes) so you don't page on transient ISR shrink. Severity: offline = critical/page now; URP = warning, escalating to page if it persists or the count keeps climbing.
go deeper
Know offline = no leader = outage = page; URP = replicas behind = durability risk = warn.
Cite the exact JMX metrics and that URP alerts should require a sustained duration to avoid restart noise.
Connect URP to min.insync.replicas and acks=all write failures; design severity tiers and maintenance-window inhibition.
Define org-wide SLOs mapping these to availability/durability budgets, and teach the ISR mechanics so teams set thresholds deliberately.
## Core concepts A Kafka **topic** is split into **partitions**; each partition is replicated onto multiple brokers for fault tolerance. The **replication factor (RF)** is how many copies exist (commonly 3). For each partition one replica is the **leader** (handles all reads/writes) and the rest are **followers** that copy data from the leader. The **in-sync replica set (ISR)** is the subset of replicas that are caught up with the leader (within `replica.lag.time.max.ms`, default 30s). A replica drops out of ISR if it stops fetching or falls behind. ## Under-replicated partitions (URP) A partition is **under-replicated** when `|ISR| < RF` — i.e. fewer than the configured number of replicas are caught up. JMX metric: `kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions`. The leader still serves the partition, so producers/consumers keep working. The risk is **reduced durability**: you've lost redundancy. If `min.insync.replicas` is set (say 2) and ISR shrinks below it, producers using `acks=all` start getting `NotEnoughReplicasException` and writes fail — so severe URP can become a write outage even before partitions go offline. Causes: a broker is down/slow, network partition, disk saturation, GC pauses, or a rolling restart in progress. URP is expected transiently during restarts, so alert on **sustained** URP (5-15 min) rather than any nonzero spike. ## Offline partitions A partition is **offline** when none of its replicas can be leader — typically all replicas are down, or the last in-sync replica failed and `unclean.leader.election.enable=false` (the safe default) prevents an out-of-sync replica from becoming leader. JMX metric: `kafka.controller:type=KafkaController,name=OfflinePartitionsCount`, reported by the active controller. Offline = no leader = clients for that partition get errors and cannot produce or consume. **This is a live outage.** ## Why these two, and severities They are the canonical health signals because they map cleanly to the two things users care about: **availability** (offline = down now) and **durability** (URP = redundancy lost, outage risk rising). - **OfflinePartitionsCount > 0 → critical, page immediately.** Any nonzero value is a customer-facing outage. - **UnderReplicatedPartitions > 0 sustained → warning that escalates.** Page if it persists, the count grows, or ISR drops below `min.insync.replicas`. ## Edge cases - During a planned rolling restart, both metrics can briefly move; silence/inhibit alerts for the maintenance window. - A single slow broker can cause URP across many partitions at once — alert on the count and on *which broker* is the common factor. - `UnderMinIsrPartitionCount` is a sharper durability alert: it counts partitions whose ISR is below `min.insync.replicas`, i.e. writes with `acks=all` are already failing.
- A rolling restart briefly shows URP > 0. How do you avoid paging on it?Use a sustained duration (e.g. 'for 10m') in the alert rule, and inhibit/silence URP and offline alerts during planned maintenance windows. Transient ISR shrink during a controlled restart is expected.
- What's the difference between UnderReplicatedPartitions and UnderMinIsrPartitionCount?URP counts partitions where ISR < replication factor (durability degraded but writes may still succeed). UnderMinIsrPartitionCount counts partitions where ISR < min.insync.replicas, meaning acks=all producers are already failing — a sharper, more urgent durability/availability signal.
saying these in an interview costs you the question
- Saying offline and under-replicated are the same thing
- Paging on any single nonzero URP reading instead of sustained URP
- Claiming URP means data is currently unavailable (the leader still serves it)
- Forgetting that severe ISR shrink + min.insync.replicas can cause write failures