skip to content

Operational Monitoring and Alerting

What to page on for Kafka, from under-replicated and offline partitions to consumer lag and controller health, and the runbook that follows. Interviewers ask to see whether your alerts map to user impact rather than noise.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

Why are under-replicated partitions (URP) and offline partitions the two most important things to alert on in a Kafka cluster, and how should the alert severities differ?

level: juniorimportance: must knowfreq 78%

answer

  1. Offline = no leader = page now
  2. URP = ISR < RF = durability lost
  3. OfflinePartitionsCount (controller)
  4. UnderReplicatedPartitions (ReplicaManager)
  5. Sustained URP, not transient spike

basics

~10 s

Offline partitions mean some partitions have no leader, so reads/writes fail — page immediately (critical). Under-replicated partitions mean replicas are behind the leader, reducing durability but still serving traffic — warn, page if sustained.

solid answer

~40 s

Alert on these because they directly map to availability and durability. Offline partitions (broker metric kafka.controller:type=KafkaController,name=OfflinePartitionsCount > 0) mean partitions have no leader: producers and consumers for those partitions get errors. That is a hard outage — page immediately, critical. Under-replicated partitions (kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions > 0) mean the in-sync replica set (ISR) is smaller than the configured replication factor: data is still served by the leader, but you've lost redundancy, so a single further failure could cause data loss or offline partitions. A brief blip during a broker restart or rolling deploy is normal; alert when URP > 0 is sustained (e.g. 5-15 minutes) so you don't page on transient ISR shrink. Severity: offline = critical/page now; URP = warning, escalating to page if it persists or the count keeps climbing.

go deeper

for a junior

Know offline = no leader = outage = page; URP = replicas behind = durability risk = warn.

for a middle

Cite the exact JMX metrics and that URP alerts should require a sustained duration to avoid restart noise.

for a senior

Connect URP to min.insync.replicas and acks=all write failures; design severity tiers and maintenance-window inhibition.

for a principal

Define org-wide SLOs mapping these to availability/durability budgets, and teach the ISR mechanics so teams set thresholds deliberately.

## Core concepts A Kafka **topic** is split into **partitions**; each partition is replicated onto multiple brokers for fault tolerance. The **replication factor (RF)** is how many copies exist (commonly 3). For each partition one replica is the **leader** (handles all reads/writes) and the rest are **followers** that copy data from the leader. The **in-sync replica set (ISR)** is the subset of replicas that are caught up with the leader (within `replica.lag.time.max.ms`, default 30s). A replica drops out of ISR if it stops fetching or falls behind. ## Under-replicated partitions (URP) A partition is **under-replicated** when `|ISR| < RF` — i.e. fewer than the configured number of replicas are caught up. JMX metric: `kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions`. The leader still serves the partition, so producers/consumers keep working. The risk is **reduced durability**: you've lost redundancy. If `min.insync.replicas` is set (say 2) and ISR shrinks below it, producers using `acks=all` start getting `NotEnoughReplicasException` and writes fail — so severe URP can become a write outage even before partitions go offline. Causes: a broker is down/slow, network partition, disk saturation, GC pauses, or a rolling restart in progress. URP is expected transiently during restarts, so alert on **sustained** URP (5-15 min) rather than any nonzero spike. ## Offline partitions A partition is **offline** when none of its replicas can be leader — typically all replicas are down, or the last in-sync replica failed and `unclean.leader.election.enable=false` (the safe default) prevents an out-of-sync replica from becoming leader. JMX metric: `kafka.controller:type=KafkaController,name=OfflinePartitionsCount`, reported by the active controller. Offline = no leader = clients for that partition get errors and cannot produce or consume. **This is a live outage.** ## Why these two, and severities They are the canonical health signals because they map cleanly to the two things users care about: **availability** (offline = down now) and **durability** (URP = redundancy lost, outage risk rising). - **OfflinePartitionsCount > 0 → critical, page immediately.** Any nonzero value is a customer-facing outage. - **UnderReplicatedPartitions > 0 sustained → warning that escalates.** Page if it persists, the count grows, or ISR drops below `min.insync.replicas`. ## Edge cases - During a planned rolling restart, both metrics can briefly move; silence/inhibit alerts for the maintenance window. - A single slow broker can cause URP across many partitions at once — alert on the count and on *which broker* is the common factor. - `UnderMinIsrPartitionCount` is a sharper durability alert: it counts partitions whose ISR is below `min.insync.replicas`, i.e. writes with `acks=all` are already failing.

  • A rolling restart briefly shows URP > 0. How do you avoid paging on it?
    Use a sustained duration (e.g. 'for 10m') in the alert rule, and inhibit/silence URP and offline alerts during planned maintenance windows. Transient ISR shrink during a controlled restart is expected.
  • What's the difference between UnderReplicatedPartitions and UnderMinIsrPartitionCount?
    URP counts partitions where ISR < replication factor (durability degraded but writes may still succeed). UnderMinIsrPartitionCount counts partitions where ISR < min.insync.replicas, meaning acks=all producers are already failing — a sharper, more urgent durability/availability signal.

saying these in an interview costs you the question

  • Saying offline and under-replicated are the same thing
  • Paging on any single nonzero URP reading instead of sustained URP
  • Claiming URP means data is currently unavailable (the leader still serves it)
  • Forgetting that severe ISR shrink + min.insync.replicas can cause write failures

context

open as a page

How would you design a consumer-lag alert that pages on-call only for real problems, and what makes raw lag a poor threshold?

level: middleimportance: must knowfreq 80%

basics

~20 s

Raw lag (offsets behind) is misleading because acceptable lag depends on throughput. Alert on time-to-drain (lag / consume rate) or sustained-and-growing lag, not a single absolute number, and require the condition to persist before paging.

open as a page

What controller-health signals should you alert on in a Kafka cluster, and how does this differ between ZooKeeper-based and KRaft clusters?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Alert if the active controller count isn't exactly 1 (0 = no controller, >1 = split brain), and on rising controller queue size or slow leader elections. In KRaft, also watch the metadata quorum: leader presence, follower lag, and unfetched metadata.

open as a page

Walk through the on-call runbook you'd follow when a single broker is down and URP has spiked across many partitions. What do you check, fix, and escalate?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Confirm the alert and identify the down broker. Check if it's process-down vs disk-full vs network-isolated. If recoverable, restart/heal it and watch URP drop as replicas rejoin ISR. If data is lost, reassign replicas. Escalate if min.insync.replicas is breached or partitions go offline.

open as a page

How would you structure Kafka alerting around SLOs to avoid alert fatigue — symptom-based vs cause-based alerts, escalation tiers, and what should and shouldn't page?

level: principalimportance: should knowfreq 42%

basics

~20 s

Page on a small set of symptom alerts tied to user-facing SLOs (availability, durability, freshness) — offline partitions, write failures, SLO-breaching lag. Route cause-based and predictive signals (disk filling, URP, rising controller queue) to tickets/dashboards, not pages, unless they breach an SLO.

open as a page