skip to content

Walk through the on-call runbook you'd follow when a single broker is down and URP has spiked across many partitions. What do you check, fix, and escalate?

level: seniorimportance: should knowfreq 50%

answer

  1. Confirm → scope → triage broker → assess durability → remediate → verify
  2. One broker down = URP across many partitions
  3. Check UnderMinIsrPartitionCount (writes failing?)
  4. Recoverable: restart, watch URP drop
  5. Never unclean.leader.election to clear URP

basics

~20 s

Confirm the alert and identify the down broker. Check if it's process-down vs disk-full vs network-isolated. If recoverable, restart/heal it and watch URP drop as replicas rejoin ISR. If data is lost, reassign replicas. Escalate if min.insync.replicas is breached or partitions go offline.

solid answer

~60 s

Step 1 — confirm and scope: verify the URP alert is real (not a maintenance window), and from the controller metrics see which broker dropped out and how many partitions are affected; one broker typically explains a fleet of URP. Step 2 — triage the broker: is the process down, the disk full, JVM in a GC death-spiral, or is it network-isolated (broker up but unreachable)? Check broker logs, disk usage, and ZK/metadata connectivity. Step 3 — assess durability: is `UnderMinIsrPartitionCount > 0`? If ISR for any partition is below `min.insync.replicas`, `acks=all` writes are already failing — this is now an availability incident, escalate. Step 4 — remediate: if the broker is recoverable, restart/fix it; replicas refetch and rejoin ISR, and URP should decline steadily. If the disk/data is lost, the broker rebuilds replicas from leaders (watch network/IO), or you reassign partitions off it with `kafka-reassign-partitions.sh`. Step 5 — verify and escalate: watch URP trend to zero. Never enable `unclean.leader.election` to clear URP — it risks data loss; escalate to the data-platform owner instead. Capture timeline for the postmortem.

go deeper

for a junior

Identify the down broker and restart it, then watch URP recover; know not to flip unclean leader election.

for a middle

Triage the failure mode (process/disk/GC/network) and check whether min.insync.replicas is breached.

for a senior

Run the full triage→durability→remediate→verify flow, manage throttled rebuilds and reassignment, and decide when to escalate.

for a principal

Author the runbook and escalation policy, encode rack-aware placement and durability SLOs so single-broker loss never reaches offline, and lead the postmortem.

## Setup: what the alert means **Under-replicated partitions (URP)** = partitions whose in-sync replica set (ISR) is smaller than the replication factor. When **one broker** dies, every partition that had a replica on it loses an ISR member, so a single broker failure typically spikes URP across *many* partitions at once. The data is still served (leaders are elsewhere), but redundancy is gone and a second failure could cause loss or offline partitions. ## Runbook **1. Confirm the alert is genuine.** - Is a planned rolling restart or maintenance window in progress? If so, expect transient URP and inhibit the page. - Is URP sustained (the alert should already require this) or a momentary blip? **2. Scope the blast radius.** - From `OfflinePartitionsCount`: are any partitions actually offline (no leader)? If yes, this is already an outage — jump to escalation. - Identify the failed broker: URP across many partitions usually points to one broker dropping out of all their ISRs. Per-broker metrics or `kafka-topics.sh --describe --under-replicated-partitions` show which replicas are missing. **3. Triage the broker's failure mode.** - **Process down**: crashed/OOM — check broker logs and the JVM (OutOfMemoryError, segfault). - **Disk full / failed disk**: `log.dirs` partition at 100% — Kafka marks the log dir offline; the broker may be partially up. - **GC death spiral**: long GC pauses make the broker miss `replica.lag.time.max.ms` and `zookeeper.session.timeout.ms`/metadata heartbeats, so it falls out of ISR while technically 'up'. - **Network isolation**: broker is alive but unreachable by peers/controller — followers can't fetch. **4. Assess durability urgency.** - Check `UnderMinIsrPartitionCount`: if any partition's ISR is below `min.insync.replicas`, producers with `acks=all` are already getting `NotEnoughReplicasException` — writes are failing. This elevates the incident from 'durability degraded' to 'partial write outage'. Escalate. - Note replication factor: with RF=3 and one broker down, ISR is 2; still above a typical `min.insync.replicas=2`, so writes continue. With RF=2 you may already be at the edge. **5. Remediate.** - **Recoverable broker**: restart the process / free the disk / fix the network. On rejoin, followers fetch from leaders and re-enter ISR; URP should trend steadily to zero. Watch it — if it doesn't fall, the replica may be stuck fetching (slow disk, throttle). - **Lost broker/data**: bring up a replacement broker with the same broker id (it rebuilds replicas from leaders) — watch replication network/IO and consider `replication.throttle` to avoid saturating leaders. Or permanently reassign partitions off the dead broker with `kafka-reassign-partitions.sh`. - **What NOT to do**: do not set `unclean.leader.election.enable=true` to 'clear' URP/offline — it promotes an out-of-sync replica and can silently drop committed data. That is an escalation decision, not a quick fix. **6. Verify and hand off.** - Confirm URP → 0 and `UnderMinIsrPartitionCount → 0`. - If you can't recover quickly, or durability is breached, escalate to the Kafka/data-platform on-call and (if customer-facing) declare an incident. - Record the timeline, root cause, and metric graphs for the postmortem. ## Edge cases - **Rack awareness**: if replica placement isn't rack-aware, a single rack/AZ outage can take multiple replicas of the same partition, turning one 'broker down' into offline partitions. - **Throttled rebuild**: an existing `follower.replication.throttled.rate` can make ISR re-entry crawl — URP stays high not because the broker is broken but because catch-up is rate-limited. - **Cascading**: a slow broker can drag down leaders it fetches from; fix the slow broker before assuming the leaders are the problem.

  • Mid-incident someone suggests setting unclean.leader.election.enable=true to clear the URP fast. Your response?
    Refuse it as a reflex. Unclean leader election promotes an out-of-sync replica as leader, discarding messages the old leader had but the new one didn't — silent data loss. It's a deliberate, escalated availability-over-durability tradeoff, not an on-call shortcut, and only relevant if partitions are actually offline.
  • The broker is restarted but URP isn't falling. What might be wrong?
    The rejoining follower may be stuck catching up: a replication throttle (follower.replication.throttled.rate) rate-limiting fetch, a slow/failing disk, or network saturation. Check fetch progress per replica and whether a throttle config is capping catch-up before assuming a deeper fault.

saying these in an interview costs you the question

  • Jumping to unclean leader election to clear URP
  • Assuming each URP partition is an independent failure rather than one common broker
  • Not checking UnderMinIsrPartitionCount before declaring 'just degraded'
  • Restarting blindly without checking disk-full or GC root cause (it'll recur)

context