skip to content

What does LeaderElectionRateAndTimeMs measure, and how would you design alerting around broker health JMX metrics as a whole?

level: principalimportance: should knowfreq 30%

answer

  1. timer: rate (how often) + histogram (how long, ms)
  2. near zero healthy; spikes on failure/restart
  3. unclean elections = separate data-loss alert
  4. tier: critical/high/warning
  5. dwell windows + correct aggregation + correlation dashboards

basics

~20 s

LeaderElectionRateAndTimeMs is a timer the controller exposes: it tracks how often partition leader elections happen and how long they take (in ms). Frequent or slow elections signal instability. For overall broker health, alert on a small curated set — offline/under-replicated partitions, controller count, thread idle, ISR churn, election rate — with thresholds and dwell times tuned to severity.

solid answer

~40 s

LeaderElectionRateAndTimeMs (kafka.controller:type=ControllerStats,name=LeaderElectionRateAndTimeMs) is a controller-side timer: the rate (Count/OneMinuteRate) tells you how frequently leader elections occur and the histogram (Mean, 99thPercentile) tells you how long they take. A baseline near zero is healthy; spikes accompany broker failures or restarts, and a sustained high rate plus long election times signals churn that adds unavailability windows. For an overall broker-health alerting design, I'd tier by severity: critical pages on OfflinePartitionsCount > 0, sum(ActiveControllerCount) != 1, and UnderMinIsrPartitionCount > 0; high-priority on UnderReplicatedPartitions > 0 sustained and elevated IsrShrinksPerSec; warning on RequestHandlerAvgIdlePercent / NetworkProcessorAvgIdlePercent below ~0.3 and rising RequestQueueTimeMs; and election-rate spikes as a correlating signal. Each needs a dwell window to suppress transient restart noise, aggregation across brokers (sums vs maxes per metric), and dashboards that correlate symptom metrics with resource metrics for root cause.

go deeper

for a junior

Knows leader elections happen when a leader fails and this metric tracks them.

for a middle

Reads it as rate + duration, knows spikes accompany failures/restarts.

for a senior

Correlates election rate/time with broker failures and distinguishes clean vs unclean elections.

for a principal

Architects the tiered alerting strategy (thresholds, dwell, aggregation, correlation, KRaft) across the org's clusters.

## LeaderElectionRateAndTimeMs A **leader election** picks a new leader for a partition when the current leader becomes unavailable (broker down, controlled shutdown, or reassignment). The controller runs these elections. `kafka.controller:type=ControllerStats,name=LeaderElectionRateAndTimeMs` is a **timer** — a combined rate + histogram: - **Rate side:** `Count`, `OneMinuteRate` — *how often* elections happen. - **Histogram side:** `Mean`, `99thPercentile`, `Max` — *how long* each election takes, in milliseconds. Healthy clusters sit near **zero elections** in steady state. Elections are expected during planned restarts and reassignments. Concerning patterns: - **High election rate sustained:** leadership is churning — usually because brokers are flapping or the cluster is unstable. Every election is a brief unavailability window for the affected partitions (clients re-discover the new leader). - **Long election times (high 99th percentile):** the controller is slow to elect, extending unavailability — points at controller overload, large metadata, or (in ZK mode) ZooKeeper latency. - A related metric, `UncleanLeaderElectionsPerSec`, specifically counts elections that promoted an **out-of-sync** replica (potential data loss) and deserves its own alert when unclean election is enabled. ## Designing broker-health alerting as a whole The goal: page on real availability/durability problems, warn on saturation trends, and suppress transient restart noise. A pragmatic tiering: **Critical (page immediately):** - `OfflinePartitionsCount > 0` — partitions unavailable. - `sum(ActiveControllerCount) != 1` — no controller or split-brain. - `UnderMinIsrPartitionCount > 0` — acks=all writes failing. **High (page if sustained):** - `UnderReplicatedPartitions > 0` for > N minutes (e.g. 5-10) — durability margin shrinking. - Elevated `IsrShrinksPerSec` OneMinuteRate sustained — replica instability. - `UncleanLeaderElectionsPerSec > 0` if unclean election is enabled. **Warning (ticket / investigate):** - `RequestHandlerAvgIdlePercent` or `NetworkProcessorAvgIdlePercent` below ~0.3 sustained — saturation. - Rising `RequestQueueTimeMs` / `TotalTimeMs` percentiles — latency degradation. - Elevated `LeaderElectionRateAndTimeMs` — leadership churn (often a correlating, not primary, alert). **Cross-cutting design rules:** 1. **Aggregation:** know whether to sum or max each metric across brokers. UnderReplicatedPartitions and OfflinePartitionsCount sum (per-leader/controller reporting); ActiveControllerCount must be summed and compared to 1; idle ratios are per-broker so alert per broker. 2. **Dwell/duration windows:** require the condition to persist (e.g. `for: 5m` in Prometheus) so rolling restarts and reassignments don't page. 3. **Severity matched to impact:** durability-loss and availability metrics page; saturation metrics warn. 4. **Correlation dashboards:** pair each symptom metric with resource metrics (CPU, disk util, network, GC pause time, page-cache) so on-call can root-cause quickly — e.g. ISR flapping next to that broker's disk and GC graphs. 5. **Pipeline:** collect via the Prometheus JMX exporter agent, alert in Prometheus/Alertmanager, visualize in Grafana; avoid open JMX RMI. 6. **KRaft awareness:** in KRaft, also track metadata-log lag and quorum health, but the broker-health intent carries over. ## Edge cases - Alerting purely on a single broker's gauge (e.g. ActiveControllerCount) misreads the cluster — aggregation is essential. - Too-tight dwell windows cause restart-driven false pages; too-loose windows delay real incidents — tune per metric. - Election-rate spikes are often **secondary** to a broker failure already caught by other alerts; treat them as confirmation/context rather than the primary page to avoid alert storms.

  • Why use a duration/dwell window (e.g. Prometheus 'for: 5m') on UnderReplicatedPartitions alerts?
    Rolling restarts, reassignments, and brief load bursts cause transient under-replication that self-heals. A dwell window requires the condition to persist before paging, filtering out expected, self-resolving spikes while still catching real sustained problems.
  • How do you decide whether to sum or max a JMX gauge across brokers when alerting?
    It depends on how the metric is reported. Per-leader metrics like UnderReplicatedPartitions and controller-reported OfflinePartitionsCount are summed for a cluster total. ActiveControllerCount is summed and compared to exactly 1. Per-broker saturation gauges like the idle percentages are evaluated per broker (max/min), since each broker's value is independently meaningful.
  • What's the difference between LeaderElectionRateAndTimeMs and UncleanLeaderElectionsPerSec, and why alert separately?
    LeaderElectionRateAndTimeMs counts all leader elections and their duration — most are clean and safe. UncleanLeaderElectionsPerSec counts only elections that promoted an out-of-sync replica, which can lose committed data. The unclean metric is a durability/data-loss alert and warrants its own, stricter alert when unclean election is enabled.

saying these in an interview costs you the question

  • Thinking all leader elections are dangerous (clean elections are routine and safe)
  • Confusing total election rate with unclean elections (only the latter risks data loss)
  • Designing alerts on single-broker values without aggregation rules
  • Omitting dwell windows so every rolling restart pages on-call
  • Treating election-rate spikes as a primary page rather than correlating context

context