What does LeaderElectionRateAndTimeMs measure, and how would you design alerting around broker health JMX metrics as a whole?
answer
- timer: rate (how often) + histogram (how long, ms)
- near zero healthy; spikes on failure/restart
- unclean elections = separate data-loss alert
- tier: critical/high/warning
- dwell windows + correct aggregation + correlation dashboards
basics
~20 sLeaderElectionRateAndTimeMs is a timer the controller exposes: it tracks how often partition leader elections happen and how long they take (in ms). Frequent or slow elections signal instability. For overall broker health, alert on a small curated set — offline/under-replicated partitions, controller count, thread idle, ISR churn, election rate — with thresholds and dwell times tuned to severity.
solid answer
~40 sLeaderElectionRateAndTimeMs (kafka.controller:type=ControllerStats,name=LeaderElectionRateAndTimeMs) is a controller-side timer: the rate (Count/OneMinuteRate) tells you how frequently leader elections occur and the histogram (Mean, 99thPercentile) tells you how long they take. A baseline near zero is healthy; spikes accompany broker failures or restarts, and a sustained high rate plus long election times signals churn that adds unavailability windows. For an overall broker-health alerting design, I'd tier by severity: critical pages on OfflinePartitionsCount > 0, sum(ActiveControllerCount) != 1, and UnderMinIsrPartitionCount > 0; high-priority on UnderReplicatedPartitions > 0 sustained and elevated IsrShrinksPerSec; warning on RequestHandlerAvgIdlePercent / NetworkProcessorAvgIdlePercent below ~0.3 and rising RequestQueueTimeMs; and election-rate spikes as a correlating signal. Each needs a dwell window to suppress transient restart noise, aggregation across brokers (sums vs maxes per metric), and dashboards that correlate symptom metrics with resource metrics for root cause.
go deeper
Knows leader elections happen when a leader fails and this metric tracks them.
Reads it as rate + duration, knows spikes accompany failures/restarts.
Correlates election rate/time with broker failures and distinguishes clean vs unclean elections.
Architects the tiered alerting strategy (thresholds, dwell, aggregation, correlation, KRaft) across the org's clusters.
## LeaderElectionRateAndTimeMs A **leader election** picks a new leader for a partition when the current leader becomes unavailable (broker down, controlled shutdown, or reassignment). The controller runs these elections. `kafka.controller:type=ControllerStats,name=LeaderElectionRateAndTimeMs` is a **timer** — a combined rate + histogram: - **Rate side:** `Count`, `OneMinuteRate` — *how often* elections happen. - **Histogram side:** `Mean`, `99thPercentile`, `Max` — *how long* each election takes, in milliseconds. Healthy clusters sit near **zero elections** in steady state. Elections are expected during planned restarts and reassignments. Concerning patterns: - **High election rate sustained:** leadership is churning — usually because brokers are flapping or the cluster is unstable. Every election is a brief unavailability window for the affected partitions (clients re-discover the new leader). - **Long election times (high 99th percentile):** the controller is slow to elect, extending unavailability — points at controller overload, large metadata, or (in ZK mode) ZooKeeper latency. - A related metric, `UncleanLeaderElectionsPerSec`, specifically counts elections that promoted an **out-of-sync** replica (potential data loss) and deserves its own alert when unclean election is enabled. ## Designing broker-health alerting as a whole The goal: page on real availability/durability problems, warn on saturation trends, and suppress transient restart noise. A pragmatic tiering: **Critical (page immediately):** - `OfflinePartitionsCount > 0` — partitions unavailable. - `sum(ActiveControllerCount) != 1` — no controller or split-brain. - `UnderMinIsrPartitionCount > 0` — acks=all writes failing. **High (page if sustained):** - `UnderReplicatedPartitions > 0` for > N minutes (e.g. 5-10) — durability margin shrinking. - Elevated `IsrShrinksPerSec` OneMinuteRate sustained — replica instability. - `UncleanLeaderElectionsPerSec > 0` if unclean election is enabled. **Warning (ticket / investigate):** - `RequestHandlerAvgIdlePercent` or `NetworkProcessorAvgIdlePercent` below ~0.3 sustained — saturation. - Rising `RequestQueueTimeMs` / `TotalTimeMs` percentiles — latency degradation. - Elevated `LeaderElectionRateAndTimeMs` — leadership churn (often a correlating, not primary, alert). **Cross-cutting design rules:** 1. **Aggregation:** know whether to sum or max each metric across brokers. UnderReplicatedPartitions and OfflinePartitionsCount sum (per-leader/controller reporting); ActiveControllerCount must be summed and compared to 1; idle ratios are per-broker so alert per broker. 2. **Dwell/duration windows:** require the condition to persist (e.g. `for: 5m` in Prometheus) so rolling restarts and reassignments don't page. 3. **Severity matched to impact:** durability-loss and availability metrics page; saturation metrics warn. 4. **Correlation dashboards:** pair each symptom metric with resource metrics (CPU, disk util, network, GC pause time, page-cache) so on-call can root-cause quickly — e.g. ISR flapping next to that broker's disk and GC graphs. 5. **Pipeline:** collect via the Prometheus JMX exporter agent, alert in Prometheus/Alertmanager, visualize in Grafana; avoid open JMX RMI. 6. **KRaft awareness:** in KRaft, also track metadata-log lag and quorum health, but the broker-health intent carries over. ## Edge cases - Alerting purely on a single broker's gauge (e.g. ActiveControllerCount) misreads the cluster — aggregation is essential. - Too-tight dwell windows cause restart-driven false pages; too-loose windows delay real incidents — tune per metric. - Election-rate spikes are often **secondary** to a broker failure already caught by other alerts; treat them as confirmation/context rather than the primary page to avoid alert storms.
- Why use a duration/dwell window (e.g. Prometheus 'for: 5m') on UnderReplicatedPartitions alerts?Rolling restarts, reassignments, and brief load bursts cause transient under-replication that self-heals. A dwell window requires the condition to persist before paging, filtering out expected, self-resolving spikes while still catching real sustained problems.
- How do you decide whether to sum or max a JMX gauge across brokers when alerting?It depends on how the metric is reported. Per-leader metrics like UnderReplicatedPartitions and controller-reported OfflinePartitionsCount are summed for a cluster total. ActiveControllerCount is summed and compared to exactly 1. Per-broker saturation gauges like the idle percentages are evaluated per broker (max/min), since each broker's value is independently meaningful.
- What's the difference between LeaderElectionRateAndTimeMs and UncleanLeaderElectionsPerSec, and why alert separately?LeaderElectionRateAndTimeMs counts all leader elections and their duration — most are clean and safe. UncleanLeaderElectionsPerSec counts only elections that promoted an out-of-sync replica, which can lose committed data. The unclean metric is a durability/data-loss alert and warrants its own, stricter alert when unclean election is enabled.
saying these in an interview costs you the question
- Thinking all leader elections are dangerous (clean elections are routine and safe)
- Confusing total election rate with unclean elections (only the latter risks data loss)
- Designing alerts on single-broker values without aggregation rules
- Omitting dwell windows so every rolling restart pages on-call
- Treating election-rate spikes as a primary page rather than correlating context