skip to content

Design a monitoring and alerting strategy for KRaft quorum health. Which signals do you collect and what do you alert on?

level: principalimportance: should knowfreq 30%

answer

  1. JMX time-series + CLI describe snapshots
  2. page: sum(ActiveControllerCount)!=1
  3. page: voters at bare majority
  4. warn: voter lag / stale caught-up / applied lag
  5. warn: metadata errors + epoch churn
  6. outage signal vs erosion signal

basics

~20 s

Collect controller JMX metrics (ActiveControllerCount, lastAppliedRecordOffset/LagMs, metadata error counts, election/epoch counters) plus periodic kafka-metadata-quorum.sh describe output. Alert on: cluster ActiveControllerCount sum != 1, any voter lagging or missing, rising applied lag, and non-zero metadata errors.

solid answer

~50 s

A robust KRaft quorum-monitoring strategy layers continuous JMX time-series with periodic CLI snapshots. From JMX I collect, per controller: ActiveControllerCount (sum must equal 1), lastAppliedRecordOffset/lastAppliedRecordLagMs (apply progress and staleness), metadata loading/apply error counts, and election-related counters/LeaderEpoch (to catch flapping). From kafka-metadata-quorum.sh describe (scraped on a schedule) I capture the voter set, per-voter Lag and LastCaughtUpTimestamp, and HighWatermark. Core alerts: (1) page on sum(ActiveControllerCount) != 1 sustained beyond a normal election window; (2) warn when any voter's Lag or applied lag exceeds a threshold or LastCaughtUpTimestamp goes stale — you're approaching loss of fault tolerance; (3) page when the online voter count drops to the bare majority (one more failure = outage); (4) warn on any non-zero/growing metadata error count; (5) warn on frequent LeaderEpoch increments (election churn). I also watch metadata-log dir disk and snapshot age, and ensure controllers are dashboarded separately so failover is observable.

go deeper

for a junior

Know the headline alert: cluster ActiveControllerCount sum should be 1.

for a middle

List the core metrics (ActiveControllerCount, applied lag, voter lag) and what each warns about.

for a senior

Combine JMX and CLI, set thresholds, and separate outage alerts from fault-tolerance-erosion alerts.

for a principal

Architect the full strategy: severity tiers, election-tolerant windows, fault-tolerance/capacity planning, snapshot/disk/clock edge cases, and observer impact on clients.

## Goal The KRaft quorum is the **control plane**: if it degrades, you can't create topics, change configs, reassign partitions, or register brokers, and eventually data-plane availability suffers. The strategy must detect both **outright loss** (no leader) and **erosion of fault tolerance** (voters falling behind so the next failure tips you over majority). ## Signals to collect ### A. Continuous JMX (per controller node) - **`ActiveControllerCount`** — 1 on leader, 0 elsewhere. Aggregate: cluster sum. - **`lastAppliedRecordOffset` / `lastAppliedRecordTimestamp` / `lastAppliedRecordLagMs`** — apply progress and staleness of each node's metadata image (controllers and brokers). - **Metadata error counts** — controller metadata loading/apply error metrics; non-zero means the node hit a problem applying records. - **LeaderEpoch / election counters** — to detect election churn (frequent re-elections / epoch bumps). - **Metadata-log dir disk usage and I/O latency** — a full or slow disk stalls appends/applies. - **Snapshot age / size** — stale or missing snapshots make recovery slow and let the log grow. ### B. Periodic CLI scrape (cluster-wide truth) Run `kafka-metadata-quorum.sh describe --status` and `--replication` on a schedule and export: - LeaderId, LeaderEpoch, HighWatermark. - Current voter and observer sets (detect a voter that dropped out). - Per-voter **Lag**, **LastFetchTimestamp**, **LastCaughtUpTimestamp**. The CLI is authoritative for voter membership/lag in a way JMX per-node metrics alone don't summarize; JMX gives the high-resolution time-series for alerting. ## What to alert on | Severity | Condition | Why | |---|---|---| | **Page** | `sum(ActiveControllerCount) != 1` sustained (e.g. > N seconds) | No (or ambiguous) controller — control plane down. Short window tolerates normal elections. | | **Page** | Online voter count at bare majority (e.g. 2 of 3) | One more failure causes total quorum loss. | | **Page** | Quorum describe unreachable / no LeaderId | Can't confirm a leader; likely partition or outage. | | **Warn** | Any voter Lag > threshold OR applied lag rising OR stale LastCaughtUpTimestamp | Fault tolerance eroding; a lagging voter may not be a useful failover target. | | **Warn** | Non-zero / growing metadata error count | Node may be diverging or unable to apply metadata. | | **Warn** | Frequent LeaderEpoch increments (election churn) | Instability — networking, GC, or overloaded controllers. | | **Warn** | Metadata-log disk near full / snapshot age too old | Risk of stalled appends and slow recovery. | ## Design principles - **Distinguish outage from erosion.** Sum!=1 is the outage signal; voter lag / bare-majority are the erosion signals you want to catch *before* an outage. - **Tolerate normal elections.** Use `for`/duration windows so rolling restarts don't page. - **Per-node dashboards.** Graph each controller separately so a controlled failover (one→0, another→1) is visible and expected. - **Combine layers.** JMX for resolution + alerting; CLI for authoritative voter membership/lag; metadata shell reserved for deep/post-mortem inspection. - **Capacity/fault-tolerance planning.** 3 voters tolerate 1 failure, 5 tolerate 2; pick voter count for your failure-domain spread, and alert relative to that majority. ## Edge cases - **Clock skew** distorts time-based lag (`lastAppliedRecordLagMs`, LastCaughtUpTimestamp) — keep NTP healthy and treat wild values as suspect. - **Brokers as observers**: their applied lag affects client routing correctness even though they don't vote — monitor it too. - **Combined-mode nodes** (broker+controller) share a process; resource exhaustion can hit both roles at once. - **Snapshot loads** cause brief offset jumps — don't alert on those transients.

  • Why alert when online voters drop to a bare majority (2 of 3) even though the cluster is still 'up'?
    At bare majority there's zero remaining fault tolerance — one more voter failure loses quorum and freezes the control plane. Alerting here gives you time to restore a voter before an outage rather than reacting to one.
  • Why combine kafka-metadata-quorum.sh scrapes with JMX rather than relying on JMX alone?
    JMX gives high-resolution per-node time-series ideal for alerting, but the CLI is authoritative for voter/observer membership and per-voter lag/caught-up timestamps in a single cluster-level view — catching a voter that dropped out of the set, which is awkward to infer from per-node JMX alone.
  • How do you avoid paging during routine rolling restarts of controllers?
    Use a sustained 'for' window on the sum(ActiveControllerCount)!=1 alert so brief election gaps don't fire, and optionally suppress/silence quorum alerts during planned maintenance windows while still watching for a failover that doesn't complete.

saying these in an interview costs you the question

  • Only alerting on the outage (sum!=1) and ignoring erosion signals like voter lag and bare-majority.
  • Firing the sum!=1 alert instantly with no tolerance for normal elections.
  • Relying solely on JMX and missing a voter dropping out of the membership set.
  • Ignoring broker (observer) applied lag, which still affects client routing correctness.
  • Trusting time-based lag metrics without verifying clock sync.

context