Design a monitoring and alerting strategy for KRaft quorum health. Which signals do you collect and what do you alert on?
answer
- JMX time-series + CLI describe snapshots
- page: sum(ActiveControllerCount)!=1
- page: voters at bare majority
- warn: voter lag / stale caught-up / applied lag
- warn: metadata errors + epoch churn
- outage signal vs erosion signal
basics
~20 sCollect controller JMX metrics (ActiveControllerCount, lastAppliedRecordOffset/LagMs, metadata error counts, election/epoch counters) plus periodic kafka-metadata-quorum.sh describe output. Alert on: cluster ActiveControllerCount sum != 1, any voter lagging or missing, rising applied lag, and non-zero metadata errors.
solid answer
~50 sA robust KRaft quorum-monitoring strategy layers continuous JMX time-series with periodic CLI snapshots. From JMX I collect, per controller: ActiveControllerCount (sum must equal 1), lastAppliedRecordOffset/lastAppliedRecordLagMs (apply progress and staleness), metadata loading/apply error counts, and election-related counters/LeaderEpoch (to catch flapping). From kafka-metadata-quorum.sh describe (scraped on a schedule) I capture the voter set, per-voter Lag and LastCaughtUpTimestamp, and HighWatermark. Core alerts: (1) page on sum(ActiveControllerCount) != 1 sustained beyond a normal election window; (2) warn when any voter's Lag or applied lag exceeds a threshold or LastCaughtUpTimestamp goes stale — you're approaching loss of fault tolerance; (3) page when the online voter count drops to the bare majority (one more failure = outage); (4) warn on any non-zero/growing metadata error count; (5) warn on frequent LeaderEpoch increments (election churn). I also watch metadata-log dir disk and snapshot age, and ensure controllers are dashboarded separately so failover is observable.
go deeper
Know the headline alert: cluster ActiveControllerCount sum should be 1.
List the core metrics (ActiveControllerCount, applied lag, voter lag) and what each warns about.
Combine JMX and CLI, set thresholds, and separate outage alerts from fault-tolerance-erosion alerts.
Architect the full strategy: severity tiers, election-tolerant windows, fault-tolerance/capacity planning, snapshot/disk/clock edge cases, and observer impact on clients.
## Goal The KRaft quorum is the **control plane**: if it degrades, you can't create topics, change configs, reassign partitions, or register brokers, and eventually data-plane availability suffers. The strategy must detect both **outright loss** (no leader) and **erosion of fault tolerance** (voters falling behind so the next failure tips you over majority). ## Signals to collect ### A. Continuous JMX (per controller node) - **`ActiveControllerCount`** — 1 on leader, 0 elsewhere. Aggregate: cluster sum. - **`lastAppliedRecordOffset` / `lastAppliedRecordTimestamp` / `lastAppliedRecordLagMs`** — apply progress and staleness of each node's metadata image (controllers and brokers). - **Metadata error counts** — controller metadata loading/apply error metrics; non-zero means the node hit a problem applying records. - **LeaderEpoch / election counters** — to detect election churn (frequent re-elections / epoch bumps). - **Metadata-log dir disk usage and I/O latency** — a full or slow disk stalls appends/applies. - **Snapshot age / size** — stale or missing snapshots make recovery slow and let the log grow. ### B. Periodic CLI scrape (cluster-wide truth) Run `kafka-metadata-quorum.sh describe --status` and `--replication` on a schedule and export: - LeaderId, LeaderEpoch, HighWatermark. - Current voter and observer sets (detect a voter that dropped out). - Per-voter **Lag**, **LastFetchTimestamp**, **LastCaughtUpTimestamp**. The CLI is authoritative for voter membership/lag in a way JMX per-node metrics alone don't summarize; JMX gives the high-resolution time-series for alerting. ## What to alert on | Severity | Condition | Why | |---|---|---| | **Page** | `sum(ActiveControllerCount) != 1` sustained (e.g. > N seconds) | No (or ambiguous) controller — control plane down. Short window tolerates normal elections. | | **Page** | Online voter count at bare majority (e.g. 2 of 3) | One more failure causes total quorum loss. | | **Page** | Quorum describe unreachable / no LeaderId | Can't confirm a leader; likely partition or outage. | | **Warn** | Any voter Lag > threshold OR applied lag rising OR stale LastCaughtUpTimestamp | Fault tolerance eroding; a lagging voter may not be a useful failover target. | | **Warn** | Non-zero / growing metadata error count | Node may be diverging or unable to apply metadata. | | **Warn** | Frequent LeaderEpoch increments (election churn) | Instability — networking, GC, or overloaded controllers. | | **Warn** | Metadata-log disk near full / snapshot age too old | Risk of stalled appends and slow recovery. | ## Design principles - **Distinguish outage from erosion.** Sum!=1 is the outage signal; voter lag / bare-majority are the erosion signals you want to catch *before* an outage. - **Tolerate normal elections.** Use `for`/duration windows so rolling restarts don't page. - **Per-node dashboards.** Graph each controller separately so a controlled failover (one→0, another→1) is visible and expected. - **Combine layers.** JMX for resolution + alerting; CLI for authoritative voter membership/lag; metadata shell reserved for deep/post-mortem inspection. - **Capacity/fault-tolerance planning.** 3 voters tolerate 1 failure, 5 tolerate 2; pick voter count for your failure-domain spread, and alert relative to that majority. ## Edge cases - **Clock skew** distorts time-based lag (`lastAppliedRecordLagMs`, LastCaughtUpTimestamp) — keep NTP healthy and treat wild values as suspect. - **Brokers as observers**: their applied lag affects client routing correctness even though they don't vote — monitor it too. - **Combined-mode nodes** (broker+controller) share a process; resource exhaustion can hit both roles at once. - **Snapshot loads** cause brief offset jumps — don't alert on those transients.
- Why alert when online voters drop to a bare majority (2 of 3) even though the cluster is still 'up'?At bare majority there's zero remaining fault tolerance — one more voter failure loses quorum and freezes the control plane. Alerting here gives you time to restore a voter before an outage rather than reacting to one.
- Why combine kafka-metadata-quorum.sh scrapes with JMX rather than relying on JMX alone?JMX gives high-resolution per-node time-series ideal for alerting, but the CLI is authoritative for voter/observer membership and per-voter lag/caught-up timestamps in a single cluster-level view — catching a voter that dropped out of the set, which is awkward to infer from per-node JMX alone.
- How do you avoid paging during routine rolling restarts of controllers?Use a sustained 'for' window on the sum(ActiveControllerCount)!=1 alert so brief election gaps don't fire, and optionally suppress/silence quorum alerts during planned maintenance windows while still watching for a failover that doesn't complete.
saying these in an interview costs you the question
- Only alerting on the outage (sum!=1) and ignoring erosion signals like voter lag and bare-majority.
- Firing the sum!=1 alert instantly with no tolerance for normal elections.
- Relying solely on JMX and missing a voter dropping out of the membership set.
- Ignoring broker (observer) applied lag, which still affects client routing correctness.
- Trusting time-based lag metrics without verifying clock sync.