skip to content

KRaft Operations and Observability

Inspecting a KRaft quorum in production: quorum describe output, the metadata shell, and the controller metrics worth alerting on. Interviewers ask what you would actually look at when metadata seems stuck.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What does kafka-metadata-quorum.sh describe show, and how do you use it to check the health of a KRaft cluster?

level: juniorimportance: must knowfreq 70%

answer

  1. describe --status vs --replication
  2. voters vote, observers (brokers) don't
  3. LeaderId, LeaderEpoch, HighWatermark
  4. Lag + LastCaughtUpTimestamp per node
  5. majority of voters = progress

basics

~10 s

kafka-metadata-quorum.sh describe reports the KRaft metadata quorum: who the current leader controller is, the list of voters and observers, and how far each replica lags behind the leader's metadata log.

solid answer

~40 s

kafka-metadata-quorum.sh is the CLI for inspecting the Raft (KRaft) metadata quorum. You point it at a bootstrap server and run `describe`. With `--status` you get a summary: LeaderId, LeaderEpoch, HighWatermark, and current voter/observer IDs. With `--replication` you get a per-replica table showing NodeId, LogEndOffset, Lag, LastFetchTimestamp, LastCaughtUpTimestamp, and whether each is a Leader/Follower/Observer. Voters are the controllers that vote in Raft elections (form the quorum); observers (e.g. brokers) replicate the metadata log but don't vote. To assess health you check that there is exactly one leader, that all voters have low Lag and recent LastCaughtUpTimestamp, and that you have a majority of voters online. A growing Lag or stale timestamp on a voter signals a falling-behind or partitioned controller.

code

bash · 8 lines
bash
# One-shot quorum summary (leader, epoch, high watermark, voters/observers)
kafka-metadata-quorum.sh --bootstrap-server localhost:9092 describe --status

# Per-replica replication table (LogEndOffset, Lag, LastCaughtUpTimestamp, status)
kafka-metadata-quorum.sh --bootstrap-server localhost:9092 describe --replication

# When the cluster has lost quorum and brokers can't route, hit controllers directly
kafka-metadata-quorum.sh --bootstrap-controller localhost:9093 describe --status

go deeper

for a junior

Know that the command exists and that it tells you the controller leader, the voters/observers, and lag.

for a middle

Distinguish --status vs --replication output fields and interpret Lag plus timestamps to judge health.

for a senior

Use --bootstrap-controller for diagnosis during lost-quorum, reason about majority math and election churn from LeaderEpoch.

for a principal

Define quorum-health SLOs (max voter lag, epoch stability) and bake describe into alerting/runbooks alongside JMX metrics.

## Background: KRaft and the metadata quorum Apache Kafka historically stored cluster metadata (topics, partitions, configs, ACLs) in **ZooKeeper**. **KRaft** (Kafka Raft) replaces ZooKeeper with an internal **Raft consensus** mechanism. A special internal topic, `__cluster_metadata` (a single-partition replicated log), holds all metadata as an ordered sequence of records. A small set of **controller** nodes form a **quorum** that replicates this log using Raft. - **Voters**: the controllers configured in `controller.quorum.voters` (or, in newer KRaft, the dynamic voter set). They participate in leader elections and must reach a **majority** for the cluster to make progress. With 3 voters you tolerate 1 failure; with 5 voters, 2. - **Leader**: exactly one voter is elected leader for a given **epoch** (a monotonically increasing term number). The leader appends new metadata records; followers replicate them. - **Observers**: nodes that replicate the metadata log but do NOT vote — typically the **brokers**, which need an up-to-date metadata cache but must not affect quorum math. ## The tool: kafka-metadata-quorum.sh This CLI (backed by the `MetadataQuorumCommand` class) inspects that quorum. You connect via `--bootstrap-server` (or `--bootstrap-controller` to talk directly to the controllers). ### `describe --status` Returns a one-shot summary: - **ClusterId** - **LeaderId** — node ID of the current Raft leader - **LeaderEpoch** — current Raft term - **HighWatermark** — the offset in the metadata log known to be committed (replicated to a majority) - **MaxFollowerLag / MaxFollowerLagTimeMs** — worst-case follower lag - **CurrentVoters** and **CurrentObservers** — the node ID sets ### `describe --replication` Returns a per-node table: | Column | Meaning | |---|---| | NodeId | replica's node ID | | LogEndOffset | last offset that replica has in its metadata log | | Lag | LeaderEndOffset − this replica's LogEndOffset | | LastFetchTimestamp | when this replica last fetched from the leader | | LastCaughtUpTimestamp | when this replica was last fully caught up | | Status | Leader / Follower / Observer | ## Reading health from the output 1. **Exactly one Leader** with a stable LeaderEpoch — frequent epoch bumps mean election churn (flapping). 2. **Low Lag on all voters** — a voter with growing Lag is falling behind; if too many lag, you risk losing majority on a leader failure. 3. **Recent LastCaughtUpTimestamp** — a stale value means that replica hasn't kept up recently even if Lag looks small at one instant. 4. **Majority of voters present** — if voters are missing from CurrentVoters, the quorum is degraded. ## Edge cases - Querying via `--bootstrap-server` works through brokers, but if the cluster can't elect a controller (lost quorum), use `--bootstrap-controller` to reach controllers directly for diagnosis. - Lag is a snapshot; combine it with the timestamps to distinguish a momentary fetch gap from a persistently-behind replica. - Observers (brokers) having lag is usually less critical than a voter lagging, since observers don't affect quorum, but a badly-lagging broker has a stale metadata view.

  • What is the difference between a voter and an observer in the output?
    Voters are controllers in the quorum who vote in Raft elections and count toward majority; observers (typically brokers) replicate the metadata log read-only and do not vote. Losing observers doesn't break consensus; losing a majority of voters does.
  • Why look at LastCaughtUpTimestamp in addition to Lag?
    Lag is an instantaneous offset difference and can momentarily read 0 right after a fetch. LastCaughtUpTimestamp shows whether the replica has actually been keeping up over time, catching a replica that periodically falls behind.

saying these in an interview costs you the question

  • Saying the tool talks to ZooKeeper — KRaft has no ZooKeeper.
  • Claiming brokers are voters — brokers are observers by default.
  • Thinking Lag=0 at one instant proves a healthy replica; you must also check timestamps.
  • Confusing HighWatermark (committed offset) with LogEndOffset (last appended offset).

context

open as a page

What does the ActiveControllerCount metric mean in a KRaft cluster, and what value should it have?

level: middleimportance: must knowfreq 60%

basics

~10 s

ActiveControllerCount is a JMX gauge that is 1 on the controller node that is the current active (leader) controller and 0 on the others. Summed across the cluster it should always equal exactly 1.

open as a page

Which controller metrics track how far the metadata log has been applied, and how do you use lastAppliedRecordOffset to detect a lagging or stuck controller?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Controllers expose lastAppliedRecordOffset (the highest metadata-log offset the node has applied to its in-memory state), plus lastAppliedRecordTimestamp and lastAppliedRecordLagMs. Compare a follower's applied offset/lag against the leader to spot a controller that is stuck or falling behind.

open as a page

Design a monitoring and alerting strategy for KRaft quorum health. Which signals do you collect and what do you alert on?

level: principalimportance: should knowfreq 30%

basics

~20 s

Collect controller JMX metrics (ActiveControllerCount, lastAppliedRecordOffset/LagMs, metadata error counts, election/epoch counters) plus periodic kafka-metadata-quorum.sh describe output. Alert on: cluster ActiveControllerCount sum != 1, any voter lagging or missing, rising applied lag, and non-zero metadata errors.

open as a page

What is kafka-metadata-shell.sh used for, and when would you reach for it instead of kafka-metadata-quorum.sh?

level: middleimportance: nice to knowfreq 25%

basics

~10 s

kafka-metadata-shell.sh opens an interactive, filesystem-like shell over the contents of the KRaft __cluster_metadata log (or a snapshot). You browse the actual metadata records (topics, brokers, configs); kafka-metadata-quorum.sh instead reports quorum/replication health.

open as a page