Which controller metrics track how far the metadata log has been applied, and how do you use lastAppliedRecordOffset to detect a lagging or stuck controller?
answer
- replicate (LogEndOffset) vs apply (lastApplied*)
- lastAppliedRecordOffset / Timestamp / LagMs
- plateaued offset = stuck applier
- low Raft Lag + high applied LagMs = apply bottleneck
- metadata error count non-zero = investigate
basics
~20 sControllers expose lastAppliedRecordOffset (the highest metadata-log offset the node has applied to its in-memory state), plus lastAppliedRecordTimestamp and lastAppliedRecordLagMs. Compare a follower's applied offset/lag against the leader to spot a controller that is stuck or falling behind.
solid answer
~50 sKRaft controllers (and brokers, as metadata observers) publish metrics under `kafka.server:type=...` / `kafka.controller` that describe how current each node's view of the metadata log is. The key ones are `lastAppliedRecordOffset` (the last `__cluster_metadata` offset this node has applied to its in-memory metadata image), `lastAppliedRecordTimestamp` (the wall-clock time embedded in that record), and `lastAppliedRecordLagMs` (how old the last applied record is relative to now). On the active controller, lastAppliedRecordOffset tracks near the log end offset; on followers and brokers it should trail only slightly. To detect trouble you watch for: a follower whose lastAppliedRecordOffset stops advancing while the leader's climbs (stuck applier or slow disk), or a rising lastAppliedRecordLagMs on a broker (its metadata view is going stale). Combine with kafka-metadata-quorum.sh describe (which shows LogEndOffset/Lag at the replication layer) — applied-offset metrics catch problems in the apply step even when fetch/replication looks fine.
go deeper
Just know there's a metric showing how far each node has applied the metadata log.
Know lastAppliedRecordOffset/Timestamp/LagMs and that rising lag means a stale metadata view.
Separate replication lag from apply lag, diagnose a stuck applier, and correlate with metadata error counts and quorum describe.
Design observability that distinguishes fetch vs apply bottlenecks, account for snapshots and clock skew, and set SLOs on broker metadata staleness for client correctness.
## Two layers: replicated vs applied KRaft metadata moves through two stages on every node: 1. **Replication** — the node fetches records of the `__cluster_metadata` log and appends them to its local log (tracked by LogEndOffset / Lag, visible in `kafka-metadata-quorum.sh describe --replication`). 2. **Apply** — the node feeds those records through its state machine to build the in-memory **metadata image** (topics, configs, ACLs). This is what the `lastApplied*` metrics measure. A node can be **replicating fine but applying slowly** (or stuck), so you need both views. ## The key metrics | Metric | Meaning | |---|---| | `lastAppliedRecordOffset` | Highest metadata-log offset the node has applied to its in-memory state. | | `lastAppliedRecordTimestamp` | The append timestamp recorded in that last-applied record. | | `lastAppliedRecordLagMs` | `now − lastAppliedRecordTimestamp`: how stale the applied view is. | These are exposed by both **controllers** (`kafka.controller` / `kafka.server` MBeans) and **brokers** (which apply metadata as observers). The exact MBean names have shifted across Kafka versions, but the semantics above are stable. ## Detecting a lagging or stuck controller 1. **Compare across nodes.** Graph `lastAppliedRecordOffset` per node. On a healthy cluster all controllers track close together and the leader leads slightly. A follower whose offset **plateaus** while others advance is stuck applying — possible causes: slow/failing disk, GC pauses, a poison record, or a thread hang. 2. **Watch `lastAppliedRecordLagMs`.** A steadily **rising** lag means the node's metadata image is aging. On a broker, a high applied lag means it may route based on a stale topic/partition map — clients can see NOT_LEADER errors or miss new topics. 3. **Cross-check with quorum describe.** If `describe --replication` shows low Lag (replication healthy) but `lastAppliedRecordLagMs` is high, the bottleneck is in the **apply** path, not the network/fetch path. Conversely high Raft Lag points at replication/fetch. 4. **Snapshot interactions.** KRaft periodically writes **snapshots** of the metadata log so it doesn't grow unbounded. A node catching up from a snapshot will briefly show a jump in applied offset; don't mistake the snapshot-load gap for being stuck. ## Metadata error counts Controllers also expose error/anomaly counters (e.g. metadata loading or apply error counts) and `MetadataErrorCount`-style metrics. A non-zero and growing error count indicates the controller hit problems applying records — often correlated with an applied offset that stops advancing. Any non-zero value here warrants investigation because a controller that can't apply metadata diverges from the rest of the cluster. ## Edge cases - **Active controller vs followers**: the active controller applies records it just committed, so its applied offset is near the high watermark; followers apply after fetching, so a small steady gap is normal. - **Brokers are observers**: their applied lag matters for data-plane correctness (routing) even though they don't vote. - **Clock skew** can distort `lastAppliedRecordLagMs` because it uses record timestamps vs local now; large negative or wildly large values may indicate NTP problems rather than real lag.
- Quorum describe shows Lag near 0 for a follower controller, but its lastAppliedRecordLagMs is climbing. What does that tell you?Replication/fetch is healthy (it has the records locally) but the apply step is the bottleneck — the state machine isn't processing records fast enough or is stuck. Look at disk I/O, GC pauses, thread dumps, and metadata error counts on that node, not the network.
- Why might a broker's lastAppliedRecordLagMs matter for client-facing correctness?Brokers serve metadata (topic/partition leadership) to clients from their applied image. A stale applied image means clients get outdated routing, causing NOT_LEADER_OR_FOLLOWER errors, missed new topics, or sends to old leaders until the broker catches up.
saying these in an interview costs you the question
- Treating lastAppliedRecordOffset and the Raft LogEndOffset as the same thing — they measure different stages (apply vs replicate).
- Assuming a controller that replicates fine must also be applying fine.
- Ignoring metadata error counts when the applied offset plateaus.
- Trusting lastAppliedRecordLagMs without checking clock sync (it uses record timestamps).