skip to content

Consumer Lag and End-to-End Latency Monitoring

Tracking consumer lag and end-to-end latency from client metrics, committed offsets, and exporters. Interviewers ask because lag is the metric product owners actually feel.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What is consumer lag in Kafka, and how do you compute it for a single partition?

level: juniorimportance: must knowfreq 80%

answer

  1. lag = LEO - committed offset
  2. per partition, summed per group
  3. rising lag = consumer slower than producer
  4. committed offset in __consumer_offsets
  5. max-partition lag often matters most

basics

~20 s

Consumer lag is how far behind a consumer is on a partition: log-end offset minus the consumer's committed offset. A lag of 0 means the consumer has read everything; a growing lag means it can't keep up with producers.

solid answer

~40 s

Consumer lag measures how many messages a consumer group still has to read on a partition. For one partition it is LogEndOffset (LEO, the offset of the next message the broker will write) minus the consumer group's committed offset for that partition. So lag = LEO - committedOffset. A group's total lag is the sum across all assigned partitions; the per-partition maximum is often what matters most because one slow partition can hold up downstream processing. Lag near 0 means the group is caught up. A steadily rising lag means consumption throughput is below production throughput, while a flat-but-nonzero lag means the consumer keeps pace but stays a fixed distance behind. Lag is the single most important consumer health signal because it directly reflects data-processing delay, independent of CPU or memory.

go deeper

for a junior

Know the formula lag = LEO - committed offset and that rising lag means the consumer can't keep up.

for a middle

Aggregate lag per group, distinguish total vs max-partition lag, and reason about flat vs rising lag regimes.

for a senior

Connect lag to retention risk, idle-producer false positives, and explain why max-partition lag drives alerting.

for a principal

Frame lag as the primary SLO signal for streaming pipelines and design alerting that separates throughput deficits from idle producers and from latency-based SLOs.

## What lag is Kafka stores each topic partition as an append-only log. Every message gets a monotonically increasing **offset** (0, 1, 2, ...). Two offsets matter for lag: - **Log-end offset (LEO)**: the offset that will be assigned to the *next* message produced to that partition. Equivalently, it is one past the last message currently on the broker. (More precisely the **high-water mark** — the last *committed/replicated* offset readable by consumers — is what consumers can reach; in healthy clusters LEO and high-water mark are effectively the same for lag math.) - **Committed offset**: the offset a consumer group has acknowledged it has processed, stored in the internal `__consumer_offsets` topic. By convention the committed offset is the offset of the *next* record to read, i.e. last-processed + 1. **Lag for one partition = LogEndOffset - committedOffset.** Example: producers have written up to offset 1000 (so LEO = 1000), and the group committed offset 970. Lag = 1000 - 970 = 30 messages still to read. ## Why it matters Lag is the most direct measure of how stale a consumer's view of the data is. Unlike CPU or memory, lag speaks the language of the business: "we are 30 messages / 5 seconds behind real time." There are three regimes: - **Lag ≈ 0**: caught up. - **Lag flat but nonzero**: consumer matches producer throughput but trails by a constant amount (e.g. a fixed batching/processing delay). - **Lag rising over time**: consumption throughput < production throughput. This is the danger signal — without intervention (more consumers, faster processing) lag grows unbounded and may eventually breach `retention.ms`, causing data loss for that group. ## Aggregating across partitions A consumer group is assigned many partitions. Total group lag = sum of per-partition lag. But the **max** per-partition lag is frequently the more actionable number, because a single hot or stuck partition can stall end-to-end pipelines even when total lag looks modest. ## Edge cases - **No committed offset yet** (brand-new group): tools may show lag as the full LEO or as unknown, depending on `auto.offset.reset`. - **Idle producer**: if no new messages arrive, lag naturally drops to 0 and stays there — that is healthy, not a stalled consumer. - **Committed offset can briefly exceed nothing** but should never legitimately exceed LEO; a negative lag usually signals a tooling/race artifact (offset read before LEO refresh) rather than a real state. - Lag is **per consumer group**, not per topic: two groups reading the same topic have independent lag.

  • If lag is flat at a nonzero value, is the consumer unhealthy?
    Not necessarily. Flat lag means the consumer keeps pace with the producer but trails by a constant amount (e.g. a fixed processing/batching delay). Rising lag is the unhealthy signal; flat lag is steady-state.
  • Why might total group lag be misleading compared to max per-partition lag?
    Total lag can look small while one partition is badly stuck. Since processing and ordering are per-partition, a single high-lag partition can stall a pipeline even if the sum across partitions is modest, so max-partition lag is often the better alert signal.

saying these in an interview costs you the question

  • Saying lag is committed offset minus LEO (the subtraction is backwards).
  • Claiming lag is per-topic rather than per consumer group.
  • Thinking lag of 0 always means a healthy fast consumer — it can just mean the producer is idle.
  • Confusing lag (messages behind) with end-to-end latency (time behind).

context

open as a page

How do you read kafka-consumer-groups --describe output, and why is offset lag not the same as end-to-end latency?

level: middleimportance: must knowfreq 60%

basics

~20 s

kafka-consumer-groups --describe shows, per partition, CURRENT-OFFSET (committed), LOG-END-OFFSET, LAG (the difference), plus consumer-id/host. Offset lag counts messages behind. End-to-end latency is the time between when a record was produced and when it is consumed — a time, not a count — so high-throughput partitions can have big offset lag but tiny time latency.

open as a page

What are the records-lag and records-lag-max client metrics, and why might they differ from what kafka-consumer-groups reports?

level: middleimportance: must knowfreq 65%

basics

~20 s

records-lag is a per-partition consumer client (JMX) metric showing how far behind that fetch is; records-lag-max is the maximum across the consumer's assigned partitions. They come from the live client, so they only exist while the consumer is fetching and reflect what it has fetched, not committed.

open as a page

How does Burrow evaluate consumer-group health, and why is its sliding-window approach better than a single lag threshold?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Burrow watches a sliding window of recent committed offsets per partition and looks at the trend, not one threshold. It flags a group as ERR/STALLED if offsets stop advancing while lag is nonzero, or WARN if lag is consistently growing, instead of alerting on an arbitrary fixed lag number.

open as a page

Compare exposing consumer lag to Prometheus via kafka-exporter versus the JMX exporter. When would you use each?

level: seniorimportance: should knowfreq 45%

basics

~20 s

kafka-exporter talks the Kafka protocol to the brokers and computes committed lag itself (it works without any live consumer). The JMX exporter scrapes MBeans from a JVM, so it surfaces client metrics like records-lag-max but only while that consumer is running. Use kafka-exporter for durable group lag; JMX exporter for in-process metrics.

open as a page