What is consumer lag in Kafka, and how do you compute it for a single partition?
answer
- lag = LEO - committed offset
- per partition, summed per group
- rising lag = consumer slower than producer
- committed offset in __consumer_offsets
- max-partition lag often matters most
basics
~20 sConsumer lag is how far behind a consumer is on a partition: log-end offset minus the consumer's committed offset. A lag of 0 means the consumer has read everything; a growing lag means it can't keep up with producers.
solid answer
~40 sConsumer lag measures how many messages a consumer group still has to read on a partition. For one partition it is LogEndOffset (LEO, the offset of the next message the broker will write) minus the consumer group's committed offset for that partition. So lag = LEO - committedOffset. A group's total lag is the sum across all assigned partitions; the per-partition maximum is often what matters most because one slow partition can hold up downstream processing. Lag near 0 means the group is caught up. A steadily rising lag means consumption throughput is below production throughput, while a flat-but-nonzero lag means the consumer keeps pace but stays a fixed distance behind. Lag is the single most important consumer health signal because it directly reflects data-processing delay, independent of CPU or memory.
go deeper
Know the formula lag = LEO - committed offset and that rising lag means the consumer can't keep up.
Aggregate lag per group, distinguish total vs max-partition lag, and reason about flat vs rising lag regimes.
Connect lag to retention risk, idle-producer false positives, and explain why max-partition lag drives alerting.
Frame lag as the primary SLO signal for streaming pipelines and design alerting that separates throughput deficits from idle producers and from latency-based SLOs.
## What lag is Kafka stores each topic partition as an append-only log. Every message gets a monotonically increasing **offset** (0, 1, 2, ...). Two offsets matter for lag: - **Log-end offset (LEO)**: the offset that will be assigned to the *next* message produced to that partition. Equivalently, it is one past the last message currently on the broker. (More precisely the **high-water mark** — the last *committed/replicated* offset readable by consumers — is what consumers can reach; in healthy clusters LEO and high-water mark are effectively the same for lag math.) - **Committed offset**: the offset a consumer group has acknowledged it has processed, stored in the internal `__consumer_offsets` topic. By convention the committed offset is the offset of the *next* record to read, i.e. last-processed + 1. **Lag for one partition = LogEndOffset - committedOffset.** Example: producers have written up to offset 1000 (so LEO = 1000), and the group committed offset 970. Lag = 1000 - 970 = 30 messages still to read. ## Why it matters Lag is the most direct measure of how stale a consumer's view of the data is. Unlike CPU or memory, lag speaks the language of the business: "we are 30 messages / 5 seconds behind real time." There are three regimes: - **Lag ≈ 0**: caught up. - **Lag flat but nonzero**: consumer matches producer throughput but trails by a constant amount (e.g. a fixed batching/processing delay). - **Lag rising over time**: consumption throughput < production throughput. This is the danger signal — without intervention (more consumers, faster processing) lag grows unbounded and may eventually breach `retention.ms`, causing data loss for that group. ## Aggregating across partitions A consumer group is assigned many partitions. Total group lag = sum of per-partition lag. But the **max** per-partition lag is frequently the more actionable number, because a single hot or stuck partition can stall end-to-end pipelines even when total lag looks modest. ## Edge cases - **No committed offset yet** (brand-new group): tools may show lag as the full LEO or as unknown, depending on `auto.offset.reset`. - **Idle producer**: if no new messages arrive, lag naturally drops to 0 and stays there — that is healthy, not a stalled consumer. - **Committed offset can briefly exceed nothing** but should never legitimately exceed LEO; a negative lag usually signals a tooling/race artifact (offset read before LEO refresh) rather than a real state. - Lag is **per consumer group**, not per topic: two groups reading the same topic have independent lag.
- If lag is flat at a nonzero value, is the consumer unhealthy?Not necessarily. Flat lag means the consumer keeps pace with the producer but trails by a constant amount (e.g. a fixed processing/batching delay). Rising lag is the unhealthy signal; flat lag is steady-state.
- Why might total group lag be misleading compared to max per-partition lag?Total lag can look small while one partition is badly stuck. Since processing and ordering are per-partition, a single high-lag partition can stall a pipeline even if the sum across partitions is modest, so max-partition lag is often the better alert signal.
saying these in an interview costs you the question
- Saying lag is committed offset minus LEO (the subtraction is backwards).
- Claiming lag is per-topic rather than per consumer group.
- Thinking lag of 0 always means a healthy fast consumer — it can just mean the producer is idle.
- Confusing lag (messages behind) with end-to-end latency (time behind).