skip to content

What is consumer lag in Kafka, and how do you compute it for a single partition?

level: juniorimportance: must knowfreq 80%

answer

  1. lag = LEO - committed offset
  2. per partition, summed per group
  3. rising lag = consumer slower than producer
  4. committed offset in __consumer_offsets
  5. max-partition lag often matters most

basics

~20 s

Consumer lag is how far behind a consumer is on a partition: log-end offset minus the consumer's committed offset. A lag of 0 means the consumer has read everything; a growing lag means it can't keep up with producers.

solid answer

~40 s

Consumer lag measures how many messages a consumer group still has to read on a partition. For one partition it is LogEndOffset (LEO, the offset of the next message the broker will write) minus the consumer group's committed offset for that partition. So lag = LEO - committedOffset. A group's total lag is the sum across all assigned partitions; the per-partition maximum is often what matters most because one slow partition can hold up downstream processing. Lag near 0 means the group is caught up. A steadily rising lag means consumption throughput is below production throughput, while a flat-but-nonzero lag means the consumer keeps pace but stays a fixed distance behind. Lag is the single most important consumer health signal because it directly reflects data-processing delay, independent of CPU or memory.

go deeper

for a junior

Know the formula lag = LEO - committed offset and that rising lag means the consumer can't keep up.

for a middle

Aggregate lag per group, distinguish total vs max-partition lag, and reason about flat vs rising lag regimes.

for a senior

Connect lag to retention risk, idle-producer false positives, and explain why max-partition lag drives alerting.

for a principal

Frame lag as the primary SLO signal for streaming pipelines and design alerting that separates throughput deficits from idle producers and from latency-based SLOs.

## What lag is Kafka stores each topic partition as an append-only log. Every message gets a monotonically increasing **offset** (0, 1, 2, ...). Two offsets matter for lag: - **Log-end offset (LEO)**: the offset that will be assigned to the *next* message produced to that partition. Equivalently, it is one past the last message currently on the broker. (More precisely the **high-water mark** — the last *committed/replicated* offset readable by consumers — is what consumers can reach; in healthy clusters LEO and high-water mark are effectively the same for lag math.) - **Committed offset**: the offset a consumer group has acknowledged it has processed, stored in the internal `__consumer_offsets` topic. By convention the committed offset is the offset of the *next* record to read, i.e. last-processed + 1. **Lag for one partition = LogEndOffset - committedOffset.** Example: producers have written up to offset 1000 (so LEO = 1000), and the group committed offset 970. Lag = 1000 - 970 = 30 messages still to read. ## Why it matters Lag is the most direct measure of how stale a consumer's view of the data is. Unlike CPU or memory, lag speaks the language of the business: "we are 30 messages / 5 seconds behind real time." There are three regimes: - **Lag ≈ 0**: caught up. - **Lag flat but nonzero**: consumer matches producer throughput but trails by a constant amount (e.g. a fixed batching/processing delay). - **Lag rising over time**: consumption throughput < production throughput. This is the danger signal — without intervention (more consumers, faster processing) lag grows unbounded and may eventually breach `retention.ms`, causing data loss for that group. ## Aggregating across partitions A consumer group is assigned many partitions. Total group lag = sum of per-partition lag. But the **max** per-partition lag is frequently the more actionable number, because a single hot or stuck partition can stall end-to-end pipelines even when total lag looks modest. ## Edge cases - **No committed offset yet** (brand-new group): tools may show lag as the full LEO or as unknown, depending on `auto.offset.reset`. - **Idle producer**: if no new messages arrive, lag naturally drops to 0 and stays there — that is healthy, not a stalled consumer. - **Committed offset can briefly exceed nothing** but should never legitimately exceed LEO; a negative lag usually signals a tooling/race artifact (offset read before LEO refresh) rather than a real state. - Lag is **per consumer group**, not per topic: two groups reading the same topic have independent lag.

  • If lag is flat at a nonzero value, is the consumer unhealthy?
    Not necessarily. Flat lag means the consumer keeps pace with the producer but trails by a constant amount (e.g. a fixed processing/batching delay). Rising lag is the unhealthy signal; flat lag is steady-state.
  • Why might total group lag be misleading compared to max per-partition lag?
    Total lag can look small while one partition is badly stuck. Since processing and ordering are per-partition, a single high-lag partition can stall a pipeline even if the sum across partitions is modest, so max-partition lag is often the better alert signal.

saying these in an interview costs you the question

  • Saying lag is committed offset minus LEO (the subtraction is backwards).
  • Claiming lag is per-topic rather than per consumer group.
  • Thinking lag of 0 always means a healthy fast consumer — it can just mean the producer is idle.
  • Confusing lag (messages behind) with end-to-end latency (time behind).

context