skip to content

What is consumer lag in Kafka, and how is it calculated for a single partition?

level: juniorimportance: must knowfreq 80%

answer

  1. LEO minus committed offset
  2. Per partition, summed per group
  3. Records not time
  4. 0 = caught up, growing = falling behind
  5. Lives on (group, partition)

basics

~10 s

Consumer lag is how far behind a consumer is on a partition: the latest message offset (log-end-offset) minus the offset the consumer has committed. Lag of 0 means fully caught up.

solid answer

~40 s

Consumer lag measures how many records a consumer still has to read on a partition. Each partition has a log-end-offset (LEO) — the offset just past the newest committed record produced. Each consumer group tracks a committed offset per partition (the position it has acknowledged processing up to). Lag = LEO - committed offset. A lag of 0 means the group has consumed everything currently in the log; a growing lag means producers are writing faster than the group consumes, or consumption stalled. Lag is per-partition; total group lag is the sum across the partitions the group is subscribed to. It is the single most important health signal for a consumer group, because it directly reflects end-to-end processing delay (when paired with throughput) and whether the group is keeping up with the topic.

go deeper

for a junior

Know the formula LEO - committed offset and that 0 means caught up.

for a middle

Explain that lag is per partition/per group, summed for a total, and stored offsets live in __consumer_offsets.

for a senior

Distinguish record-lag from time-lag, discuss high watermark vs LEO, and reason about what growing vs flat-high lag implies operationally.

for a principal

Frame lag as the core SLO signal feeding alerting/autoscaling and teach the team why time-based lag matters more than raw record counts.

## Core definitions A Kafka **topic** is split into **partitions**; each partition is an append-only ordered log of records. Every record in a partition has a monotonically increasing integer **offset** (0, 1, 2, ...). - **Log-end-offset (LEO)**: the offset that will be assigned to the *next* record appended to the partition — i.e. one past the last record. If a partition holds records at offsets 0..99, the LEO is 100. (Strictly, consumers can only read up to the **high watermark**, the highest offset replicated to all in-sync replicas; for a healthy partition LEO and high watermark are equal or very close.) - **Committed offset**: a **consumer group** periodically writes, to the internal `__consumer_offsets` topic, the offset it has finished processing for each partition. This is the *next* offset it intends to fetch. If the group has committed offset 90, it has processed records 0..89. ## The lag formula ``` lag(partition) = log-end-offset - committed-offset ``` If LEO = 100 and committed = 90, lag = 10: ten records are produced but not yet consumed by that group. **Total group lag** = sum of per-partition lag over all assigned partitions. ## Why it matters Lag is the primary measure of whether a consumer group is *keeping up*. - **Lag ≈ 0 and stable** → healthy, consuming as fast as produced. - **Lag growing** → producers outpace consumers, or consumers stalled (slow processing, crash, rebalance, GC pause, downstream dependency slow). - **Lag flat but high** → consuming at the same rate as production but with a permanent backlog (under-provisioned). ## Edge cases and nuances - **Lag is in records, not time.** 1,000 records of lag could be milliseconds or hours of delay depending on throughput. Time-based lag (estimating *when* the lagged record was produced) is more meaningful but needs extra computation (e.g. Burrow, Kafka Lag Exporter). - **Per group, not per consumer.** Lag is a property of the (group, partition) pair. The current owning consumer can change on rebalance, but the committed offset persists in `__consumer_offsets`. - **Negative or odd readings** can appear transiently if the committed offset is stale relative to a just-advanced LEO read at a slightly different instant, or with transactional/aborted records inflating offsets via control records. - **No committed offset yet** (brand-new group) → tooling often shows lag as unknown/`-`, because there is no committed position to subtract. ## How you observe it The CLI `kafka-consumer-groups.sh --describe --group <g>` prints, per partition: `CURRENT-OFFSET` (committed), `LOG-END-OFFSET`, and `LAG` (the difference). The client-side JMX metric `records-lag-max` exposes the worst partition's lag from inside the consumer.

  • Where does Kafka store the committed offset that lag is computed against?
    In the internal compacted topic __consumer_offsets, keyed by (group, topic, partition). The broker reads it to answer kafka-consumer-groups.sh --describe.
  • Why can 1,000 records of lag be harmless in one topic but alarming in another?
    Lag is counted in records, not time. At high throughput 1,000 records may be milliseconds behind; at low throughput it could be minutes. Time-lag is the more meaningful SLO signal.

saying these in an interview costs you the question

  • Saying lag is measured in seconds/time by default — it is records by default.
  • Confusing committed offset with the current consumer's in-memory position (uncommitted reads don't reduce lag).
  • Thinking lag is per-consumer rather than per (group, partition).
  • Believing lag of 0 guarantees low latency regardless of throughput.

context