What is consumer lag in Kafka, and how is it calculated for a single partition?
answer
- LEO minus committed offset
- Per partition, summed per group
- Records not time
- 0 = caught up, growing = falling behind
- Lives on (group, partition)
basics
~10 sConsumer lag is how far behind a consumer is on a partition: the latest message offset (log-end-offset) minus the offset the consumer has committed. Lag of 0 means fully caught up.
solid answer
~40 sConsumer lag measures how many records a consumer still has to read on a partition. Each partition has a log-end-offset (LEO) — the offset just past the newest committed record produced. Each consumer group tracks a committed offset per partition (the position it has acknowledged processing up to). Lag = LEO - committed offset. A lag of 0 means the group has consumed everything currently in the log; a growing lag means producers are writing faster than the group consumes, or consumption stalled. Lag is per-partition; total group lag is the sum across the partitions the group is subscribed to. It is the single most important health signal for a consumer group, because it directly reflects end-to-end processing delay (when paired with throughput) and whether the group is keeping up with the topic.
go deeper
Know the formula LEO - committed offset and that 0 means caught up.
Explain that lag is per partition/per group, summed for a total, and stored offsets live in __consumer_offsets.
Distinguish record-lag from time-lag, discuss high watermark vs LEO, and reason about what growing vs flat-high lag implies operationally.
Frame lag as the core SLO signal feeding alerting/autoscaling and teach the team why time-based lag matters more than raw record counts.
## Core definitions A Kafka **topic** is split into **partitions**; each partition is an append-only ordered log of records. Every record in a partition has a monotonically increasing integer **offset** (0, 1, 2, ...). - **Log-end-offset (LEO)**: the offset that will be assigned to the *next* record appended to the partition — i.e. one past the last record. If a partition holds records at offsets 0..99, the LEO is 100. (Strictly, consumers can only read up to the **high watermark**, the highest offset replicated to all in-sync replicas; for a healthy partition LEO and high watermark are equal or very close.) - **Committed offset**: a **consumer group** periodically writes, to the internal `__consumer_offsets` topic, the offset it has finished processing for each partition. This is the *next* offset it intends to fetch. If the group has committed offset 90, it has processed records 0..89. ## The lag formula ``` lag(partition) = log-end-offset - committed-offset ``` If LEO = 100 and committed = 90, lag = 10: ten records are produced but not yet consumed by that group. **Total group lag** = sum of per-partition lag over all assigned partitions. ## Why it matters Lag is the primary measure of whether a consumer group is *keeping up*. - **Lag ≈ 0 and stable** → healthy, consuming as fast as produced. - **Lag growing** → producers outpace consumers, or consumers stalled (slow processing, crash, rebalance, GC pause, downstream dependency slow). - **Lag flat but high** → consuming at the same rate as production but with a permanent backlog (under-provisioned). ## Edge cases and nuances - **Lag is in records, not time.** 1,000 records of lag could be milliseconds or hours of delay depending on throughput. Time-based lag (estimating *when* the lagged record was produced) is more meaningful but needs extra computation (e.g. Burrow, Kafka Lag Exporter). - **Per group, not per consumer.** Lag is a property of the (group, partition) pair. The current owning consumer can change on rebalance, but the committed offset persists in `__consumer_offsets`. - **Negative or odd readings** can appear transiently if the committed offset is stale relative to a just-advanced LEO read at a slightly different instant, or with transactional/aborted records inflating offsets via control records. - **No committed offset yet** (brand-new group) → tooling often shows lag as unknown/`-`, because there is no committed position to subtract. ## How you observe it The CLI `kafka-consumer-groups.sh --describe --group <g>` prints, per partition: `CURRENT-OFFSET` (committed), `LOG-END-OFFSET`, and `LAG` (the difference). The client-side JMX metric `records-lag-max` exposes the worst partition's lag from inside the consumer.
- Where does Kafka store the committed offset that lag is computed against?In the internal compacted topic __consumer_offsets, keyed by (group, topic, partition). The broker reads it to answer kafka-consumer-groups.sh --describe.
- Why can 1,000 records of lag be harmless in one topic but alarming in another?Lag is counted in records, not time. At high throughput 1,000 records may be milliseconds behind; at low throughput it could be minutes. Time-lag is the more meaningful SLO signal.
saying these in an interview costs you the question
- Saying lag is measured in seconds/time by default — it is records by default.
- Confusing committed offset with the current consumer's in-memory position (uncommitted reads don't reduce lag).
- Thinking lag is per-consumer rather than per (group, partition).
- Believing lag of 0 guarantees low latency regardless of throughput.