skip to content

Consumer Lag and Measurement

Measuring lag as log-end offset minus committed offset, and the tooling that exposes it per partition. Interviewers ask because lag is the primary health and autoscaling signal for consumers.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

How do you inspect consumer lag from the command line, and what do the columns of kafka-consumer-groups.sh --describe mean?

level: juniorimportance: must knowfreq 70%

answer

  1. kafka-consumer-groups.sh --describe --group
  2. CURRENT-OFFSET = committed, LOG-END-OFFSET = latest
  3. LAG column = difference
  4. '-' lag = no committed offset
  5. Snapshot only, not a trend

basics

~10 s

Run kafka-consumer-groups.sh --bootstrap-server host:port --describe --group <group>. It lists each partition with CURRENT-OFFSET (committed), LOG-END-OFFSET (latest), and LAG = the difference, plus the owning consumer.

solid answer

~40 s

The standard tool is kafka-consumer-groups.sh (or kafka-consumer-groups on some distros). To see lag: `kafka-consumer-groups.sh --bootstrap-server broker:9092 --describe --group my-group`. Per (topic, partition) it shows CURRENT-OFFSET (the group's committed offset), LOG-END-OFFSET (the partition's latest offset), LAG (LOG-END-OFFSET minus CURRENT-OFFSET), and CONSUMER-ID / HOST / CLIENT-ID identifying which member currently owns the partition. A LAG value of '-' usually means no committed offset exists yet; a CONSUMER-ID of '-' means the partition is assigned to no active member (the group may be empty or rebalancing). You can also `--list` all groups, `--describe --members` to see assignments, and `--reset-offsets` to rewind/skip. It's a point-in-time snapshot — trend over time needs the JMX metric or an exporter.

go deeper

for a junior

Be able to run --describe and name what CURRENT-OFFSET, LOG-END-OFFSET and LAG mean.

for a middle

Interpret special values (dashes), know related subcommands like --reset-offsets and --state.

for a senior

Use the CLI to diagnose skew/stuck partitions and explain why a snapshot is insufficient for production monitoring.

for a principal

Position CLI as debugging only; mandate exporters/JMX for SLOs and codify offset-reset runbooks for incident recovery.

## The tool `kafka-consumer-groups.sh` (Linux/macOS) or `kafka-consumer-groups.bat` (Windows), shipped in Kafka's `bin/` directory, is the canonical admin CLI for consumer groups. Newer packaging may name it `kafka-consumer-groups`. It talks to the cluster via the **bootstrap server**, not ZooKeeper (the old `--zookeeper` flag is removed in modern Kafka). ## Describing a group ``` kafka-consumer-groups.sh \ --bootstrap-server broker1:9092 \ --describe \ --group my-app ``` Output is one row per assigned **(topic, partition)**: | Column | Meaning | |---|---| | `GROUP` | The consumer group id. | | `TOPIC` | Topic name. | | `PARTITION` | Partition number. | | `CURRENT-OFFSET` | The group's **committed** offset for this partition (what it has acknowledged). | | `LOG-END-OFFSET` | The partition's **latest** offset (one past the newest record). | | `LAG` | `LOG-END-OFFSET - CURRENT-OFFSET` — records not yet consumed. | | `CONSUMER-ID` | The member (instance) currently owning the partition. | | `HOST` | The member's host/IP. | | `CLIENT-ID` | The client.id the member set. | ## Reading special values - **`LAG = -`**: there is no committed offset for that partition yet (new group, or it was reset/expired), so lag can't be computed. - **`CONSUMER-ID = -` / `HOST = -`**: the partition is assigned to no live member — the group is empty (offsets retained) or mid-rebalance. Lag may still show, computed against retained committed offsets. - A group with many partitions showing high lag on a *few* partitions and zero on others indicates **partition skew** (uneven key distribution or a slow/stuck consumer on those partitions). ## Related subcommands - `--list` — list all group ids. - `--describe --members` / `--describe --members --verbose` — show membership and per-member assignments. - `--describe --state` — show group state (Stable, PreparingRebalance, Empty, Dead, ...). - `--reset-offsets --to-earliest|--to-latest|--shift-by N|--to-datetime ... --execute` — rewind/skip a group's offsets (must be stopped/empty). - `--delete-offsets` — drop committed offsets for specific partitions. ## Limitations The `--describe` output is a **single snapshot**. It's perfect for ad-hoc debugging but useless for trending, alerting, or autoscaling — for that you need the in-process JMX metric `records-lag-max` or an external exporter (Burrow, Kafka Lag Exporter) feeding a time-series database.

  • Why does the CONSUMER-ID column sometimes show a dash even though lag is reported?
    The partition has retained committed offsets but no live member currently owns it — the group is empty or rebalancing. Lag is computed against the stored committed offset, so it still displays.
  • The CLI uses --bootstrap-server, not --zookeeper. Why?
    Consumer offsets moved from ZooKeeper into the internal __consumer_offsets topic years ago, and modern Kafka (KRaft) has no ZooKeeper. The tool queries the brokers' group coordinator directly.

saying these in an interview costs you the question

  • Using the removed --zookeeper flag instead of --bootstrap-server.
  • Treating the CLI snapshot as a monitoring solution for alerting/trends.
  • Reading CURRENT-OFFSET as the latest record offset (it is the committed/consumed offset).
  • Assuming LAG = '-' means zero lag (it means no committed offset exists).

context

open as a page

What is consumer lag in Kafka, and how is it calculated for a single partition?

level: juniorimportance: must knowfreq 80%

basics

~10 s

Consumer lag is how far behind a consumer is on a partition: the latest message offset (log-end-offset) minus the offset the consumer has committed. Lag of 0 means fully caught up.

open as a page

What is the records-lag-max JMX metric, where does it come from, and what are its strengths and blind spots compared to broker-side lag?

level: middleimportance: must knowfreq 60%

basics

~10 s

records-lag-max is a client-side JMX metric exposed by each consumer reporting the maximum lag across the partitions it currently fetches. It's cheap and real-time but only sees assigned partitions of running consumers.

open as a page

Why use a dedicated lag monitor like Burrow or Kafka Lag Exporter instead of just the CLI or records-lag-max, and how do their approaches differ?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Dedicated monitors continuously read every group's committed offsets and partition LEOs externally, so they see lag even when consumers are dead, and expose it to Prometheus/alerting. Burrow adds a threshold-free status evaluation; Lag Exporter adds time-based lag estimates.

open as a page

How do you diagnose partition skew from lag data, and how would you design lag-driven autoscaling for a consumer group?

level: principalimportance: should knowfreq 35%

basics

~20 s

Skew shows as lag concentrated on a few partitions while others sit at zero — usually hot keys or a stuck/slow consumer. For autoscaling, scale consumers on total/time lag, but cap replicas at the partition count, throttle on rebalances, and prefer time-lag with hysteresis to avoid flapping.

open as a page