skip to content

How do you read kafka-consumer-groups --describe output, and why is offset lag not the same as end-to-end latency?

level: middleimportance: must knowfreq 60%

answer

  1. --describe: CURRENT-OFFSET, LOG-END-OFFSET, LAG, CONSUMER-ID/HOST
  2. dash = no commit / no active owner
  3. lag = message count; e2e latency = time
  4. convert via throughput (msgs/sec)
  5. measure latency from record.timestamp() to processing time

basics

~20 s

kafka-consumer-groups --describe shows, per partition, CURRENT-OFFSET (committed), LOG-END-OFFSET, LAG (the difference), plus consumer-id/host. Offset lag counts messages behind. End-to-end latency is the time between when a record was produced and when it is consumed — a time, not a count — so high-throughput partitions can have big offset lag but tiny time latency.

solid answer

~50 s

Run kafka-consumer-groups --bootstrap-server <b> --describe --group <g>. Each row is a topic-partition with CURRENT-OFFSET (the group's committed offset from __consumer_offsets), LOG-END-OFFSET (the broker's LEO), LAG = LOG-END-OFFSET - CURRENT-OFFSET, plus CONSUMER-ID, HOST, and CLIENT-ID. A '-' in CURRENT-OFFSET/LAG means no commit yet; a '-' in the consumer columns means no active member owns that partition. Offset lag is a *count of messages*, while end-to-end latency is the wall-clock time from when a record was appended (its timestamp) to when the consumer processes it. The two differ because they have different units: a partition doing 100k msg/s with 100k lag is only ~1 second behind in time, whereas a partition at 1 msg/s with lag 100 is ~100 seconds behind. To measure true e2e latency you compare the record's timestamp (or an embedded produce time) to consume time, since offset count alone can't tell you elapsed time.

go deeper

for a junior

Read the columns: CURRENT-OFFSET, LOG-END-OFFSET, LAG, and know LAG is a message count.

for a middle

Interpret dashes, and explain why offset lag is not end-to-end latency and how throughput links them.

for a senior

Measure true e2e latency via record timestamps, account for CreateTime vs LogAppendTime and clock skew.

for a principal

Decide which signal backs which SLO (backlog/retention vs latency SLA) and design measurement that survives transformations and skew.

## The CLI ``` kafka-consumer-groups.sh --bootstrap-server host:9092 --describe --group my-group ``` For every partition the group is responsible for, you get columns: - **GROUP** — the consumer group id. - **TOPIC** / **PARTITION** — the topic-partition. - **CURRENT-OFFSET** — the group's **committed** offset (from `__consumer_offsets`); the offset of the next record it intends to read. - **LOG-END-OFFSET** — the broker's **LEO** for that partition (next offset to be produced). - **LAG** — `LOG-END-OFFSET - CURRENT-OFFSET` (the message-count lag). - **CONSUMER-ID** — the member instance currently assigned the partition. - **HOST** / **CLIENT-ID** — where that member runs and its client id. Reading rules: - A `-` (dash) in **CURRENT-OFFSET / LAG** means the group has **no committed offset** for that partition yet (brand-new group, or it never reached that partition). - A `-` in **CONSUMER-ID / HOST** means **no active member owns** the partition right now (group is empty/stopped, or the partition is unassigned mid-rebalance). The CLI still prints CURRENT-OFFSET, LOG-END-OFFSET, and LAG from stored data, which is how you inspect a stopped group. - Useful variants: `--all-groups`, `--state`, `--members --verbose`, and `--reset-offsets` (for repositioning). ## Why offset lag ≠ end-to-end latency These are **different quantities with different units**: - **Offset lag** = a **count of messages** the consumer hasn't read (LEO - committed). - **End-to-end latency** = a **duration**: the wall-clock time from when a record was **produced/appended** (its log-append or create timestamp) until it is **consumed/processed**. The conversion between them depends on **throughput**: - 100,000 msg/s with offset lag 100,000 → roughly **1 second** of time latency. - 1 msg/s with offset lag 100 → roughly **100 seconds** of time latency. So identical offset lag can mean wildly different time delays. Conversely a small offset lag on a slow topic can hide a large staleness. ## Measuring true end-to-end latency Because offset lag can't be converted to time without knowing throughput, measure latency directly: - Use the record's **timestamp** (`ConsumerRecord.timestamp()` / `timestampType` — either `CreateTime` set by the producer or `LogAppendTime` set by the broker) and compare it to the **time you finish processing** the record: `e2eLatency = processingTime - record.timestamp()`. - Or embed an explicit produce timestamp in the payload/header for end-to-end (producer-to-final-sink) timing that survives transformations. - Watch for **clock skew** between producer and consumer hosts when using CreateTime; LogAppendTime uses the broker clock and avoids producer skew but only measures broker-to-consumer, not true producer-to-consumer. ## When to use which - **Offset lag** is cheap, always available (CLI/exporter), and great for trend/backlog and retention-risk alarms. - **End-to-end latency** is what SLAs are usually written in ("events processed within 2s"), so for latency SLOs you must measure time directly rather than infer it from lag. - Together: rising offset lag warns of a building backlog; the latency metric tells you whether that backlog already breaches the time SLA.

  • In --describe output, what do dashes in CONSUMER-ID and HOST mean while LAG still shows a number?
    No active consumer member currently owns that partition (group stopped or mid-rebalance), but CURRENT-OFFSET/LOG-END-OFFSET/LAG are still printed from stored committed offsets and broker LEO, so you can inspect backlog of an idle group.
  • How would you measure true end-to-end latency rather than inferring it from offset lag?
    Compare the record's timestamp (ConsumerRecord.timestamp(), CreateTime or LogAppendTime) — or an embedded produce time in a header — against the wall-clock time the consumer finishes processing it. Mind producer/consumer clock skew when using CreateTime.

saying these in an interview costs you the question

  • Treating LAG (a message count) and end-to-end latency (a time) as interchangeable.
  • Assuming a high LAG always means a serious time delay regardless of throughput.
  • Thinking the CLI cannot show lag for a stopped group (it can, from stored offsets).
  • Ignoring clock skew when computing latency from CreateTime timestamps.

context