How do you diagnose consumer lag and group state using kafka-consumer-groups --describe?
answer
- LAG = LOG-END-OFFSET - CURRENT-OFFSET
- CONSUMER-ID '-' = unassigned partition
- CURRENT-OFFSET '-' = never committed
- --state shows Stable/Empty/Rebalancing + coordinator
- watch lag trend, not a single snapshot
basics
~10 sRun kafka-consumer-groups.sh --bootstrap-server host --describe --group g. It shows per-partition CURRENT-OFFSET, LOG-END-OFFSET, and LAG (end minus current), plus which consumer/member owns each partition. High LAG means consumers are falling behind.
solid answer
~40 skafka-consumer-groups.sh --bootstrap-server host:9092 --describe --group payment-processor lists each assigned topic-partition with CURRENT-OFFSET (last committed offset for the group), LOG-END-OFFSET (the high water mark — next offset to be produced), and LAG = LOG-END-OFFSET - CURRENT-OFFSET. It also shows CONSUMER-ID, HOST, and CLIENT-ID so you can see which member owns each partition. Growing lag means consumption is slower than production. A partition with lag but no owner (CONSUMER-ID '-') indicates an unassigned partition — often a group that's down or rebalancing. --state shows the group's coordinator and state (Stable, PreparingRebalance, Empty, Dead). --members --verbose shows each member's assignment. --all-groups describes every group. '-' in CURRENT-OFFSET means the group never committed for that partition.
go deeper
Run --describe and identify the LAG column and that high lag is bad.
Interpret CURRENT/LOG-END/LAG, spot unassigned partitions and per-partition skew.
Distinguish commit-driven lag artifacts, rebalance states, and provisioning limits (partition count).
Define lag-trend alerting/SLOs and capacity rules across many groups, accounting for LSO and offset retention.
## The model A **consumer group** is a set of consumers that cooperatively read a set of topics; Kafka assigns each partition to exactly one member. The group's read position is stored as **committed offsets** in the internal `__consumer_offsets` topic, keyed by (group, topic, partition). Key offset terms: - **CURRENT-OFFSET** — the last offset the group has *committed* for a partition (where it will resume). - **LOG-END-OFFSET (LEO / high water mark)** — the offset just past the last record, i.e. the next offset a producer will write. - **LAG** = `LOG-END-OFFSET - CURRENT-OFFSET` — how many records are produced but not yet processed (or at least not yet committed) by the group. ## The command ``` kafka-consumer-groups.sh --bootstrap-server broker:9092 \ --describe --group payment-processor ``` Output columns: ``` GROUP TOPIC PARTITION CURRENT-OFFSET LOG-END-OFFSET LAG CONSUMER-ID HOST CLIENT-ID ``` - **CONSUMER-ID = '-'** → the partition is currently **unassigned** (no live member owns it). Lag will grow until a consumer picks it up. Common during an outage, a stuck rebalance, or when the group is Empty. - **CURRENT-OFFSET = '-'** → the group has **never committed** for that partition. - Sudden lag on *one* partition (others fine) points at a hot key / skewed partition or a single slow/stuck consumer instance, not a global slowdown. ## Useful variants - `--state` — prints the group's **coordinator** broker and **state**: `Stable` (assigned and running), `PreparingRebalance` / `CompletingRebalance` (mid-rebalance), `Empty` (offsets exist but no members), `Dead` (no offsets, will be removed). - `--members` and `--members --verbose` — per-member assignment list (which partitions each member owns); useful to spot an imbalanced assignment. - `--all-groups` — describe every group at once. - `--offsets` (the default for --describe) vs `--state` vs `--members` select which view you get. ## Reading lag correctly - Lag is computed from **committed** offsets, so a consumer that processes fast but commits slowly can *look* laggy. Conversely, auto-commit can make lag look healthy while records aren't actually processed. - Lag alone isn't an SLA; **lag trend over time** matters. Flat-but-nonzero lag is fine; monotonically rising lag means you're under-provisioned (add consumers up to the partition count, or speed up processing). - You can only add useful consumers up to the **partition count**; beyond that, extra members sit idle. ## Edge cases - After offsets expire (`offsets.retention.minutes`) for an inactive group, the group can show Empty / current-offset '-'. - During a rebalance, describe output can momentarily show no owners; re-run after it stabilizes. - Transactional/read-committed consumers lag is measured against the **last stable offset (LSO)**, not raw LEO, so open transactions can affect apparent lag.
- You see lag on one partition while the rest are at zero. What does that suggest?A skewed or 'hot' partition (key distribution sending most traffic to one partition) or a single stuck/slow consumer instance owning that partition. It's a local problem, not a global throughput shortage — adding more consumers won't help if all the load is on one partition.
- Can lag be misleading? How?Yes. Lag is based on committed offsets, so slow/batched commits inflate it even when processing keeps up; auto-commit can deflate it (commits advance before/without successful processing). Read-committed consumers measure against the last stable offset, so open transactions also affect apparent lag.
saying these in an interview costs you the question
- Defining LAG as CURRENT minus LOG-END (sign reversed).
- Saying adding consumers beyond the partition count reduces lag.
- Treating a single lag snapshot as an SLA instead of watching the trend.
- Not recognizing CONSUMER-ID '-' as an unowned partition.