skip to content

Why use a dedicated lag monitor like Burrow or Kafka Lag Exporter instead of just the CLI or records-lag-max, and how do their approaches differ?

level: seniorimportance: should knowfreq 45%

answer

  1. External = sees lag even when consumers are dead
  2. Burrow: consumes __consumer_offsets, window status OK/WARN/ERR/STALL, threshold-free
  3. Lag Exporter: Prometheus + time-lag seconds via interpolation
  4. Time-lag > record-lag for SLOs
  5. Feeds Grafana + KEDA autoscaling

basics

~20 s

Dedicated monitors continuously read every group's committed offsets and partition LEOs externally, so they see lag even when consumers are dead, and expose it to Prometheus/alerting. Burrow adds a threshold-free status evaluation; Lag Exporter adds time-based lag estimates.

solid answer

~40 s

The CLI is a manual snapshot and records-lag-max goes silent when a consumer dies, so neither suits continuous SLOs. Burrow (LinkedIn) and Kafka Lag Exporter (Lightbend/Sean Glover) run as external services that consume __consumer_offsets (or poll via admin) to track committed offsets and partition LEOs for all groups, regardless of whether any consumer is alive — catching stalled/absent consumers. Burrow's distinctive idea is evaluating consumer health by a sliding-window analysis of offset-commit trends (status OK/WARN/ERR/STALL) instead of a fixed lag threshold, avoiding noisy per-topic tuning. Kafka Lag Exporter focuses on Prometheus: it exposes per-partition lag and, crucially, an interpolated time-lag (estimated seconds behind) by mapping offsets to timestamps, which is far more actionable than raw record counts and feeds Grafana dashboards and HPA/KEDA autoscaling. Both decouple lag observability from the consumer process itself.

go deeper

for a junior

Know these are external tools that monitor lag and feed dashboards.

for a middle

Explain they read committed offsets and LEOs independently of consumers and export to Prometheus.

for a senior

Contrast Burrow's threshold-free status model with Lag Exporter's time-lag metric and justify when to use each.

for a principal

Design the full lag-observability stack — in-app metric, external time-lag exporter, alert semantics — and weigh the operational cost and cluster load they add.

## Why not just CLI or records-lag-max? - **CLI (`kafka-consumer-groups.sh --describe`)**: a one-off snapshot. No history, no alerting, no aggregation across groups. Fine for debugging, useless as a monitoring backbone. - **`records-lag-max` (JMX)**: real-time and great for a *running* app, but it is **client-side** — when consumers crash it goes silent while real lag keeps climbing. It also only sees each consumer's own assigned partitions. Production monitoring needs an **external** view that sees *all* groups and *all* partitions **even when consumers are down**, with history and alerting. That is what dedicated lag monitors provide. ## Burrow (LinkedIn) Burrow is a standalone Go service. Key ideas: - It **consumes the internal `__consumer_offsets` topic** to learn every group's committed offsets in near real time, and tracks broker LEOs — so it covers all groups automatically, no per-group config. - Its signature feature is **threshold-free evaluation**: instead of you guessing 'alert if lag > 10000', Burrow keeps a **sliding window** of recent offset commits per partition and classifies consumer **status** as OK / WARNING / ERROR, detecting conditions like **STALLED** (committing offsets but the position isn't advancing) and **STOPPED** (no commits at all). This avoids the false positives of static thresholds during legitimate traffic bursts. - It exposes an HTTP/JSON API; you bolt on notifiers (email, HTTP) or a Prometheus bridge. ## Kafka Lag Exporter (Lightbend / Sean Glover) A JVM service purpose-built for **Prometheus**: - Polls committed offsets and LEOs via the **admin/consumer-group APIs** for configured clusters and exposes metrics like `kafka_consumergroup_group_lag` (records) and **`kafka_consumergroup_group_lag_seconds`** (estimated time lag). - **Time-lag interpolation** is its standout: it periodically samples (offset, timestamp) points per partition and linearly interpolates to estimate *how many seconds behind* the committed offset is. Time-lag is far more meaningful than record counts because it normalizes across wildly different throughputs and maps directly to an end-to-end latency SLO. - Ships Grafana dashboards and is commonly wired to **KEDA**/HPA for **lag-driven autoscaling** of consumers (e.g. on Kubernetes). ## How their approaches differ | Aspect | Burrow | Kafka Lag Exporter | |---|---|---| | Health model | Window-based status (OK/WARN/ERR/STALL), threshold-free | Numeric metrics; you set thresholds/alerts | | Time-lag | Not native | First-class (`*_lag_seconds`) | | Native sink | HTTP/JSON API + notifiers | Prometheus metrics endpoint | | Discovery | Auto via `__consumer_offsets` | Polls admin API for configured clusters | ## The common, important property Both read lag **independently of the consumer process**, so a dead/stalled consumer still produces a signal — closing the exact blind spot of `records-lag-max`. The trade-off is added moving parts (a service to run), polling load on the cluster, and (for offset-topic consumers like Burrow) the need to keep up with `__consumer_offsets`. ## Putting it together A mature setup typically uses: `records-lag-max` for fast in-app/autoscaling signals, **plus** an external exporter (time-lag) for SLO alerting and absent-consumer detection, with Burrow-style status logic or `*_lag_seconds` thresholds to avoid noisy alerts.

  • What problem does Burrow's sliding-window status model solve compared to a fixed lag threshold?
    Fixed thresholds cause false alarms during legitimate bursts and miss slow stalls. Burrow evaluates the trend of offset commits over a window, classifying STALLED/STOPPED/OK without per-topic threshold tuning.
  • How does Kafka Lag Exporter compute time-lag in seconds?
    It samples (offset, timestamp) pairs per partition over time and linearly interpolates: given the committed offset, it estimates the timestamp of that record and subtracts from now to get seconds-behind.
  • What new operational cost do these monitors add?
    An extra service to run and secure, polling/consuming load on the cluster (especially keeping up with __consumer_offsets for Burrow), and the usual metrics pipeline (Prometheus, Grafana, alert rules).

saying these in an interview costs you the question

  • Saying these tools replace records-lag-max entirely — they complement it (external vs in-app).
  • Claiming the CLI is adequate for production alerting/trending.
  • Asserting record-count lag is as good as time-lag for cross-topic SLOs.
  • Thinking Burrow requires per-topic lag thresholds (its whole point is avoiding them).

context