Why use a dedicated lag monitor like Burrow or Kafka Lag Exporter instead of just the CLI or records-lag-max, and how do their approaches differ?
answer
- External = sees lag even when consumers are dead
- Burrow: consumes __consumer_offsets, window status OK/WARN/ERR/STALL, threshold-free
- Lag Exporter: Prometheus + time-lag seconds via interpolation
- Time-lag > record-lag for SLOs
- Feeds Grafana + KEDA autoscaling
basics
~20 sDedicated monitors continuously read every group's committed offsets and partition LEOs externally, so they see lag even when consumers are dead, and expose it to Prometheus/alerting. Burrow adds a threshold-free status evaluation; Lag Exporter adds time-based lag estimates.
solid answer
~40 sThe CLI is a manual snapshot and records-lag-max goes silent when a consumer dies, so neither suits continuous SLOs. Burrow (LinkedIn) and Kafka Lag Exporter (Lightbend/Sean Glover) run as external services that consume __consumer_offsets (or poll via admin) to track committed offsets and partition LEOs for all groups, regardless of whether any consumer is alive — catching stalled/absent consumers. Burrow's distinctive idea is evaluating consumer health by a sliding-window analysis of offset-commit trends (status OK/WARN/ERR/STALL) instead of a fixed lag threshold, avoiding noisy per-topic tuning. Kafka Lag Exporter focuses on Prometheus: it exposes per-partition lag and, crucially, an interpolated time-lag (estimated seconds behind) by mapping offsets to timestamps, which is far more actionable than raw record counts and feeds Grafana dashboards and HPA/KEDA autoscaling. Both decouple lag observability from the consumer process itself.
go deeper
Know these are external tools that monitor lag and feed dashboards.
Explain they read committed offsets and LEOs independently of consumers and export to Prometheus.
Contrast Burrow's threshold-free status model with Lag Exporter's time-lag metric and justify when to use each.
Design the full lag-observability stack — in-app metric, external time-lag exporter, alert semantics — and weigh the operational cost and cluster load they add.
## Why not just CLI or records-lag-max? - **CLI (`kafka-consumer-groups.sh --describe`)**: a one-off snapshot. No history, no alerting, no aggregation across groups. Fine for debugging, useless as a monitoring backbone. - **`records-lag-max` (JMX)**: real-time and great for a *running* app, but it is **client-side** — when consumers crash it goes silent while real lag keeps climbing. It also only sees each consumer's own assigned partitions. Production monitoring needs an **external** view that sees *all* groups and *all* partitions **even when consumers are down**, with history and alerting. That is what dedicated lag monitors provide. ## Burrow (LinkedIn) Burrow is a standalone Go service. Key ideas: - It **consumes the internal `__consumer_offsets` topic** to learn every group's committed offsets in near real time, and tracks broker LEOs — so it covers all groups automatically, no per-group config. - Its signature feature is **threshold-free evaluation**: instead of you guessing 'alert if lag > 10000', Burrow keeps a **sliding window** of recent offset commits per partition and classifies consumer **status** as OK / WARNING / ERROR, detecting conditions like **STALLED** (committing offsets but the position isn't advancing) and **STOPPED** (no commits at all). This avoids the false positives of static thresholds during legitimate traffic bursts. - It exposes an HTTP/JSON API; you bolt on notifiers (email, HTTP) or a Prometheus bridge. ## Kafka Lag Exporter (Lightbend / Sean Glover) A JVM service purpose-built for **Prometheus**: - Polls committed offsets and LEOs via the **admin/consumer-group APIs** for configured clusters and exposes metrics like `kafka_consumergroup_group_lag` (records) and **`kafka_consumergroup_group_lag_seconds`** (estimated time lag). - **Time-lag interpolation** is its standout: it periodically samples (offset, timestamp) points per partition and linearly interpolates to estimate *how many seconds behind* the committed offset is. Time-lag is far more meaningful than record counts because it normalizes across wildly different throughputs and maps directly to an end-to-end latency SLO. - Ships Grafana dashboards and is commonly wired to **KEDA**/HPA for **lag-driven autoscaling** of consumers (e.g. on Kubernetes). ## How their approaches differ | Aspect | Burrow | Kafka Lag Exporter | |---|---|---| | Health model | Window-based status (OK/WARN/ERR/STALL), threshold-free | Numeric metrics; you set thresholds/alerts | | Time-lag | Not native | First-class (`*_lag_seconds`) | | Native sink | HTTP/JSON API + notifiers | Prometheus metrics endpoint | | Discovery | Auto via `__consumer_offsets` | Polls admin API for configured clusters | ## The common, important property Both read lag **independently of the consumer process**, so a dead/stalled consumer still produces a signal — closing the exact blind spot of `records-lag-max`. The trade-off is added moving parts (a service to run), polling load on the cluster, and (for offset-topic consumers like Burrow) the need to keep up with `__consumer_offsets`. ## Putting it together A mature setup typically uses: `records-lag-max` for fast in-app/autoscaling signals, **plus** an external exporter (time-lag) for SLO alerting and absent-consumer detection, with Burrow-style status logic or `*_lag_seconds` thresholds to avoid noisy alerts.
- What problem does Burrow's sliding-window status model solve compared to a fixed lag threshold?Fixed thresholds cause false alarms during legitimate bursts and miss slow stalls. Burrow evaluates the trend of offset commits over a window, classifying STALLED/STOPPED/OK without per-topic threshold tuning.
- How does Kafka Lag Exporter compute time-lag in seconds?It samples (offset, timestamp) pairs per partition over time and linearly interpolates: given the committed offset, it estimates the timestamp of that record and subtracts from now to get seconds-behind.
- What new operational cost do these monitors add?An extra service to run and secure, polling/consuming load on the cluster (especially keeping up with __consumer_offsets for Burrow), and the usual metrics pipeline (Prometheus, Grafana, alert rules).
saying these in an interview costs you the question
- Saying these tools replace records-lag-max entirely — they complement it (external vs in-app).
- Claiming the CLI is adequate for production alerting/trending.
- Asserting record-count lag is as good as time-lag for cross-topic SLOs.
- Thinking Burrow requires per-topic lag thresholds (its whole point is avoiding them).