How does Burrow evaluate consumer-group health, and why is its sliding-window approach better than a single lag threshold?
answer
- consumes __consumer_offsets, tracks window of commits
- trend-based, threshold-free
- OK / WARN(lag growing) / STALL(no progress, lag>0) / STOP(no commits)
- group status = worst partition
- beats fixed threshold across throughput tiers
basics
~20 sBurrow watches a sliding window of recent committed offsets per partition and looks at the trend, not one threshold. It flags a group as ERR/STALLED if offsets stop advancing while lag is nonzero, or WARN if lag is consistently growing, instead of alerting on an arbitrary fixed lag number.
solid answer
~50 sBurrow is a consumer-lag monitor that consumes __consumer_offsets and, per partition, keeps a sliding window of the last N committed offsets plus their lag and timestamps. Its evaluation rules classify each partition's status by looking at the window's trend rather than a single threshold. Key statuses: OK; WARN when lag is consistently increasing across the window; ERR/STOP/STALL when the committed offset is not advancing while lag is nonzero (the consumer is stuck), or when the consumer hasn't committed within the expected interval. A group's status is the worst of its partitions. The advantage over a fixed threshold is that it is self-tuning: a high-throughput topic legitimately runs at high absolute lag, so a static '> 10000' rule produces false alarms, while a low-throughput topic can be stalled at lag 50. Burrow instead asks 'is this consumer making progress and keeping pace?', which is threshold-free and works across heterogeneous topics.
go deeper
Know Burrow watches consumer offsets and flags groups that fall behind, without you setting a lag number.
Describe the sliding window of commits and the OK/WARN/ERR statuses based on offset and lag movement.
Articulate why trend-based evaluation beats thresholds across throughput tiers and how STALL differs from WARN.
Design multi-tenant lag alerting strategy, weighing window size vs responsiveness, pairing Burrow status with a TSDB, and choosing committed-lag evaluation semantics.
## What Burrow is Burrow (originally from LinkedIn) is a standalone service that monitors Kafka consumer lag **without requiring you to pick lag thresholds**. It connects to the cluster, **consumes the internal `__consumer_offsets` topic** to learn every group's committed offsets in real time, and separately tracks broker log-end offsets. It exposes status over an HTTP API (and can push notifications). ## The sliding window For each (group, topic, partition) Burrow maintains a ring buffer of the most recent **offset commits**. Each entry records: the committed offset, the lag at that moment (LEO - committed), and the timestamp. The window length is configurable (`intervals`), e.g. the last 10 commits. Evaluation runs over this window so the decision is based on a **trend**, not an instantaneous reading. ## The evaluation rules (consumer status) Burrow assigns each partition one of several statuses; the group's overall status is the **worst** partition status. The core rules: - **OK** — the consumer is committing and either lag is zero or lag is not consistently growing. - **WARN (lag growing)** — across the whole window, lag is **monotonically increasing**: every sample's lag is higher than the previous. The consumer is committing but falling behind. - **STALL / ERR (stopped)** — the **committed offset is not changing** across the window (no progress) **while lag is nonzero**. The consumer is alive enough to have a window but is not actually advancing — classic stuck-processing or deadlock case. - **STOP / ERR (not committing)** — no new commits have arrived within the expected window time; the consumer may be dead or its commit interval is misconfigured. - **REWIND** — committed offset went backwards (offset reset / replay), surfaced as an anomaly. Note: rules combine **offset movement** and **lag movement**. A consumer that is far behind but steadily catching up (lag decreasing) is **OK**, which a threshold rule would wrongly alarm on. ## Why this beats a single threshold A static rule like "alert if lag > 10000" fails in two opposite ways: - **False positives on high-throughput topics**: a topic doing 200k msg/s may sit at 50k lag and be perfectly healthy (50k is ~0.25s of data). The threshold screams constantly. - **False negatives on low-throughput topics**: a topic doing 1 msg/s stuck at lag 50 has been dead for ~50 seconds, yet 50 < 10000 so no alert fires. Burrow's trend-based logic is **dimensionless**: it asks whether the consumer is *making progress* and *keeping pace*, which is correct regardless of the topic's absolute throughput. This makes it ideal for large multi-tenant clusters where you cannot hand-tune a threshold per group. ## Trade-offs / edge cases - **Window size vs. responsiveness**: a longer window is more stable but slower to react; a short window is jumpier. - **Sparse / bursty commits**: groups that commit rarely give Burrow few samples, making trend evaluation noisier — tune commit interval awareness accordingly. - **Burrow tracks committed lag** (from __consumer_offsets), so like the CLI it sees committed, not in-flight read, position. - **It does not store history** long-term; it is an evaluator, not a time-series DB. Many teams pair Burrow's status with Prometheus (via burrow_exporter) for dashboards and history. - A group with **no movement and zero lag** is OK (idle but caught up) — Burrow correctly does not alarm because lag is zero.
- A consumer is 50,000 messages behind but lag is decreasing every commit. What status does Burrow give and why?OK. Burrow evaluates the trend, not the absolute number; decreasing lag means the consumer is catching up and keeping pace, so it does not alarm even though absolute lag is large.
- How does Burrow distinguish a stalled consumer from one that is merely slow?A stalled consumer's committed offset does not advance across the window while lag stays nonzero (STALL/ERR). A merely slow consumer keeps committing and advancing offsets but its lag grows (WARN). The discriminator is whether the offset is moving at all.
saying these in an interview costs you the question
- Saying Burrow uses a configurable fixed lag threshold for alerting — its whole point is to avoid thresholds.
- Claiming Burrow stores long-term lag history (it evaluates status; pair it with Prometheus for history).
- Forgetting that a large but decreasing lag is OK, not WARN.
- Confusing STALL (offset not moving, lag>0) with WARN (offset moving but lag growing).