skip to content

Compare exposing consumer lag to Prometheus via kafka-exporter versus the JMX exporter. When would you use each?

level: seniorimportance: should knowfreq 45%

answer

  1. kafka-exporter: Kafka protocol, committed lag, no live consumer needed
  2. kafka_consumergroup_lag / _lag_sum
  3. JMX exporter: MBeans from a JVM, records-lag-max, only while alive
  4. exporter = durable cluster-wide; JMX = per-instance internals
  5. run both; watch cardinality

basics

~20 s

kafka-exporter talks the Kafka protocol to the brokers and computes committed lag itself (it works without any live consumer). The JMX exporter scrapes MBeans from a JVM, so it surfaces client metrics like records-lag-max but only while that consumer is running. Use kafka-exporter for durable group lag; JMX exporter for in-process metrics.

solid answer

~50 s

kafka-exporter (danielqsj) is a standalone Prometheus exporter that connects to the brokers over the Kafka wire protocol, lists groups, and emits kafka_consumergroup_lag and kafka_consumergroup_lag_sum from committed offsets vs LEO. Because it derives lag from __consumer_offsets and broker metadata, it reports lag even when consumers are dead and needs no JVM access — ideal for cluster-wide, durable group-lag monitoring. The JMX exporter (jmx_exporter) is a Java agent or HTTP scraper that reads MBeans from a specific JVM and translates them to Prometheus; for consumers it surfaces client metrics like records-lag-max, fetch rates, commit rates from consumer-fetch-manager-metrics, but only while that consumer JVM is alive and only for the partitions it owns. So: kafka-exporter = external committed-lag truth across all groups; JMX exporter = rich per-instance client internals. In practice you run both — kafka-exporter for alerting on durable lag, JMX exporter for diagnosing why a specific consumer is slow.

go deeper

for a junior

Know kafka-exporter gives Prometheus the group lag; JMX exporter gives client metrics from a running consumer.

for a middle

Name kafka_consumergroup_lag(_sum) and records-lag-max, and know which survives a consumer crash.

for a senior

Explain the protocol-vs-MBean data sources, the dead-consumer blind spot, and run-both strategy.

for a principal

Design a lag observability stack across exporters weighing cardinality, security/ACLs, staleness, and when to add Burrow for threshold-free evaluation.

## Goal You want consumer lag (and related metrics) in **Prometheus** so you can graph, alert, and correlate. Two common paths exist, with different data sources and blind spots. ## Path A: kafka-exporter (protocol-level external exporter) `kafka-exporter` (commonly the danielqsj implementation) runs as a separate process and **speaks the Kafka client protocol to the brokers**. It: - Lists consumer groups and their committed offsets (effectively reading the same data as `kafka-consumer-groups --describe`, i.e. `__consumer_offsets`). - Fetches partition log-end offsets. - Computes and exposes Prometheus metrics, notably: - **`kafka_consumergroup_lag{consumergroup,topic,partition}`** — per-partition committed lag. - **`kafka_consumergroup_lag_sum{consumergroup,topic}`** — summed per topic. - **`kafka_topic_partition_current_offset`** (LEO) and **`kafka_consumergroup_current_offset`** (committed). Properties: - **Works with no live consumer** — it reads stored committed offsets, so a crashed consumer's lag still grows and stays visible (critical for catching outages). - **Cluster-wide** — sees every group/topic/partition, not just one JVM. - **No access to client internals** — it cannot tell you fetch rate, commit latency, or per-partition read position; only committed lag. - Adds a small, steady load on brokers (listing offsets); cardinality can be large on big clusters (partition-level series). ## Path B: JMX exporter (jmx_exporter on the consumer JVM) The **Prometheus JMX exporter** runs either as a **`-javaagent`** inside the consumer process or as a standalone HTTP scraper against a JMX port. It reads **MBeans** and maps them to Prometheus metrics via a YAML config (whitelist/regex patterns). For a consumer it can expose: - **`records-lag-max`**, `records-lag-avg`, per-partition `records-lag` (from `kafka.consumer:type=consumer-fetch-manager-metrics`). - Fetch throughput (`fetch-rate`, `bytes-consumed-rate`), commit metrics, rebalance metrics, etc. Properties: - **Rich client internals** — exactly the metrics that explain *why* a consumer is slow. - **Only while the JVM is alive** — metrics vanish on crash; you lose lag visibility at the worst time. - **Only the partitions that instance owns**; aggregating a group means scraping every instance. - Requires JVM/JMX access and a metric-mapping config; mind cardinality from per-partition attribute names. ## Choosing | Need | Use | |---|---| | Durable per-group lag, survives consumer crashes | kafka-exporter | | Cluster-wide lag for all groups with one deployment | kafka-exporter | | records-lag-max, fetch/commit rates, per-instance diagnosis | JMX exporter | | Understand *why* a consumer is slow | JMX exporter | **Best practice: run both.** Alert on `kafka_consumergroup_lag_sum` / max from kafka-exporter (it is the durable source of truth), and use JMX-exporter client metrics to drill into a flagged consumer. A third option is **Burrow + burrow_exporter** when you want threshold-free trend evaluation feeding Prometheus. ## Edge cases & gotchas - **Cardinality**: per-partition lag series multiply (groups x topics x partitions). On large clusters, aggregate with recording rules or scrape `*_lag_sum` to control TSDB load. - **Staleness vs. accuracy**: kafka-exporter's lag is as fresh as its scrape interval and the group's commit interval — between commits it shows committed lag, which trails real read position. - **Broker-side MBeans** (request rates, under-replicated partitions) belong to the broker JMX leaf, not here — the JMX exporter on a *broker* is a different deployment than on a *consumer*. - **Security**: kafka-exporter needs ACLs to describe groups and read offsets; JMX exporter needs a secured JMX endpoint (do not expose JMX unauthenticated).

  • Which approach still reports lag after the consumer crashes, and why?
    kafka-exporter, because it derives lag from committed offsets in __consumer_offsets and broker LEO over the Kafka protocol — it needs no live consumer JVM. The JMX exporter would lose the metric since records-lag-max only exists in the running consumer.
  • What metric-cardinality risk does per-partition lag introduce in Prometheus, and how do you mitigate it?
    Series multiply as groups x topics x partitions, which can bloat the TSDB. Mitigate by scraping aggregated kafka_consumergroup_lag_sum, using Prometheus recording rules to pre-aggregate, or dropping per-partition labels you do not need.

saying these in an interview costs you the question

  • Saying kafka-exporter reads MBeans (it speaks the Kafka protocol, not JMX).
  • Claiming the JMX exporter can report lag for a dead consumer.
  • Forgetting kafka-exporter needs no live consumer because it uses committed offsets.
  • Ignoring per-partition metric cardinality on large clusters.

context