skip to content

What are the most important built-in JMX metrics for a Kafka producer and consumer, and what does each tell you about client health?

level: juniorimportance: must knowfreq 70%

answer

  1. producer: send-rate, request-latency-avg, buffer-available-bytes
  2. consumer: records-lag-max, fetch-latency-avg, records-consumed-rate
  3. buffer-available-bytes -> backpressure / send blocks
  4. io-wait-ratio = idle vs busy network thread
  5. windowed sampling: metrics.sample.window.ms (30s)

basics

~10 s

Kafka clients expose metrics over JMX. Key producer metrics: record-send-rate, request-latency-avg, buffer-available-bytes. Key consumer metrics: records-consumed-rate, fetch-latency-avg, records-lag-max. They show throughput, latency, and buffering/lag health.

solid answer

~40 s

Kafka clients (KafkaProducer/KafkaConsumer) publish metrics via JMX MBeans under domains like kafka.producer and kafka.consumer. On the producer, record-send-rate is throughput, request-latency-avg is broker round-trip time, buffer-available-bytes shows free RecordAccumulator memory (near-zero means producer.send() will block), and io-wait-ratio shows how much of the network thread's time is spent waiting for data versus working. On the consumer, records-consumed-rate and fetch-latency-avg track throughput/latency, while records-lag-max (per-partition consumer lag) is the single most-watched signal that the consumer is falling behind production. Most metrics live under per-client-id and per-topic/per-node MBeans, so you can drill down. These are the same metrics whether scraped via JMX, Jolokia, or a MetricsReporter.

go deeper

for a junior

Know the names: record-send-rate, request-latency-avg, buffer-available-bytes for producers; records-lag-max for consumers, and that they come via JMX.

for a middle

Map each metric to a failure mode (lag rising, buffer exhaustion, latency spikes) and know rates are windowed.

for a senior

Reason about backpressure from buffer-available-bytes + io-wait-ratio together, and know client-side vs broker-side lag tradeoffs.

for a principal

Define an SLO dashboard and alerting strategy across these metrics and decide where lag/throttle signals should be authoritative.

Kafka's Java clients are heavily instrumented, and by default they expose every metric through JMX (Java Management Extensions), a standard JVM mechanism for exposing manageable attributes as MBeans (managed beans). Each metric becomes an attribute on an MBean whose ObjectName encodes a domain (e.g. `kafka.producer`, `kafka.consumer`), the `client-id`, and sometimes a `topic` or `node-id` so you can view aggregate and per-entity values. **Producer metrics that matter:** - `record-send-rate` — records sent per second per topic; your raw throughput. - `request-latency-avg` / `request-latency-max` — average/max time for a produce request to be acknowledged by the broker; rising latency signals broker or network pressure. - `buffer-available-bytes` — free bytes left in the producer's in-memory buffer (the RecordAccumulator, sized by `buffer.memory`, default 32 MB). If this trends toward zero, `send()` will block for up to `max.block.ms` and then throw `TimeoutException`. A classic backpressure signal. - `record-queue-time-avg` — how long records wait in the accumulator before being sent (affected by `batch.size`/`linger.ms`). - `io-wait-ratio` / `io-ratio` — fraction of the network (Sender) thread's time spent waiting on the selector versus doing I/O work; high io-wait-ratio means the client is idle waiting, low means it's saturated. - `compression-rate-avg`, `record-error-rate`, `record-retry-rate` — encoding efficiency and failure signals. **Consumer metrics that matter:** - `records-consumed-rate` / `bytes-consumed-rate` — consumer throughput. - `fetch-latency-avg` — round-trip time of fetch requests. - `records-lag-max` and `records-lag` (per partition) — how many records behind the log end offset the consumer is; THE health signal for stream processing. `records-lead-min` is the inverse safety margin before data is deleted by retention. - `fetch-rate`, `fetch-size-avg`, `fetch-throttle-time-avg` — fetch behavior and broker-side quota throttling. - `commit-latency-avg`, `commit-rate` — offset commit cost. - Coordinator metrics like `rebalance-rate-per-hour` and `last-rebalance-seconds-ago` reveal group instability. **Common-client metrics** (shared MBeans under `kafka.<client>:type=...-metrics`) include `connection-count`, `connection-creation-rate`, `network-io-rate`, and selector stats. **Edge cases / gotchas:** Metrics are sampled over a configurable window (`metrics.sample.window.ms`, default 30s, across `metrics.num.samples` samples), so rates and averages are windowed, not instantaneous. Per-topic and per-node MBeans only appear once traffic flows to them, so missing MBeans early on are normal. `records-lag-max` reported by the client is computed from fetch responses, so a fully idle consumer can show stale lag — many teams also compute lag broker-side from committed offsets.

  • Your producer's buffer-available-bytes is trending to zero — what happens and why?
    The RecordAccumulator (buffer.memory, default 32MB) is full. Subsequent send() calls block up to max.block.ms then throw TimeoutException. It means the producer can't ship to brokers as fast as the app produces — slow brokers, network, too-large batches, or acks=all waiting.
  • Why might records-lag-max look fine even when consumers are actually behind?
    Client-side lag is derived from fetch responses, so a stalled or idle consumer may report stale lag. It's safer to compute lag broker-side from the difference between the log-end-offset and the committed offset (e.g. via kafka-consumer-groups or Burrow).

saying these in an interview costs you the question

  • Claiming Kafka clients only expose metrics through a paid agent — they expose JMX out of the box.
  • Saying request-latency-avg measures end-to-end delivery time — it's only the produce request round-trip to the broker.
  • Treating metric rates as instantaneous — they're windowed over metrics.sample.window.ms.
  • Confusing buffer-available-bytes (producer accumulator) with OS/socket buffers.

context