skip to content

Client Metrics Instrumentation and Reporters

Client-side metrics that reveal producer and consumer health, plus the reporters that ship them to Prometheus or OpenTelemetry. Interviewers ask because broker metrics alone never explain a slow application.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What are the most important built-in JMX metrics for a Kafka producer and consumer, and what does each tell you about client health?

level: juniorimportance: must knowfreq 70%

answer

  1. producer: send-rate, request-latency-avg, buffer-available-bytes
  2. consumer: records-lag-max, fetch-latency-avg, records-consumed-rate
  3. buffer-available-bytes -> backpressure / send blocks
  4. io-wait-ratio = idle vs busy network thread
  5. windowed sampling: metrics.sample.window.ms (30s)

basics

~10 s

Kafka clients expose metrics over JMX. Key producer metrics: record-send-rate, request-latency-avg, buffer-available-bytes. Key consumer metrics: records-consumed-rate, fetch-latency-avg, records-lag-max. They show throughput, latency, and buffering/lag health.

solid answer

~40 s

Kafka clients (KafkaProducer/KafkaConsumer) publish metrics via JMX MBeans under domains like kafka.producer and kafka.consumer. On the producer, record-send-rate is throughput, request-latency-avg is broker round-trip time, buffer-available-bytes shows free RecordAccumulator memory (near-zero means producer.send() will block), and io-wait-ratio shows how much of the network thread's time is spent waiting for data versus working. On the consumer, records-consumed-rate and fetch-latency-avg track throughput/latency, while records-lag-max (per-partition consumer lag) is the single most-watched signal that the consumer is falling behind production. Most metrics live under per-client-id and per-topic/per-node MBeans, so you can drill down. These are the same metrics whether scraped via JMX, Jolokia, or a MetricsReporter.

go deeper

for a junior

Know the names: record-send-rate, request-latency-avg, buffer-available-bytes for producers; records-lag-max for consumers, and that they come via JMX.

for a middle

Map each metric to a failure mode (lag rising, buffer exhaustion, latency spikes) and know rates are windowed.

for a senior

Reason about backpressure from buffer-available-bytes + io-wait-ratio together, and know client-side vs broker-side lag tradeoffs.

for a principal

Define an SLO dashboard and alerting strategy across these metrics and decide where lag/throttle signals should be authoritative.

Kafka's Java clients are heavily instrumented, and by default they expose every metric through JMX (Java Management Extensions), a standard JVM mechanism for exposing manageable attributes as MBeans (managed beans). Each metric becomes an attribute on an MBean whose ObjectName encodes a domain (e.g. `kafka.producer`, `kafka.consumer`), the `client-id`, and sometimes a `topic` or `node-id` so you can view aggregate and per-entity values. **Producer metrics that matter:** - `record-send-rate` — records sent per second per topic; your raw throughput. - `request-latency-avg` / `request-latency-max` — average/max time for a produce request to be acknowledged by the broker; rising latency signals broker or network pressure. - `buffer-available-bytes` — free bytes left in the producer's in-memory buffer (the RecordAccumulator, sized by `buffer.memory`, default 32 MB). If this trends toward zero, `send()` will block for up to `max.block.ms` and then throw `TimeoutException`. A classic backpressure signal. - `record-queue-time-avg` — how long records wait in the accumulator before being sent (affected by `batch.size`/`linger.ms`). - `io-wait-ratio` / `io-ratio` — fraction of the network (Sender) thread's time spent waiting on the selector versus doing I/O work; high io-wait-ratio means the client is idle waiting, low means it's saturated. - `compression-rate-avg`, `record-error-rate`, `record-retry-rate` — encoding efficiency and failure signals. **Consumer metrics that matter:** - `records-consumed-rate` / `bytes-consumed-rate` — consumer throughput. - `fetch-latency-avg` — round-trip time of fetch requests. - `records-lag-max` and `records-lag` (per partition) — how many records behind the log end offset the consumer is; THE health signal for stream processing. `records-lead-min` is the inverse safety margin before data is deleted by retention. - `fetch-rate`, `fetch-size-avg`, `fetch-throttle-time-avg` — fetch behavior and broker-side quota throttling. - `commit-latency-avg`, `commit-rate` — offset commit cost. - Coordinator metrics like `rebalance-rate-per-hour` and `last-rebalance-seconds-ago` reveal group instability. **Common-client metrics** (shared MBeans under `kafka.<client>:type=...-metrics`) include `connection-count`, `connection-creation-rate`, `network-io-rate`, and selector stats. **Edge cases / gotchas:** Metrics are sampled over a configurable window (`metrics.sample.window.ms`, default 30s, across `metrics.num.samples` samples), so rates and averages are windowed, not instantaneous. Per-topic and per-node MBeans only appear once traffic flows to them, so missing MBeans early on are normal. `records-lag-max` reported by the client is computed from fetch responses, so a fully idle consumer can show stale lag — many teams also compute lag broker-side from committed offsets.

  • Your producer's buffer-available-bytes is trending to zero — what happens and why?
    The RecordAccumulator (buffer.memory, default 32MB) is full. Subsequent send() calls block up to max.block.ms then throw TimeoutException. It means the producer can't ship to brokers as fast as the app produces — slow brokers, network, too-large batches, or acks=all waiting.
  • Why might records-lag-max look fine even when consumers are actually behind?
    Client-side lag is derived from fetch responses, so a stalled or idle consumer may report stale lag. It's safer to compute lag broker-side from the difference between the log-end-offset and the committed offset (e.g. via kafka-consumer-groups or Burrow).

saying these in an interview costs you the question

  • Claiming Kafka clients only expose metrics through a paid agent — they expose JMX out of the box.
  • Saying request-latency-avg measures end-to-end delivery time — it's only the produce request round-trip to the broker.
  • Treating metric rates as instantaneous — they're windowed over metrics.sample.window.ms.
  • Confusing buffer-available-bytes (producer accumulator) with OS/socket buffers.

context

open as a page

How do you integrate Kafka client metrics into a Micrometer/OpenTelemetry-based observability stack (e.g. a Spring Boot service exporting to Prometheus)?

level: middleimportance: should knowfreq 40%

basics

~20 s

Bind Kafka's client metrics into Micrometer using KafkaClientMetrics (or Spring Boot's auto-binding for KafkaTemplate/listener containers). Micrometer then exports them to your backend (Prometheus, OTel). It reads the client's metrics() map and registers them as Micrometer gauges.

open as a page

How does the MetricsReporter SPI work, and how would you use it to ship Kafka client metrics to an external system?

level: middleimportance: should knowfreq 45%

basics

~20 s

MetricsReporter is a pluggable interface (org.apache.kafka.common.metrics.MetricsReporter). You implement it, register it via the metric.reporters client config, and Kafka calls your code as metrics are created/changed/removed so you can forward them (e.g. to Prometheus or a custom sink).

open as a page

Explain KIP-714 client telemetry push: what problem it solves, how the protocol works, and how it relates to the MetricsReporter SPI.

level: seniorimportance: should knowfreq 35%

basics

~20 s

KIP-714 lets brokers collect standardized client metrics by having clients push them to the broker over the Kafka protocol, instead of operators scraping each client's JMX. The broker subscribes clients to metrics, clients push them as OpenTelemetry-encoded payloads on an interval.

open as a page

A producer shows high io-wait-ratio and low record-send-rate while CPU is idle. Walk through how you'd use client metrics to diagnose whether the bottleneck is the application, the producer config, or the broker/network.

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

High io-wait-ratio with low send-rate means the producer's network thread is mostly idle waiting for data — it's not the bottleneck. Check record-queue-time, batch-size-avg, request-latency-avg, and buffer-available-bytes to see if the app, batching config, or broker is the limit.

open as a page