skip to content

Which broker-side metric tells you the end-to-end time a Kafka request took, and where do you find it?

level: juniorimportance: must knowfreq 60%

answer

  1. TotalTimeMs = end-to-end
  2. kafka.network:type=RequestMetrics
  3. request=<ApiKey> (Produce/Fetch/...)
  4. histogram → read 99thPercentile / Max, not just Mean
  5. JMX-exposed

basics

~10 s

TotalTimeMs, exposed via JMX under kafka.network:type=RequestMetrics, name=TotalTimeMs, request=<ApiKey> (like Produce or Fetch). It's a histogram, so you read Mean, 99thPercentile, and Max.

solid answer

~40 s

The broker reports per-request timing through JMX under the MBean kafka.network:type=RequestMetrics, with a name attribute for each phase and a request attribute for the API (Produce, Fetch, FetchConsumer, FetchFollower, Metadata, etc.). The end-to-end time is name=TotalTimeMs. Because it's a histogram, you don't read a single number — you look at Mean, 99thPercentile, and Max to understand the distribution and tail latency. To localize a slow request you then read the per-phase metrics under the same MBean (RequestQueueTimeMs, LocalTimeMs, RemoteTimeMs, ResponseQueueTimeMs, ResponseSendTimeMs), since they sum to TotalTimeMs. The metrics are per ApiKey, so you can ask 'are produces slow or fetches slow?' separately.

go deeper

for a junior

Name TotalTimeMs and that it lives under kafka.network RequestMetrics in JMX, read as a histogram.

for a middle

Know it's per ApiKey and decomposes into the per-phase metrics that sum to it.

for a senior

Drive percentile-based dashboards/alerts and pivot to per-phase decomposition.

for a principal

Standardize JMX scraping + per-ApiKey p99 SLOs across the fleet.

## The metric Kafka brokers publish request timing over **JMX** (Java Management Extensions), the standard way Java apps expose runtime metrics. The relevant MBean is: ``` kafka.network:type=RequestMetrics,name=TotalTimeMs,request=<ApiKey> ``` - `type=RequestMetrics` — the family of per-request timing metrics. - `name=TotalTimeMs` — the specific metric: end-to-end time for the request. - `request=<ApiKey>` — which API: `Produce`, `Fetch`, `FetchConsumer`, `FetchFollower`, `Metadata`, `ApiVersions`, etc. Each API has its own set of timings. ## It's a histogram, not a single value TotalTimeMs is a **histogram** (a distribution of many samples), so the MBean exposes attributes like `Mean`, `50thPercentile`, `95thPercentile`, `99thPercentile`, `999thPercentile`, `Max`, and `Count`. For latency you care most about the **tail** — the 99thPercentile and Max — because the mean hides slow outliers that hurt real users. ## How to read it Monitoring tools (Prometheus JMX exporter, Datadog, jconsole, jmxterm) scrape these MBeans. A typical dashboard plots `TotalTimeMs.99thPercentile` per ApiKey over time. ## Next step after TotalTimeMs TotalTimeMs only tells you *that* a request was slow, not *why*. Under the same `RequestMetrics` MBean there are per-phase metrics — `RequestQueueTimeMs`, `LocalTimeMs`, `RemoteTimeMs`, `ThrottleTimeMs`, `ResponseQueueTimeMs`, `ResponseSendTimeMs` — that sum to TotalTimeMs. You compare them to find the dominant phase. ## Edge cases - Different ApiKeys behave very differently (a FetchConsumer's high TotalTimeMs is often benign long-polling; a Produce's is not). - A near-zero Mean with a high Max means occasional spikes — always check percentiles, not just the average.

  • Why look at the 99th percentile of TotalTimeMs instead of the mean?
    The mean hides outliers. Tail latency (p99, Max) is what real users and SLAs feel; a healthy mean can coexist with severe spikes that the average masks. Latency problems live in the tail.

saying these in an interview costs you the question

  • Thinking TotalTimeMs is a single instantaneous value rather than a histogram with percentiles.
  • Forgetting the metric is per ApiKey (request=) so Produce and Fetch are separate.
  • Reading only the Mean and ignoring p99/Max.

context