skip to content

What do RequestHandlerAvgIdlePercent and NetworkProcessorAvgIdlePercent measure, and how do you interpret low values?

level: seniorimportance: should knowfreq 45%

answer

  1. network threads (num.network.threads) vs I/O threads (num.io.threads)
  2. idle 0.0 = saturated, 1.0 = idle
  3. keep > ~0.2-0.3
  4. handler idle low -> raise num.io.threads / disk
  5. correlate with RequestQueueTimeMs

basics

~20 s

Both are saturation gauges from 0.0 to 1.0. RequestHandlerAvgIdlePercent is the fraction of time the I/O (request handler) threads are idle; NetworkProcessorAvgIdlePercent is the fraction of time the network threads are idle. Low values (near 0) mean those thread pools are saturated and the broker is a bottleneck.

solid answer

~40 s

Kafka brokers split request work across two thread pools. The network threads (num.network.threads) read/write bytes on the socket and place requests on a queue; the request handler / I/O threads (num.io.threads) do the actual work (writing to the log, building fetch responses). NetworkProcessorAvgIdlePercent (kafka.network:type=SocketServer,name=NetworkProcessorAvgIdlePercent) and RequestHandlerAvgIdlePercent (kafka.server:type=KafkaRequestHandlerPool,name=RequestHandlerAvgIdlePercent) report the average fraction of time those pools are idle, each ranging 0.0 (fully saturated) to 1.0 (fully idle). A sustained handler idle near 0 means num.io.threads is the bottleneck (often disk-bound) — increase it or relieve disk pressure. Network idle near 0 points at num.network.threads or raw bandwidth/TLS overhead. As a rule of thumb you want idle comfortably above ~0.2-0.3; drops toward 0 correlate with rising request queue time and latency.

go deeper

for a junior

Knows these are 0-1 idle gauges and low means the broker is busy/saturated.

for a middle

Maps each metric to its thread pool and the config that sizes it (num.network.threads / num.io.threads).

for a senior

Diagnoses which pool is the bottleneck, correlates with RequestQueueTimeMs, and tunes appropriately.

for a principal

Sets capacity-planning baselines and saturation SLOs, balancing thread counts against CPU/disk/NIC and TLS cost.

## The broker request pipeline A Kafka broker processes every produce/fetch/admin request through two distinct thread pools: 1. **Network threads** (count = `num.network.threads`, default 3): they own the sockets. They read incoming request bytes off the wire, hand the request to a shared **request queue**, and later write response bytes back to the client. They also do TLS encryption/decryption work. 2. **Request handler / I/O threads** (count = `num.io.threads`, default 8): they pull requests off the queue and do the real work — appending to the partition log (disk I/O), reading log segments to satisfy fetches, updating ISR, etc. ## The two idle metrics - `kafka.network:type=SocketServer,name=NetworkProcessorAvgIdlePercent` — average fraction of time the **network** threads sit idle waiting for work. - `kafka.server:type=KafkaRequestHandlerPool,name=RequestHandlerAvgIdlePercent` — average fraction of time the **I/O/handler** threads sit idle. Both are **gauges in [0.0, 1.0]**: `1.0` = the pool is completely idle (lots of headroom); `0.0` = the pool is fully saturated (no spare capacity). Note: older Kafka versions reported `RequestHandlerAvgIdlePercent` as a rate that could exceed 1.0 (summed across threads); modern versions normalize to a per-thread fraction. Always confirm the scale on your version. ## Interpreting low values - **RequestHandlerAvgIdlePercent near 0:** the I/O threads are the bottleneck. The handler queue backs up, `RequestQueueTimeMs` rises, end-to-end latency climbs. Causes: too few `num.io.threads`, or the disk can't keep up (so threads block on I/O). Fix by raising `num.io.threads` (rule of thumb: number of disks * something), or by relieving disk pressure (faster disks, fewer partitions per broker, compression). - **NetworkProcessorAvgIdlePercent near 0:** the network layer is the bottleneck. Causes: too few `num.network.threads`, saturated NIC bandwidth, or heavy TLS overhead. Fix by raising `num.network.threads`, scaling out brokers, or offloading TLS. ## Thresholds and correlation A common operational guideline: keep both idle ratios **above ~0.2-0.3**. Anything trending toward 0 for sustained periods is a saturation alert. These metrics are most useful **correlated** with request-latency breakdowns: `kafka.network:type=RequestMetrics,name=RequestQueueTimeMs` (time waiting in queue — rises when handlers are saturated), `LocalTimeMs` (processing time), `ResponseQueueTimeMs`, and `TotalTimeMs`. Falling idle + rising queue time is the textbook saturation signature. ## Edge cases - Increasing thread counts beyond CPU/disk capacity doesn't help and can hurt via contention — these metrics tell you *which* pool to tune, not to blindly raise both. - A broker can show high idle yet high latency if the bottleneck is elsewhere (e.g. a slow downstream replica, GC pauses, or page-cache misses) — idle metrics localize thread-pool saturation, not all latency. - During TLS-heavy workloads network threads can saturate well before raw bandwidth limits.

  • RequestHandlerAvgIdlePercent has dropped to ~0.05 and RequestQueueTimeMs is climbing. What's your diagnosis and remediation?
    The I/O/handler thread pool is saturated, so requests pile up in the queue (hence rising queue time). Either num.io.threads is too low for the load or the disk can't keep up so handlers block on I/O. Remediate by increasing num.io.threads if CPU allows, and/or relieving disk pressure (faster storage, rebalancing partitions off the broker). Confirm by checking disk utilization.
  • Why split work into network threads and I/O threads at all?
    It decouples cheap socket I/O (reading/writing bytes, TLS) from expensive request processing (disk reads/writes). The network threads stay responsive and never block on disk; a bounded queue between them provides backpressure and lets you scale each pool independently for its bottleneck.

saying these in an interview costs you the question

  • Reading the metric backwards — thinking high idle is bad
  • Assuming the scale is always 0-1 (older versions could exceed 1.0 for handler idle)
  • Blindly increasing both thread pools instead of identifying which is saturated
  • Confusing thread-pool saturation with all latency sources (GC, page cache, slow replicas can cause latency at high idle)

context