What do the network-processor and request-handler idle-percent metrics tell you, and how do you use them to spot thread-pool saturation behind tail-latency spikes?
answer
- idle 0..1, lower = busier
- NetworkProcessorAvgIdlePercent + RequestHandlerAvgIdlePercent
- keep idle > ~0.3
- low idle -> queue time climbs -> p99 up
- blocked threads vs too few threads
basics
~20 sIdle percent is the fraction of time a thread pool sits with no work. NetworkProcessorAvgIdlePercent and RequestHandlerAvgIdlePercent near 0 mean the network or I/O threads are saturated, which causes requests to queue up and tail latency to spike.
solid answer
~40 sKafka exposes two key saturation gauges: kafka.network:type=SocketServer,name=NetworkProcessorAvgIdlePercent (the network/socket threads that read and write bytes) and kafka.server:type=KafkaRequestHandlerPool,name=RequestHandlerAvgIdlePercent (the I/O / request-handler threads that do the actual log work). Each ranges 0.0–1.0; it is the average fraction of time those threads are idle. When either trends toward 0, that pool is the bottleneck: requests pile up in the request or response queue and RequestQueueTimeMs / ResponseQueueTimeMs climb, dragging p99. Rule of thumb: keep idle above ~0.2–0.3. If RequestHandler idle is low, raise num.io.threads or find what blocks I/O threads (GC, disk, locks). If NetworkProcessor idle is low, raise num.network.threads or check network saturation. Crucially, low idle is a *symptom* — confirm with the phase metrics before just adding threads.
go deeper
Know there are two thread pools and that an idle percent near zero means that pool is saturated and latency will suffer.
Name both idle gauges, tie them to num.network.threads / num.io.threads, and connect low idle to rising queue-time phases.
Distinguish under-provisioned pools from blocked threads, use idle as a leading saturation signal alongside phase metrics, and set sensible alert thresholds.
Reason about the pools as queueing systems near saturation, set capacity and autoscaling policy, and teach why tail latency diverges before the mean.
## What 'idle percent' means Kafka's broker runs two fixed-size thread pools: - **Network (processor) threads** — `num.network.threads`. They read request bytes off sockets and write response bytes back. They do *no* heavy work; they shuttle bytes between the socket and the request/response queues. - **Request-handler (I/O) threads** — `num.io.threads`. They pull requests off the request queue and perform the real work: appending to the log, reading from the log/page cache, building responses. For each pool Kafka measures **average idle percent**: the fraction of wall-clock time, averaged across the pool's threads, that they had nothing to do. Values run 0.0 (always busy = saturated) to 1.0 (always idle). - `kafka.network:type=SocketServer,name=NetworkProcessorAvgIdlePercent` - `kafka.server:type=KafkaRequestHandlerPool,name=RequestHandlerAvgIdlePercent` ## Why it predicts tail latency A fixed thread pool with low idle is a queueing system at high utilization. By queueing theory, as utilization approaches 1.0, queue length and wait time blow up *non-linearly* — and the **tail** (p99/p999) blows up before the mean does. So idle percent dropping from 0.5 to 0.1 can take p99 from 5 ms to hundreds of ms even though the mean barely moves. Concretely: if request-handler threads are saturated, requests sit in the request queue longer → **RequestQueueTimeMs** rises. If network threads are saturated, responses can't be written promptly → **ResponseQueueTimeMs / ResponseSendTimeMs** rise. The idle gauge tells you *which pool*; the phase metric confirms the *effect*. ## How to use it 1. Watch both idle gauges per broker. A common alert threshold is idle < 0.3 (warn) / < 0.1 (critical). 2. **RequestHandler idle low** → I/O threads are the bottleneck. Either they're genuinely under-provisioned (raise `num.io.threads`, typically toward the disk count / a multiple of cores) **or** they're *blocked* — by GC pauses, fsync/disk stalls, or lock contention. Adding threads won't help a blocking problem, so check GC and disk first. 3. **NetworkProcessor idle low** → network threads saturated. Raise `num.network.threads`, or investigate NIC/bandwidth saturation and large response sizes. ## Edge cases & gotchas - **Low idle from blocking ≠ low idle from too few threads.** If I/O threads spend their non-idle time *waiting* on disk fsync, they show as busy; adding threads just adds more threads all stuck on the same disk. Correlate with LocalTimeMs and disk metrics. - **It's an average.** One hot thread among many can be masked. Per-broker granularity matters; a single saturated broker can dominate cluster p99. - **Network idle can drop due to TLS/compression CPU**, not just byte volume. - Don't blindly max out thread counts — oversized pools waste CPU on context switching and can worsen latency.
- RequestHandlerAvgIdlePercent is near zero so you bump num.io.threads, but latency doesn't improve. What now?The threads were busy because they were blocked, not under-provisioned. Check LocalTimeMs, GC pause times, and disk fsync/await — I/O threads stalled on disk or GC stay 'busy' no matter how many you add. Fix the underlying stall instead.
- Which phase metric would you expect to rise when RequestHandlerAvgIdlePercent drops toward zero?RequestQueueTimeMs — requests back up in the request queue waiting for a free I/O thread, so their queue-wait time grows and feeds into the p99 spike.
saying these in an interview costs you the question
- Reading idle percent as a percentage where higher means worse (it's the opposite: low idle = saturated)
- Always responding to low RequestHandler idle by adding threads without checking whether threads are blocked on GC/disk
- Ignoring that the metric is a pool average that can hide a single hot thread or a single bad broker
- Confusing the network thread pool with the I/O thread pool