How do num.io.threads and num.network.threads affect broker throughput, and how would you decide whether to increase one of them?
answer
- network threads = socket I/O + TLS; io threads = KafkaApis work
- RequestHandlerAvgIdlePercent / NetworkProcessorAvgIdlePercent near 0 = bottleneck
- RequestQueueSize + RequestQueueTimeMs -> raise num.io.threads
- high LocalTimeMs = downstream (disk), threads won't help
- tune one pool at a time, re-measure
basics
~20 snum.network.threads sizes the network (processor) pool that does socket I/O; num.io.threads sizes the request-handler pool that does the actual work (log reads/writes). Use the idle-percent metrics: if a pool's average idle percent is low, it is saturated and you raise its thread count.
solid answer
~40 snum.network.threads (default 3) controls Processor threads doing non-blocking socket I/O and request parsing; num.io.threads (default 8) controls request-handler threads that run KafkaApis logic — log append/read, replication, metadata. To decide, watch the idle-percent JMX metrics: NetworkProcessorAvgIdlePercent for processors and RequestHandlerAvgIdlePercent for handlers. A value near 0 means that pool is the bottleneck; near 1 means it is mostly idle. If handler idle is low and network idle is high, raise num.io.threads (often toward the number of disks/cores, e.g. 8–16+). If network idle is low, raise num.network.threads. Over-provisioning wastes context-switching and memory; under-provisioning grows RequestQueueSize and request latency. Also watch the per-request time breakdown (RequestQueueTimeMs, LocalTimeMs) to confirm where time is spent before changing thread counts.
go deeper
Know the two configs exist and that one is for network I/O and one for processing work.
Map each config to its pool and name the idle-percent metric used to size it.
Drive the decision from idle-percent + queue-time/local-time metrics and avoid over-provisioning.
Distinguish thread-starvation from downstream (disk/replication) saturation and set tuning policy, including dynamic reconfiguration, across a fleet.
Kafka's two tunable broker thread pools map directly onto the two halves of the request pipeline, and tuning them is about matching pool size to where work actually queues. **`num.network.threads` (default 3) — Processor pool.** These threads do **only network I/O**: read bytes off sockets, parse into requests, enqueue on the shared request queue, and later write responses. They are non-blocking NIO event loops. They become the bottleneck when there is a lot of *connection/byte* work: many clients, high request rate, TLS encryption/decryption (which happens here), or large response serialization. **`num.io.threads` (default 8) — request-handler (I/O) pool.** These threads dequeue requests and run **`KafkaApis`** logic: appending to the log, reading log segments (page-cache or disk), driving replication, handling metadata. They become the bottleneck when there is a lot of *work per request*: heavy produce/fetch volume, slow disks, compression/decompression, or many partitions. **How to decide — the idle-percent metrics.** Each pool exposes an average idle fraction in `[0,1]`: - **`RequestHandlerAvgIdlePercent`** (kafka.server:type=KafkaRequestHandlerPool) — fraction of time handler threads sit idle waiting for work. - **`NetworkProcessorAvgIdlePercent`** (kafka.network:type=SocketServer) — same for processors. A value near **0.0** means that pool is saturated (the bottleneck); near **1.0** means it has spare capacity. Rule of thumb: keep idle above ~0.3; if it persistently approaches 0, add threads to *that* pool. **Corroborating signals before you change anything:** - **`RequestQueueSize`** climbing toward `queued.max.requests` (500) plus high **`RequestQueueTimeMs`** means requests are waiting for a handler → raise `num.io.threads`. - High **`LocalTimeMs`** (time the handler spends doing the work itself) with low queue time means the bottleneck is downstream (disk/replication), and more threads won't help — fix I/O or partitioning instead. - Low processor idle + high `ResponseQueueTimeMs` points at the network pool → raise `num.network.threads`. **Sizing heuristics.** `num.io.threads` is often raised toward the number of CPU cores or disks (commonly 8–16, higher on big NVMe boxes). `num.network.threads` is raised when many connections/TLS dominate (e.g. 3 → 6–8). Both can be changed dynamically as cluster-wide dynamic configs without a restart in modern Kafka. **Cautions.** Over-provisioning increases context switching, lock contention on shared structures, and memory; it can *reduce* throughput. Always tune one pool at a time, driven by the idle metric for that pool, and re-measure. Adding handler threads cannot overcome a disk- or network-bound broker — those need hardware, better partitioning, or producer batching changes.
- RequestHandlerAvgIdlePercent is 0.05 but disk utilization is already 100%. Should you raise num.io.threads?No. The handlers are blocked on a saturated disk, not starved of threads. More threads add contention without throughput; the fix is faster/more disks, better partitioning, or reducing write amplification.
- Can these be changed without a broker restart?Yes, both num.io.threads and num.network.threads are dynamically updatable cluster-wide/broker-level configs via kafka-configs.sh; the pools are resized live.
saying these in an interview costs you the question
- Always increasing thread counts to fix latency without checking idle-percent/queue metrics.
- Confusing which pool num.network.threads vs num.io.threads sizes.
- Adding I/O threads when LocalTimeMs (disk) is the real bottleneck.