A producer shows high io-wait-ratio and low record-send-rate while CPU is idle. Walk through how you'd use client metrics to diagnose whether the bottleneck is the application, the producer config, or the broker/network.
answer
- io-wait-ratio high = sender thread idle, client not the bottleneck
- low queue-time + tiny batch-size -> app/linger under-feeding
- high request-latency / throttle-time -> broker/network/quota
- buffer near-zero + high queue-time -> broker drain-bound (acks)
- compare per-node-id MBeans to find one slow broker
basics
~20 sHigh io-wait-ratio with low send-rate means the producer's network thread is mostly idle waiting for data — it's not the bottleneck. Check record-queue-time, batch-size-avg, request-latency-avg, and buffer-available-bytes to see if the app, batching config, or broker is the limit.
solid answer
~50 sio-wait-ratio near 1 means the Sender (network) thread spends most of its time blocked on the selector waiting for work, so the producer client itself is not saturated — the limit is upstream (app produces slowly or batches poorly) or downstream (broker is fast but you're not feeding it). I'd correlate: if `record-queue-time-avg` is low and `batch-size-avg` is small with low `record-send-rate`, the application simply isn't calling send() enough or `linger.ms` is too low to form batches — an app/config issue. If `request-latency-avg`/`request-latency-max` are high, or `produce-throttle-time-avg` is nonzero (broker quota), or `request-rate` is capped while `buffer-available-bytes` is healthy, the broker/network is the constraint. If `buffer-available-bytes` is near zero with high `record-queue-time-avg`, the producer can't drain — broker/acks bound. io-ratio (the inverse) and connection-count round out the picture. Idle CPU plus high io-wait-ratio almost always rules out client-side CPU saturation.
go deeper
Recognize io-wait-ratio relates to how busy the producer's network thread is; high means mostly waiting.
Combine io-wait-ratio with batch-size, queue-time, and request-latency to point at app vs broker.
Build the full decision tree across producer metrics and use throttle-time and per-node breakdowns to localize the bottleneck.
Codify this into runbooks/alerts: which metric combos page, how SLOs gate on throughput+latency, and tuning guidance (linger/batch/in-flight).
**Define the signal.** The Kafka producer runs a dedicated background thread (the Sender) that drains the in-memory RecordAccumulator and sends produce requests via an NIO selector. Two complementary metrics describe how that thread spends time: - `io-ratio` — fraction of time doing actual I/O work. - `io-wait-ratio` — fraction of time *blocked waiting* on the selector for something to do (or for socket readiness). They roughly sum toward 1. A **high io-wait-ratio** means the thread is mostly idle waiting — so the producer machinery is *not* the bottleneck. **Combined with low record-send-rate and idle CPU**, the conclusion is: data isn't arriving fast enough for the producer to do work, OR each round-trip is slow so the thread waits on the broker. Diagnose by partitioning the pipeline: **1. Is the application under-feeding the producer?** - `record-send-rate` low + `batch-size-avg` small + `record-queue-time-avg` low ⇒ few records queued, sent almost immediately ⇒ the app isn't calling `send()` fast enough, or `linger.ms`≈0 so no batching happens. This is an application/config issue, not Kafka. Tuning `linger.ms` up (e.g. 5–20ms) and `batch.size` lets the Sender amortize round-trips and raises throughput, lowering io-wait-ratio. - `records-per-request-avg` low confirms tiny requests. **2. Is producer config throttling itself?** - `buffer-available-bytes` healthy but throughput low rules out buffer exhaustion. - `max.in.flight.requests.per.connection` very low (1, e.g. for ordering) limits pipelining; the thread sends one, waits, sends next — shows as io-wait + low rate even when batches exist. **3. Is the broker/network the bottleneck?** - `request-latency-avg` / `request-latency-max` high ⇒ each produce request is slow to ack (slow disks, `acks=all` with slow followers, replication lag). The Sender waits on the response, inflating io-wait-ratio while throughput stays low. - `produce-throttle-time-avg` (a.k.a. throttle-time) nonzero ⇒ broker is applying a **quota** and deliberately delaying you. - `request-rate` capped while batches are full ⇒ network round-trip bound; check `outgoing-byte-rate` vs link capacity. - `record-queue-time-avg` high *with* full batches and low `buffer-available-bytes` ⇒ records pile up because the broker can't ack fast enough — drain-side (broker) bound. **4. Connection/selector layer.** - `connection-count`, `connection-creation-rate` spiking, or frequent `select-rate` churn can indicate reconnects (e.g. idle connection reaps, broker restarts) that stall progress. **Putting it together (decision tree):** - High io-wait-ratio + low queue-time + small batches → **app/linger config** under-feeding. - High io-wait-ratio + high request-latency / nonzero throttle-time → **broker/network/quota**. - High io-wait-ratio + low buffer-available + high queue-time → **broker drain-bound** (acks/replication). **Edge cases:** A producer that's genuinely idle (low traffic by design) will show high io-wait-ratio harmlessly — context matters; pair the ratio with absolute rates. Setting `metrics.recording.level=DEBUG` exposes more granular per-node metrics to localize a single slow broker via per-`node-id` request-latency. Always compare per-`node-id` MBeans to spot a single hot/slow broker rather than a fleet-wide issue.
- io-wait-ratio is high but that's because the app legitimately produces only a trickle of data. How do you avoid a false alarm?Pair the ratio with absolute rates (record-send-rate, request-rate). A high io-wait-ratio at low traffic is expected and benign — the thread idles because there's little to do. Alert on the ratio only in conjunction with a throughput SLO or rising latency, not in isolation.
- Which single metric would most cleanly confirm the broker is deliberately slowing your producer?produce-throttle-time-avg (the throttle-time metric). A nonzero value means the broker is enforcing a client/user quota and delaying responses — a deliberate broker-side limit, distinct from organic latency.
saying these in an interview costs you the question
- Reading high io-wait-ratio as the producer being overloaded — it means the opposite (mostly idle/waiting).
- Diagnosing throughput without checking batch-size-avg / record-queue-time and linger.ms.
- Ignoring produce-throttle-time-avg and assuming all latency is organic.
- Not drilling into per-node-id metrics to isolate one slow broker.
- Treating a high ratio at intentionally low traffic as a problem.