skip to content

What does queued.max.requests control, and what happens when the request queue fills up? Explain the backpressure mechanism.

level: seniorimportance: should knowfreq 40%

answer

  1. queued.max.requests default 500 = request queue capacity
  2. full queue -> processors stop reading sockets
  3. TCP receive buffers fill, window shrinks -> clients slow
  4. RequestQueueSize at cap + RequestQueueTimeMs high
  5. raising it buffers more, doesn't add capacity

basics

~20 s

queued.max.requests (default 500) is the capacity of the shared request queue between network threads and request-handler threads. When it's full, network threads stop reading new requests off sockets, which slows clients down — that is the backpressure.

solid answer

~40 s

queued.max.requests (default 500) bounds the RequestChannel.requestQueue, the single ArrayBlockingQueue where Processor (network) threads place parsed requests for the request-handler (I/O) threads to consume. If handlers can't keep up, the queue fills. When full, Processors mute / stop reading from their sockets (they don't enqueue more), so request bytes back up in OS TCP receive buffers; the TCP window shrinks and producers/consumers experience slower acks and eventually block or back off. This is intentional backpressure: it prevents unbounded memory growth and protects the broker from being overwhelmed. You observe it via the RequestQueueSize metric (sitting at the cap) and rising RequestQueueTimeMs. The right response is usually to add num.io.threads or fix the downstream bottleneck, not to blindly raise queued.max.requests (which just buffers more and raises latency/memory).

go deeper

for a junior

It's the size of the queue between network and worker threads; when full, the broker slows down intake.

for a middle

Explain that processors stop reading sockets and clients are flow-controlled via TCP.

for a senior

Tie it to RequestQueueSize/RequestQueueTimeMs and argue why raising it isn't a real fix.

for a principal

Position it within an end-to-end backpressure/flow-control strategy and contrast with quotas and queued.max.request.bytes.

`queued.max.requests` is the size knob for the buffer that decouples the two broker thread pools, and it is the broker's primary in-process backpressure mechanism. **Where the queue sits.** Processor (network) threads parse incoming bytes into requests and put them on a **single shared `requestQueue`** inside `RequestChannel`. Request-handler (I/O) threads (`num.io.threads`) take requests off it and run `KafkaApis`. The queue is a bounded `ArrayBlockingQueue` whose capacity is **`queued.max.requests` (default 500)**. **Decoupling + bounding.** A bounded queue is deliberate: it absorbs short bursts (so handlers stay busy even if arrivals are spiky) while capping the maximum number of in-flight, undispatched requests so memory can't grow without limit. **What happens when it's full.** When handlers fall behind, the queue reaches its cap. A Processor that wants to enqueue a freshly-parsed request can't. The Processor then **stops reading / mutes the channels** it owns — it does not accept more request bytes. Consequences cascade outward: 1. Bytes pile up in the **OS TCP receive buffer** for those connections. 2. The TCP **receive window** advertised to the client shrinks toward zero. 3. The client's send is **flow-controlled**: producer sends and consumer fetches slow down or block; in-flight request limits (`max.in.flight.requests.per.connection`) are hit; eventually clients back off and may time out. This chain is **backpressure**: instead of crashing or OOMing, the broker pushes the pressure back to clients, who naturally slow their offered load. **Note on throttling vs this queue.** Quota-based throttling (client/user quotas) is a *separate* mechanism that delays responses; it is not the same as request-queue backpressure, though both reduce effective client rate. **Observability.** - **`RequestQueueSize`** pinned at `queued.max.requests` ⇒ the queue is the choke point. - High **`RequestQueueTimeMs`** ⇒ requests wait a long time before a handler picks them up. - Low **`RequestHandlerAvgIdlePercent`** confirms handlers are saturated. **Tuning guidance.** Raising `queued.max.requests` does **not** add processing capacity — it just lets more requests wait, increasing latency and memory. The real fixes are: increase `num.io.threads` if handlers are merely thread-starved; or address the downstream bottleneck (disk, replication, GC) if `LocalTimeMs` is high. Lowering it makes the broker apply backpressure sooner (fail-fast, lower latency tails) at the cost of less burst tolerance. There is also `queued.max.request.bytes` to bound by memory rather than count.

  • Is raising queued.max.requests a good fix when the queue is constantly full?
    Usually no. It only lets more requests wait, raising latency and memory use. Fix the cause: add num.io.threads if handlers are thread-starved, or resolve the downstream disk/replication bottleneck.
  • How is request-queue backpressure different from quota throttling?
    Backpressure is implicit and global to a broker: processors stop reading when the bounded queue fills. Quota throttling is per-client/user: the broker deliberately delays the response to keep a client within its configured byte/request rate.

saying these in an interview costs you the question

  • Saying a full queue drops/rejects requests (it doesn't drop; it stops reading and applies TCP backpressure).
  • Treating queued.max.requests as a throughput dial.
  • Confusing it with consumer-side fetch queues or quota throttling.

context