What is buffer.memory, and what happens when a high-volume producer outpaces the brokers?
answer
- buffer.memory = total unsent-record buffer (32MB default)
- full -> send() blocks up to max.block.ms (60s)
- then TimeoutException
- intentional backpressure, prevents OOM
- size > batch.size * active partitions
- monitor buffer-available-bytes, record-queue-time
basics
~10 sbuffer.memory is the total bytes the producer can use to buffer unsent records. When it fills (the producer outpaces brokers), send() blocks up to max.block.ms, then throws TimeoutException. It's the producer's backpressure mechanism.
solid answer
~50 sbuffer.memory (default 33554432, 32MB) is the total memory the producer uses to buffer records waiting to be sent across all partitions, plus compression buffers. The sender thread drains this buffer to brokers. If the application produces faster than the brokers accept (slow brokers, network limits, acks=all latency), the buffer fills. Then send() blocks for up to max.block.ms (default 60s, which also covers metadata fetch time); if space doesn't free up in time, send() throws a TimeoutException. This blocking is intentional backpressure — it stops an unbounded producer from OOMing. To tune for high throughput you raise buffer.memory so bursts can be absorbed, but the real fix for sustained overrun is more broker/partition capacity, since memory only buffers transient spikes. buffer.memory must comfortably exceed batch.size times the number of active partitions so each partition can hold in-flight batches. Monitor buffer-available-bytes and record-queue-time to detect saturation.
go deeper
Know buffer.memory is the producer's send buffer and that a full buffer blocks/throws.
Explain the send() block up to max.block.ms then TimeoutException as backpressure.
Size buffer.memory vs batch.size x partitions and distinguish burst absorption from sustained overrun.
Architect end-to-end backpressure and capacity strategy, choose block vs fail-fast per latency SLA, and wire saturation metrics.
## What buffer.memory is `buffer.memory` (default **33554432 = 32MB**) is the total amount of memory the producer client may use to **buffer records that have not yet been sent** to brokers, including space used for compression. Records from `send()` land in this shared pool (the `RecordAccumulator`), organized into per-partition batches, and a background **sender thread** drains ready batches to brokers. ## The producer/broker speed mismatch The application thread calling `send()` and the sender thread draining to brokers run at independent speeds. If the application produces faster than the cluster can absorb — because of slow/overloaded brokers, network saturation, `acks=all` replication latency, or too few partitions — unsent batches accumulate and the buffer **fills up**. ## What happens when the buffer is full (backpressure) When there is no free space for a new record: 1. `send()` **blocks** the calling thread, waiting for the sender to free space by completing in-flight requests. 2. It waits at most **`max.block.ms`** (default **60000ms**). Note this same timeout also bounds time spent waiting for **topic metadata**. 3. If space is not freed within `max.block.ms`, `send()` throws a **TimeoutException** ('Failed to allocate memory within ... ms'). This blocking is **deliberate backpressure**: it throttles the producer to the cluster's real capacity and prevents the client from buffering unboundedly and running out of heap. ## Tuning for high throughput - **Raise buffer.memory** to absorb larger bursts (e.g., 64-256MB) when traffic is spiky but sustainable on average. - **Sizing rule of thumb**: buffer.memory should comfortably exceed `batch.size * (number of partitions being actively produced to)`, so every active partition can hold at least one full batch (ideally several for pipelining). - A bigger buffer hides spikes but does **not** cure a *sustained* overrun — if you persistently produce faster than brokers accept, the buffer just delays the inevitable block. The real fixes are more partitions, more/faster brokers, lighter acks, or compression to cut bytes. ## Monitoring - `buffer-available-bytes` (JMX) — how much buffer is free; trending to 0 means saturation. - `record-queue-time-avg/max` — how long records wait in the buffer before being sent; rising values signal backpressure. - `bufferpool-wait-ratio` — fraction of time appenders waited for buffer space. ## Edge cases - Setting `max.block.ms=0` makes `send()` fail fast instead of blocking when the buffer is full or metadata is missing. - A very large single record plus a too-small buffer can deadlock allocation; buffer.memory must be at least batch.size. - Blocking in `send()` can stall the calling (often request-handling) thread — important for latency-sensitive apps, which may prefer fail-fast plus a dead-letter path.
- Does increasing buffer.memory fix a producer that is persistently faster than the brokers?No. A larger buffer only absorbs transient bursts; under sustained overrun it fills up anyway and send() still blocks. The durable fix is adding broker/partition capacity, reducing per-record cost (compression, lighter acks), or throttling the producer at the source.
- Which config bounds how long send() blocks when the buffer is full, and what does it do on expiry?max.block.ms (default 60000ms) bounds the block; it also covers metadata wait time. On expiry, send() throws a TimeoutException for that record. Setting it to 0 makes send() fail fast instead of blocking.
saying these in an interview costs you the question
- Saying a full buffer drops records silently (it blocks, then throws)
- Confusing buffer.memory (total client buffer) with batch.size (per-partition batch)
- Claiming bigger buffer.memory cures sustained overrun
- Forgetting max.block.ms also covers metadata wait
- Thinking buffer.memory is per-partition