skip to content

Why does a writer group many records into one batch before sending them to a broker node?

level: juniorimportance: must knowfreq 72%

answer

  1. the cost is per request, not per byte
  2. one round trip carries many records
  3. fixed overhead amortised over the group
  4. throughput bought with a little delay
  5. the writer sets it, the cluster pays for it

basics

~20 s

Because a large part of what a request costs a node is fixed, whatever the request carries: a round trip, a queue slot, a handler, per-request bookkeeping. A batch spreads that fixed cost over many records, so the same hardware accepts far more records per second.

solid answer

~50 s

A broker node's work is not spread evenly over bytes. Each request costs it a network round trip, a slot in its bounded request queue, a handler from a finite pool, a routing decision and a response — and that cost is nearly the same for one record as for a thousand. A **batch** is the group of records the writer puts into one request, so the fixed cost is paid once instead of once per record. Where a cluster is saturated by request count rather than by bytes, this is the single cheapest throughput lever available, and it is also what gives a compressor a useful unit to work on. The price is that a record may wait in the writer for company, so its own end-to-end delay rises, and that the writer holds unsent records in memory.

go deeper

for a junior

Recall that a request costs a node a fixed amount whatever it carries, so putting many records in one request raises throughput. Also recall the price: the record that starts a batch waits for the rest before anything is sent.

for a middle

Explain the fixed costs by name — round trip, queue slot, handler, accounting — and say why a cluster saturated by request count benefits far more than one saturated by bytes. Note that the return flattens as batches grow.

for a senior

Show that you can read the ratio of request rate to byte rate per caller and spot the writer sending tiny requests. Then talk about what changes when a node must unpack the group rather than keep it whole.

for a principal

Frame it as an economics question: grouping is the cheapest headroom on a shared cluster, it is configured by people who do not pay the cluster's costs, and a platform that wants it has to make efficient defaults the easy path for new writers.

## The cost a node pays per request, not per record A broker node's work does not scale smoothly with the bytes it receives. Every request that arrives carries an overhead that is roughly the same whether it holds one record or a thousand: - a **network round trip**, with its own latency, from the writer to the node and back; - a slot in the node's bounded **request queue**, the place accepted-but-unprocessed work waits; - a **handler** from a finite pool of threads that do the real request work; - a routing decision — which stream, and which part of it, these records belong to; - per-request accounting: metrics, authorization checks, and whatever the node records about the caller; - a response written back to the client. On a cluster carrying small records, that fixed overhead — not the byte volume — is usually what saturates first. Two writers can push identical data volumes and cost the cluster wildly different amounts, purely because one sends a thousand requests a second and the other sends ten. ## What a batch is, and what it buys A **batch** is simply the group of records a writer accumulates in memory and places into a single request. Nothing about the records changes: the same content, in the same order, with the same destination. Only the packaging changes. What that packaging buys: 1. **Amortised fixed cost.** One round trip, one queue slot, one handler, one accounting entry for many records. This is the whole of the throughput argument. 2. **Fewer stalls on network latency.** A writer that waits for an answer before sending again is limited by round-trip time; carrying more per round trip removes that ceiling without touching the network. 3. **A compressible unit.** The batch is what a writer compresses, which is a separate lever with its own trade — but it only exists because the records were grouped in the first place. 4. **Steadier load.** A node handling fewer, larger requests spends less time switching between pieces of work. ## What it costs - **Delay on the records that arrive first.** A record handed to the writer early sits and waits for company. Batching raises throughput by spending latency, and the record that started the batch pays the most of it. - **Memory on the writer's host.** Unsent records live in the sending application's memory until the batch goes. - **Coarser retry granularity.** A failed request takes its whole batch with it, so a retry re-sends every record's bytes, not just the one that had trouble. - **Coarser attribution.** A single slow request now represents many records, which makes a per-record timing story harder to read. ## Where platforms differ Grouping is universal; what happens to the group after it arrives is not. | Design | What the batch is once it arrives | |---|---| | A store that keeps the stream as the writer framed it | The group survives as a stored unit and is often handed to readers the same way, so the saving reaches storage and the read path too | | A broker that unpacks on arrival and tracks each message separately | The batch was a transport optimisation only; per-record work at the node still scales with the number of records | | A node obliged to inspect each record | The round trip is saved, the per-record work is not — this is where batching disappoints people | This is why a blanket claim like *bigger batches always help* is wrong: they always help the request count, and they help the rest only where the node treats the group as a unit. ## What an operator can see, and what they can do An operator can measure it. The ratio of request rate to byte rate, per caller, tells you the average number of records per request without needing anything from the writer's team: a caller with a high request rate and a modest byte rate is sending tiny requests, and is buying a disproportionate share of the cluster's fixed costs. What an operator cannot do is change it. Batch size, the willingness to wait for a batch to fill, and the compression choice are all set in the writer's own configuration. The cluster side of this conversation is evidence and a request, not a setting — which is why an interviewer who asks *why batch* is usually working towards *and what would you do about a writer that does not*.

  • If batching is that cheap, why not make every batch enormous?
    Because the return flattens. Once the batch is large enough that the fixed per-request cost is negligible against the bytes, further growth buys almost nothing, while the delay the first record waits, the memory the writer holds, and the amount re-sent on a failed request all keep rising. The lever has a knee, and past it you are paying latency for nothing.
  • A stream carries only a handful of records a second. What does batching buy there?
    Very little. There is nothing to group: batches stay small however long the writer is willing to wait, so the request count barely falls and the main effect is added delay on every record. Batching pays where record rate is high and records are small; on a low-rate stream the honest answer is that this lever is not the one to reach for.

A courier charges for the stop, not for the envelope. Sending fifty envelopes one at a time buys fifty stops; putting them in one sack buys one. The sack is cheaper per envelope — and the envelope that went in first waited the longest for the van to leave.

saying these in an interview costs you the question

  • Says batching makes each individual record arrive sooner.
  • Thinks a node does the same amount of work per record either way.
  • Believes a bigger batch is always better, with no cost at all.
  • Assumes an operator can set a writer's batch size from the cluster.
  • Confuses grouping records for transport with a delivery guarantee.
  • Thinks grouping records by itself reduces the bytes that are stored.