Why does compressing a batch of many small log records beat compressing and sending each record on its own?
answer
- fixed costs per message, not per record
- the compressor starts with empty history
- one envelope amortised over many records
- a tiny payload can grow when compressed
- batch size trades bytes for latency
basics
~20 sTwo fixed costs are paid once per message rather than once per record: the envelope around the message and the compressed stream's own framing. A compressor also has no earlier bytes to reference until it has seen some.
solid answer
~40 sPer-message overhead does not scale down with the record. A transport envelope carries its own fields, and a compressed member carries fixed framing — a gzip member, for example, spends ten bytes of header and eight of trailer regardless of payload. Send a 60-byte record alone and that framing rivals the data; a short payload also gives the compressor nothing to reference, so the output can come out larger than the input. Batch a thousand records and both fixed costs divide by a thousand, while every record after the first finds its field names already in the compressor's history. The bytes are bought with delay until the batch fills, a larger unit of loss and retry, and memory held per producer — so batch size is set with a flush deadline, not left unbounded.
go deeper
Know that many tiny messages cost more than the same records sent together, because each message carries its own fixed overhead regardless of how little it holds.
Separate the two effects when explaining it: fixed framing amortised across records, and a compressor that can only reference byte runs it has already seen in this stream.
Set the batch size from a measured curve and bound it with a flush deadline, then state what was traded away: added delay, a larger retry unit, and buffer memory on every producer.
The batch boundary is also the failure and ordering unit, so it belongs in the delivery-guarantee conversation rather than being tuned purely against the byte bill.
## Two fixed costs per message Every message on a link carries overhead that does not shrink when the payload does: - **The envelope.** Whatever wraps the payload — routing fields, a length, an identifier, a timestamp, headers — is charged once per message. For a request-shaped hop the envelope is frequently larger than a single log record. - **The compressed member's framing.** A compressed stream is not only its compressed bytes. A gzip member spends a **ten-byte header and an eight-byte trailer** (a checksum and the uncompressed length), eighteen bytes before any data, plus the framing of the coded block itself. Suppose those two together come to 58 bytes for a given hop. Then the framing cost per record depends only on how many records share the message: | records per message | fixed framing per record | |---|---| | 1 | 58 bytes | | 10 | 5.8 bytes | | 100 | 0.58 bytes | | 1,000 | 0.058 bytes | Against an 89-byte record, that first row is a 65% surcharge; the last is noise. Nothing about the records changed — only how many of them share one envelope. ## What the compressor's history buys The second effect is larger and less obvious. A compressor replaces a run of bytes with a reference to an **earlier occurrence in the same stream**. At the start of a stream there is no history, so the first record compresses badly. Record two repeats record one's field names, punctuation and often several of its values, and from there on the skeleton of each record is nearly free. Compressing each record separately throws that away every time: each message restarts the history at empty. Worse, a very short payload can come out **larger than it went in**, because the fixed framing exceeds whatever the coder saves on a payload with no repeats. Encoders usually detect this and fall back to storing the bytes uncompressed — which avoids expansion of the data but still leaves the framing on the bill. ## What batching costs Batching is not free, and the costs are not measured in bytes: - **Latency.** A record waits until the batch fills. Without a **flush deadline**, a low-traffic producer can hold a record for a very long time. - **Blast radius.** The batch is the unit of loss, retry and duplication. One rejected message now discards or replays a thousand records instead of one. - **Memory.** Every producer holds a buffer; multiply by the fleet, and by the number of destinations each producer writes to. - **Head-of-line effects.** A large batch occupies the link and the consumer for longer, and a poison record inside it can stall the whole batch. ## Choosing a batch size 1. **Bound it on both axes** — a maximum record count or byte size, and a maximum time before flush, whichever comes first. The time bound is what protects the low-traffic producer. 2. **Measure the curve rather than guessing.** Compressed bytes per record against batch size flattens quickly: once the fixed framing per record is well under a byte and every repeated name has occurred, additional size buys almost nothing. 3. **Pick the knee, not the maximum.** Past the point where the curve flattens, you are trading real delay and a wider failure unit for negligible bytes. 4. **Re-check after any change to the record shape**, since the amount of cross-record repetition is what the compressor was feeding on. One boundary: the processor cost of compressing, and where in the pipeline it is best spent, is a separate subject from the byte accounting here. So is the risk side of an expansion ratio, which belongs with decode-path resource limits rather than with the bill.
- Can compressing a single small record make it larger?Yes. A compressed member carries fixed framing — around eighteen bytes for a gzip member — and a short payload offers almost nothing to reference, so the output can exceed the input. Encoders typically detect this and fall back to storing the bytes uncompressed, which stops the data expanding but still leaves the framing to pay for.
- What limits the gain from making batches ever larger?Two things flatten the curve. The fixed framing per record is already a fraction of a byte after a few hundred records, and the compressor's history saturates once every repeated name and common value has occurred at least once. Beyond that knee you are buying negligible bytes with real delay, more memory and a wider failure unit.
- Does batching help an uncompressed link at all?Yes, but only through the first effect. The envelope and any per-message framing still amortise across the records, which on short records is a substantial share. What disappears is the larger effect — there is no compressor history to share — so the gain is bounded by the envelope size rather than by the records' repetition.
saying these in an interview costs you the question
- Thinks the compression ratio is independent of how much input it sees
- Believes compressing a 60-byte record must make it smaller
- Counts payload bytes and ignores the per-message envelope
- Batches with no deadline, so quiet producers hold records indefinitely
- Assumes larger batches keep paying, with no diminishing return
- Forgets the batch becomes the unit of loss and retry