skip to content

What is a Kafka record batch, and why does Kafka store messages in batches rather than one record at a time?

level: juniorimportance: must knowfreq 55%

answer

  1. batch = unit of log I/O
  2. shared 61-byte header
  3. compression once per batch
  4. batch.size + linger.ms
  5. enables zero-copy / sendfile

basics

~20 s

A record batch is a group of messages stored together as one unit on disk. Kafka batches records to reduce per-message overhead, improve compression, and write/read more efficiently — fewer, larger I/O operations instead of many tiny ones.

solid answer

~40 s

A record batch (the v2 on-disk format) is the smallest unit Kafka writes to and reads from the log. The producer accumulates records and ships them as a batch; the broker appends that batch to the partition log mostly as-is. Batching amortizes fixed overhead — a 61-byte batch header is shared across all records in the batch, compression is applied once over the whole batch (better ratios than per-record), and network/disk I/O happens in large chunks instead of per record. Producer-side knobs are batch.size (max bytes per batch per partition) and linger.ms (how long to wait to fill a batch). On the consumer side, an entire batch is fetched and decompressed as a unit. This batch-as-the-unit design is also what enables zero-copy transfer of compressed batches straight from page cache to the network.

go deeper

for a junior

Know that a batch is a group of messages stored together and that batching improves throughput/compression.

for a middle

Connect batching to batch.size/linger.ms and to better compression ratios.

for a senior

Explain the v2 header amortization, single-partition rule, and how batching enables zero-copy.

for a principal

Reason about throughput/latency trade-offs of batch tuning and how the batch-as-unit design shapes broker append and consumer fetch paths.

## What a record batch is Kafka stores messages in an append-only **log** per partition. The physical unit written to and read from that log is a **record batch**, not an individual message. A batch is a container: a fixed **batch header** followed by one or more **records** (the actual key/value messages). Since Kafka 0.11 the on-disk format is **RecordBatch v2** (also called the message format v2, `magic` byte = 2). One batch holds N records that a single producer sent together to one partition. ## Why batch instead of storing one record at a time 1. **Amortized overhead.** Each batch carries a ~61-byte header (base offset, length, CRC, producer id/epoch, timestamps, etc.). If every message had its own header, tiny 10-byte messages would be dominated by metadata. Sharing one header across many records slashes overhead per message. 2. **Better compression.** `compression.type` is applied **once over the entire batch payload**, not per record. Compressing many similar records together (e.g. JSON with repeated field names) yields far better ratios than compressing each in isolation. 3. **Fewer, larger I/O operations.** Disks and networks are vastly more efficient with big sequential writes/reads than with many small ones. Batching turns thousands of tiny writes into a few large appends. 4. **Zero-copy.** Because a batch is an opaque, contiguous blob on disk, the broker can stream it from the OS page cache directly to the network socket via `sendfile()` without copying into the JVM heap or decompressing. ## How batches get formed The **producer** accumulates records per partition in an in-memory buffer (the RecordAccumulator) and flushes a batch when either: - the batch reaches **`batch.size`** bytes (default 16 KB), or - **`linger.ms`** elapses (default 0 — send immediately, but >0 lets batches fill). The broker appends the producer's batch to the log largely **as received** (it may re-validate/assign offsets but does not unpack and re-pack each record in the common path). The **consumer** fetches whole batches, then decompresses and iterates the records inside. ## Edge cases - A batch can legitimately contain a **single record** (e.g. low throughput, `linger.ms=0`) — you still pay the header cost but lose batching benefits. - A batch never spans more than one partition. - `max.message.bytes` (broker/topic) and `max.request.size` (producer) bound batch size; an oversized batch is rejected. - Records inside a batch use **relative offsets** and **delta timestamps** off the batch header, which is why the unit matters for the format.

  • Which producer configs control how batches are formed?
    batch.size (max bytes per partition batch, default 16 KB) and linger.ms (how long to wait to fill a batch, default 0). Larger values mean fuller batches, higher throughput, slightly more latency.
  • Can a single batch contain records for two different partitions?
    No. A batch is always for exactly one topic-partition. The producer maintains a separate accumulator per partition.

saying these in an interview costs you the question

  • Saying each Kafka message is stored and compressed individually on disk.
  • Claiming the broker unpacks and recompresses every record on append (it normally appends the batch as-is).
  • Confusing batch.size (bytes) with a record count.

context