skip to content

Record Batches and Message Format

The on-disk record batch format: batch headers, relative offsets, producer id/epoch/sequence, CRC, timestamps, and per-batch compression. Comes up when discussing compression, idempotence, or why batching dominates Kafka's throughput.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is a Kafka record batch, and why does Kafka store messages in batches rather than one record at a time?

level: juniorimportance: must knowfreq 55%

answer

  1. batch = unit of log I/O
  2. shared 61-byte header
  3. compression once per batch
  4. batch.size + linger.ms
  5. enables zero-copy / sendfile

basics

~20 s

A record batch is a group of messages stored together as one unit on disk. Kafka batches records to reduce per-message overhead, improve compression, and write/read more efficiently — fewer, larger I/O operations instead of many tiny ones.

solid answer

~40 s

A record batch (the v2 on-disk format) is the smallest unit Kafka writes to and reads from the log. The producer accumulates records and ships them as a batch; the broker appends that batch to the partition log mostly as-is. Batching amortizes fixed overhead — a 61-byte batch header is shared across all records in the batch, compression is applied once over the whole batch (better ratios than per-record), and network/disk I/O happens in large chunks instead of per record. Producer-side knobs are batch.size (max bytes per batch per partition) and linger.ms (how long to wait to fill a batch). On the consumer side, an entire batch is fetched and decompressed as a unit. This batch-as-the-unit design is also what enables zero-copy transfer of compressed batches straight from page cache to the network.

go deeper

for a junior

Know that a batch is a group of messages stored together and that batching improves throughput/compression.

for a middle

Connect batching to batch.size/linger.ms and to better compression ratios.

for a senior

Explain the v2 header amortization, single-partition rule, and how batching enables zero-copy.

for a principal

Reason about throughput/latency trade-offs of batch tuning and how the batch-as-unit design shapes broker append and consumer fetch paths.

## What a record batch is Kafka stores messages in an append-only **log** per partition. The physical unit written to and read from that log is a **record batch**, not an individual message. A batch is a container: a fixed **batch header** followed by one or more **records** (the actual key/value messages). Since Kafka 0.11 the on-disk format is **RecordBatch v2** (also called the message format v2, `magic` byte = 2). One batch holds N records that a single producer sent together to one partition. ## Why batch instead of storing one record at a time 1. **Amortized overhead.** Each batch carries a ~61-byte header (base offset, length, CRC, producer id/epoch, timestamps, etc.). If every message had its own header, tiny 10-byte messages would be dominated by metadata. Sharing one header across many records slashes overhead per message. 2. **Better compression.** `compression.type` is applied **once over the entire batch payload**, not per record. Compressing many similar records together (e.g. JSON with repeated field names) yields far better ratios than compressing each in isolation. 3. **Fewer, larger I/O operations.** Disks and networks are vastly more efficient with big sequential writes/reads than with many small ones. Batching turns thousands of tiny writes into a few large appends. 4. **Zero-copy.** Because a batch is an opaque, contiguous blob on disk, the broker can stream it from the OS page cache directly to the network socket via `sendfile()` without copying into the JVM heap or decompressing. ## How batches get formed The **producer** accumulates records per partition in an in-memory buffer (the RecordAccumulator) and flushes a batch when either: - the batch reaches **`batch.size`** bytes (default 16 KB), or - **`linger.ms`** elapses (default 0 — send immediately, but >0 lets batches fill). The broker appends the producer's batch to the log largely **as received** (it may re-validate/assign offsets but does not unpack and re-pack each record in the common path). The **consumer** fetches whole batches, then decompresses and iterates the records inside. ## Edge cases - A batch can legitimately contain a **single record** (e.g. low throughput, `linger.ms=0`) — you still pay the header cost but lose batching benefits. - A batch never spans more than one partition. - `max.message.bytes` (broker/topic) and `max.request.size` (producer) bound batch size; an oversized batch is rejected. - Records inside a batch use **relative offsets** and **delta timestamps** off the batch header, which is why the unit matters for the format.

  • Which producer configs control how batches are formed?
    batch.size (max bytes per partition batch, default 16 KB) and linger.ms (how long to wait to fill a batch, default 0). Larger values mean fuller batches, higher throughput, slightly more latency.
  • Can a single batch contain records for two different partitions?
    No. A batch is always for exactly one topic-partition. The producer maintains a separate accumulator per partition.

saying these in an interview costs you the question

  • Saying each Kafka message is stored and compressed individually on disk.
  • Claiming the broker unpacks and recompresses every record on append (it normally appends the batch as-is).
  • Confusing batch.size (bytes) with a record count.

context

open as a page

Explain CreateTime vs LogAppendTime in Kafka. How does message.timestamp.type affect what is stored in a record batch?

level: middleimportance: must knowfreq 45%

basics

~10 s

CreateTime is the timestamp the producer set when the message was created. LogAppendTime is when the broker appended it. The topic config message.timestamp.type (CreateTime or LogAppendTime) decides which one Kafka keeps.

open as a page

How do the producerId, producerEpoch, and baseSequence fields in a record batch enable idempotent (exactly-once-into-the-log) production, and how does the broker use them to detect duplicates?

level: seniorimportance: must knowfreq 40%

basics

~20 s

Each batch header carries a producerId (PID), producerEpoch, and baseSequence number. The broker tracks the last sequence it accepted per (PID, partition). If a retried batch repeats sequences, the broker drops it as a duplicate, so retries don't create duplicate records.

open as a page

In the RecordBatch v2 format, how are record offsets stored, and how does the broker derive each record's absolute offset?

level: middleimportance: should knowfreq 35%

basics

~10 s

The batch header stores one absolute baseOffset. Each record inside stores only a small offsetDelta. A record's absolute offset = baseOffset + offsetDelta. This avoids repeating the full 8-byte offset on every record.

open as a page

How does compression.type work at the record-batch level, and what happens when producer, topic, and broker compression settings differ?

level: seniorimportance: should knowfreq 28%

basics

~20 s

compression.type sets the codec (none, gzip, snappy, lz4, zstd) applied to the whole batch's record section, recorded in the batch header. The producer normally compresses; the broker keeps it as-is unless the topic's compression.type forces a different codec, which makes the broker recompress.

open as a page

What does the batch-level CRC protect in RecordBatch v2, and how does that format design enable broker-side zero-copy (sendfile) of compressed data to consumers?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Each batch has one CRC-32C checksum covering the batch body (records + most of the header) to detect corruption. Because the whole compressed batch is a self-contained, opaque blob, the broker can stream it straight from the OS page cache to the network with sendfile — no decompression or JVM-heap copy.

open as a page