skip to content

What does the batch-level CRC protect in RecordBatch v2, and how does that format design enable broker-side zero-copy (sendfile) of compressed data to consumers?

level: seniorimportance: should knowfreq 30%

answer

  1. one CRC-32C per batch (hardware-accelerated)
  2. covers body after CRC; excludes baseOffset/length/leaderEpoch
  3. exclusion → re-stamp offsets without recompute
  4. opaque self-contained blob → sendfile/transferTo
  5. broker no decompress; consumer verifies CRC + decompresses

basics

~20 s

Each batch has one CRC-32C checksum covering the batch body (records + most of the header) to detect corruption. Because the whole compressed batch is a self-contained, opaque blob, the broker can stream it straight from the OS page cache to the network with sendfile — no decompression or JVM-heap copy.

solid answer

~50 s

RecordBatch v2 stores a single **CRC-32C** (Castagnoli) checksum per batch. It covers everything **after** the CRC field — the rest of the header attributes plus all records — so it detects on-disk or in-flight corruption of the batch as a unit (the leading baseOffset/length/partitionLeaderEpoch/CRC bytes are excluded so the broker can update offsets/leader epoch without recomputing the CRC). The broker verifies the CRC on append and can on fetch. Crucially, the batch is **self-describing and opaque**: compression, CRC, offsets, and timestamps are all encoded so a batch can be moved without unpacking it. That's what enables **zero-copy**: on a consumer fetch the broker uses `FileChannel.transferTo()` → the OS `sendfile()` syscall to copy bytes directly from the **page cache** to the **socket**, never landing in the JVM heap and never decompressing. Decompression and CRC verification happen on the **consumer**. This is a core reason Kafka sustains high throughput — the hot read path does almost no per-record CPU work on the broker.

go deeper

for a junior

Know there's a checksum per batch for corruption detection, and Kafka streams data efficiently.

for a middle

Explain CRC-32C per batch and the basic idea of sendfile/zero-copy from page cache.

for a senior

Detail CRC scope/exclusions, transferTo/sendfile, and that the consumer decompresses + verifies.

for a principal

Reason about when zero-copy breaks (TLS, down-conversion), throughput implications, and format-version/compatibility trade-offs.

## The CRC A **CRC (cyclic redundancy check)** is a checksum that detects accidental corruption (bad disk sectors, flipped network bits). RecordBatch v2 stores **one CRC per batch**, using the **CRC-32C (Castagnoli)** polynomial — chosen because modern CPUs have a hardware instruction (`crc32` / SSE4.2) for it, making it cheap. **What it covers:** the CRC sits early in the header, and it checksums **all bytes after the CRC field** — the remaining header attributes (timestampType, compression bits, producerId/epoch/baseSequence, etc.) **and** the entire (possibly compressed) records section. **What it deliberately excludes:** the bytes *before and including* the CRC — **baseOffset**, **batchLength**, **partitionLeaderEpoch**, and the CRC itself. This is intentional: the broker assigns/updates the baseOffset and partitionLeaderEpoch at append/leadership time, and excluding them means it can do so **without recomputing the CRC**. (This is a key difference from the older v0/v1 format where the per-message CRC made re-stamping expensive.) The broker validates the CRC when a batch is produced (to reject corrupt input); the consumer validates it after fetching to detect corruption in transit or on disk. ## Why the format enables zero-copy **Zero-copy** means moving file bytes to a socket without copying them through the application (JVM) address space. The traditional path is: disk → page cache → JVM heap → socket buffer → NIC (4 copies, 2 user/kernel context switches). With **`sendfile()`** (exposed in Java as `FileChannel.transferTo()`), it's: disk/page cache → socket → NIC, staying entirely in the kernel. Kafka can use this **only because a record batch is an opaque, contiguous, self-contained blob**: - It is stored on disk **exactly** in the wire format the consumer expects. - It carries its **own CRC, compression codec, offsets (base + deltas), and timestamps** inside the header, so nothing about it needs to be recomputed to send it. - The broker therefore does **not** decompress, does **not** deserialize records, and does **not** copy into the heap on the read path. It just hands the byte range from the page cache to `sendfile()`. **Consequence:** **decompression and CRC verification are the consumer's job**, not the broker's. The broker's fetch path is almost pure I/O. This — combined with sequential disk writes and reliance on the OS page cache — is a primary reason a single broker can serve very high read throughput. ## Edge cases / caveats - **Zero-copy is bypassed** when the broker must touch the bytes: TLS/SSL encryption (`sendfile` can't encrypt), down-conversion to an older message format for old clients (`message.format.version` mismatch), or some quota/throttling paths. Keeping clients on v2 and avoiding format down-conversion preserves zero-copy. - **CRC scope confusion:** candidates often think the CRC covers the whole batch including baseOffset — it does not, precisely so offsets/leader-epoch can change cheaply. - **Per-batch, not per-record:** v2 has one CRC per batch; the old v0/v1 had a CRC per message, which was costlier and re-stamped on offset assignment.

  • Why does the CRC deliberately exclude baseOffset and partitionLeaderEpoch?
    Because the broker assigns/updates those at append and on leadership changes. Excluding them lets the broker rewrite the offset and leader epoch without recomputing the batch CRC, keeping the append/replication path cheap.
  • Name two situations where Kafka cannot use zero-copy on the fetch path.
    TLS/SSL encryption (sendfile can't encrypt the bytes) and message-format down-conversion for old consumers (the broker must rewrite batches to an older magic, forcing a heap copy). Some throttling paths also bypass it.
  • Who decompresses a compressed batch — broker or consumer?
    The consumer. The broker stores and streams the compressed batch as-is (enabling zero-copy); decompression and CRC verification happen client-side.

saying these in an interview costs you the question

  • Saying the broker decompresses batches to serve consumers (it streams them compressed).
  • Claiming the CRC covers baseOffset/batchLength (those are excluded on purpose).
  • Thinking v2 has a CRC per record (it's one per batch).
  • Asserting zero-copy always applies even with TLS or format down-conversion.

context