skip to content

In the RecordBatch v2 format, how are record offsets stored, and how does the broker derive each record's absolute offset?

level: middleimportance: should knowfreq 35%

answer

  1. baseOffset in header (absolute, 8 bytes)
  2. record stores offsetDelta (varint)
  3. abs = base + delta
  4. lastOffsetDelta → next offset cheaply
  5. broker rewrites only the base

basics

~10 s

The batch header stores one absolute baseOffset. Each record inside stores only a small offsetDelta. A record's absolute offset = baseOffset + offsetDelta. This avoids repeating the full 8-byte offset on every record.

solid answer

~50 s

RecordBatch v2 stores a single absolute **baseOffset** (8 bytes) in the batch header. Every record inside stores just an **offsetDelta** (a varint, usually 1-2 bytes) relative to that base, so the record's absolute offset is `baseOffset + offsetDelta`. The header also carries **lastOffsetDelta**, which lets the broker compute the batch's last offset (baseOffset + lastOffsetDelta) and therefore the next offset without scanning records. This relative encoding is a major space saving: instead of an 8-byte offset per record, you store one base plus tiny deltas. The same delta technique is used for timestamps (firstTimestamp + timestampDelta). When the broker assigns offsets at append time, it sets baseOffset to the partition's current log-end offset; the deltas were already laid out by the producer (0,1,2,…) and don't change. This is why batches are largely immutable and copyable — the broker only rewrites the base, not every record.

go deeper

for a junior

Know that offsets increase per record and the batch has a starting offset.

for a middle

Explain baseOffset + offsetDelta and why delta encoding saves space.

for a senior

Add lastOffsetDelta, varint encoding, and why this keeps broker append/zero-copy cheap.

for a principal

Discuss immutability of batches, offset re-assignment without decompression, and interaction with compaction/idempotence.

## The problem relative offsets solve Every message in a Kafka partition has a monotonically increasing **offset** (a 64-bit/8-byte long). If each record on disk stored its full absolute offset, a batch of 1000 tiny records would waste ~8 KB just on offsets. RecordBatch v2 avoids this with **relative (delta) encoding**. ## How it is laid out The **batch header** contains: - **baseOffset** (int64, absolute) — the offset of the *first* record in the batch. - **lastOffsetDelta** (int32) — delta of the last record from the base; the batch's last offset is `baseOffset + lastOffsetDelta`. Each **record** inside the batch contains: - **offsetDelta** (varint) — its position relative to baseOffset (0 for the first record, 1 for the second, …). So for any record: `absoluteOffset = baseOffset + offsetDelta`. Because deltas are small integers, they are stored as **varints** (variable-length, typically 1-2 bytes) instead of fixed 8 bytes. A 1000-record batch stores one 8-byte base plus 1000 tiny deltas — a huge saving. ## Why lastOffsetDelta matters The broker can compute the **next offset** to assign (and validate batch contiguity) purely from the header: `nextOffset = baseOffset + lastOffsetDelta + 1`. It does not need to decompress or iterate the records, which keeps the append path cheap and supports zero-copy reads. ## Broker offset assignment When a producer sends a batch, the deltas are already `0..N-1`. The broker sets **baseOffset = current log-end offset** of the partition at append time. The internal deltas are untouched. This is why a batch is essentially immutable: assigning offsets means rewriting one header field, not every record. (One subtlety: with compression, if the broker must re-assign offsets it can do so without decompressing because the base lives in the uncompressed header.) ## Edge cases - **Aborted/compacted records:** offsets can become non-contiguous across batches after log compaction, but within a batch the deltas remain 0..lastOffsetDelta as originally produced. - **Idempotent/transactional producers** rely on this stable structure together with producerId/sequence (a separate field). - **Timestamps use the same trick:** the header stores firstTimestamp and maxTimestamp; each record stores a timestampDelta.

  • How does the broker know the next offset to assign without reading the records?
    From the header: nextOffset = baseOffset + lastOffsetDelta + 1. Both fields live in the uncompressed batch header, so no decompression or per-record scan is needed.
  • Is the same delta technique used for anything besides offsets?
    Yes — timestamps. The header holds firstTimestamp (and maxTimestamp); each record stores a timestampDelta, so each record's timestamp = firstTimestamp + timestampDelta.

saying these in an interview costs you the question

  • Claiming each record stores its full 8-byte absolute offset on disk.
  • Saying the broker must decompress the batch to figure out offsets.
  • Confusing baseOffset (per batch) with the partition's log-start offset (per partition).

context