In the RecordBatch v2 format, how are record offsets stored, and how does the broker derive each record's absolute offset?
answer
- baseOffset in header (absolute, 8 bytes)
- record stores offsetDelta (varint)
- abs = base + delta
- lastOffsetDelta → next offset cheaply
- broker rewrites only the base
basics
~10 sThe batch header stores one absolute baseOffset. Each record inside stores only a small offsetDelta. A record's absolute offset = baseOffset + offsetDelta. This avoids repeating the full 8-byte offset on every record.
solid answer
~50 sRecordBatch v2 stores a single absolute **baseOffset** (8 bytes) in the batch header. Every record inside stores just an **offsetDelta** (a varint, usually 1-2 bytes) relative to that base, so the record's absolute offset is `baseOffset + offsetDelta`. The header also carries **lastOffsetDelta**, which lets the broker compute the batch's last offset (baseOffset + lastOffsetDelta) and therefore the next offset without scanning records. This relative encoding is a major space saving: instead of an 8-byte offset per record, you store one base plus tiny deltas. The same delta technique is used for timestamps (firstTimestamp + timestampDelta). When the broker assigns offsets at append time, it sets baseOffset to the partition's current log-end offset; the deltas were already laid out by the producer (0,1,2,…) and don't change. This is why batches are largely immutable and copyable — the broker only rewrites the base, not every record.
go deeper
Know that offsets increase per record and the batch has a starting offset.
Explain baseOffset + offsetDelta and why delta encoding saves space.
Add lastOffsetDelta, varint encoding, and why this keeps broker append/zero-copy cheap.
Discuss immutability of batches, offset re-assignment without decompression, and interaction with compaction/idempotence.
## The problem relative offsets solve Every message in a Kafka partition has a monotonically increasing **offset** (a 64-bit/8-byte long). If each record on disk stored its full absolute offset, a batch of 1000 tiny records would waste ~8 KB just on offsets. RecordBatch v2 avoids this with **relative (delta) encoding**. ## How it is laid out The **batch header** contains: - **baseOffset** (int64, absolute) — the offset of the *first* record in the batch. - **lastOffsetDelta** (int32) — delta of the last record from the base; the batch's last offset is `baseOffset + lastOffsetDelta`. Each **record** inside the batch contains: - **offsetDelta** (varint) — its position relative to baseOffset (0 for the first record, 1 for the second, …). So for any record: `absoluteOffset = baseOffset + offsetDelta`. Because deltas are small integers, they are stored as **varints** (variable-length, typically 1-2 bytes) instead of fixed 8 bytes. A 1000-record batch stores one 8-byte base plus 1000 tiny deltas — a huge saving. ## Why lastOffsetDelta matters The broker can compute the **next offset** to assign (and validate batch contiguity) purely from the header: `nextOffset = baseOffset + lastOffsetDelta + 1`. It does not need to decompress or iterate the records, which keeps the append path cheap and supports zero-copy reads. ## Broker offset assignment When a producer sends a batch, the deltas are already `0..N-1`. The broker sets **baseOffset = current log-end offset** of the partition at append time. The internal deltas are untouched. This is why a batch is essentially immutable: assigning offsets means rewriting one header field, not every record. (One subtlety: with compression, if the broker must re-assign offsets it can do so without decompressing because the base lives in the uncompressed header.) ## Edge cases - **Aborted/compacted records:** offsets can become non-contiguous across batches after log compaction, but within a batch the deltas remain 0..lastOffsetDelta as originally produced. - **Idempotent/transactional producers** rely on this stable structure together with producerId/sequence (a separate field). - **Timestamps use the same trick:** the header stores firstTimestamp and maxTimestamp; each record stores a timestampDelta.
- How does the broker know the next offset to assign without reading the records?From the header: nextOffset = baseOffset + lastOffsetDelta + 1. Both fields live in the uncompressed batch header, so no decompression or per-record scan is needed.
- Is the same delta technique used for anything besides offsets?Yes — timestamps. The header holds firstTimestamp (and maxTimestamp); each record stores a timestampDelta, so each record's timestamp = firstTimestamp + timestampDelta.
saying these in an interview costs you the question
- Claiming each record stores its full 8-byte absolute offset on disk.
- Saying the broker must decompress the batch to figure out offsets.
- Confusing baseOffset (per batch) with the partition's log-start offset (per partition).