skip to content

How are segment files named, and what does the filename tell you about the records inside?

level: middleimportance: should knowfreq 45%

answer

  1. filename = base offset
  2. 20 digits, zero-padded
  3. lexical sort == numeric sort
  4. floor lookup to find segment
  5. .index stores relative offsets (4 bytes)

basics

~20 s

A segment's files are named after its base offset — the offset of the first record in that segment — zero-padded to 20 digits, e.g. 00000000000000006000.log. So the filename tells you the lowest offset that segment can contain.

solid answer

~40 s

Each segment's files share a base name equal to the segment's base offset: the offset of the first record it holds, formatted as a 20-digit zero-padded decimal. So `00000000000000000000.log` starts at offset 0, and the next segment might be `00000000000000006000.log` if the previous one ended at offset 5999. The .index and .timeindex companions reuse the same base. Because base offsets are strictly increasing across segments and offsets are monotonic within a partition, Kafka can binary-search the directory listing to find which segment contains a target offset, then use the per-segment .index to jump to the byte position. The 20-digit padding exists purely so lexical sort order matches numeric order. Note .index stores offsets RELATIVE to the base offset (so they fit in 4 bytes), and the absolute offset is base + relative.

go deeper

for a junior

Know the filename equals the first offset in the segment.

for a middle

Explain 20-digit padding, floor lookup to locate a segment, and that base offsets strictly increase.

for a senior

Connect filename base offset to relative offsets in .index and the seek path.

for a principal

Discuss how naming + monotonic offsets enable O(log n) segment lookup and tooling/recovery built on it.

## Base offset = the segment's identity When Kafka rolls a new segment, the first record written into it has some offset N. That N becomes the segment's **base offset**, and it is used as the **filename** for every file in that segment, padded with leading zeros to **20 decimal digits**: ``` 00000000000000000000.log 00000000000000000000.index 00000000000000000000.timeindex ``` The next segment begins at whatever offset comes after the previous segment's last record. If the first segment held offsets 0..5999, the next segment's base offset is 6000: ``` 00000000000000006000.log 00000000000000006000.index 00000000000000006000.timeindex ``` ## Why 20 digits, zero-padded Offsets are 64-bit; the max value 9,223,372,036,854,775,807 is 19 digits, so 20 digits guarantees room. Zero-padding makes the **lexical (string) sort order equal the numeric order**, so a plain directory listing is already sorted by offset — convenient for binary search. ## What the name tells you The filename is the **lowest offset** the segment can contain. To find the segment holding offset X, Kafka takes the segments whose base offset is <= X and picks the one with the largest such base offset (a floor lookup, doable by binary search since base offsets increase). It then consults that segment's **.index** file. ## Relative offsets inside the index The .index does not store absolute offsets. It stores **relative offsets** = absolute_offset - base_offset, which fit in 4 bytes, paired with a 4-byte physical byte position into the .log. The absolute offset is reconstructed as base_offset + relative. This is exactly why the base offset is baked into the filename: it's the anchor for every relative offset in that segment. ## Edge cases - An empty freshly rolled active segment exists with a base offset but no records yet; reads simply return nothing past the log end. - Compacted topics keep the same base-offset naming; cleaned segments may have gaps in offsets but the base-offset rule still holds. - Tooling like `kafka-dump-log.sh --files 00000000000000006000.log` reads a single segment by name.

  • Why does the .index store relative rather than absolute offsets?
    Relative offsets (absolute minus base) fit in 4 bytes paired with a 4-byte file position, halving index entry size versus 8-byte absolute offsets. The base offset from the filename is added back to recover the absolute offset.

saying these in an interview costs you the question

  • Saying the filename is the LAST offset in the segment — it's the FIRST (base) offset.
  • Thinking the padding is cosmetic — it makes lexical order match numeric order for binary search.
  • Claiming the .index stores absolute offsets — it stores relative-to-base offsets.

context