skip to content

Offset and Time Indexes

The per-segment offset and time indexes that let a broker jump straight to a record instead of scanning the log. Interviewers use it to check whether you understand how lookups by offset or timestamp stay cheap with almost no memory.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What are the .index and .timeindex files that Kafka maintains alongside each log segment, and what does each one map?

level: juniorimportance: must knowfreq 60%

answer

  1. .index = offset -> byte position
  2. .timeindex = timestamp -> offset
  3. one per segment, base-offset named
  4. sparse: every ~index.interval.bytes
  5. binary search + short scan

basics

~20 s

Each Kafka log segment has a .index file mapping a message offset to its physical byte position in the .log file, and a .timeindex file mapping a timestamp to an offset. They let Kafka jump near a record without scanning the whole segment.

solid answer

~40 s

A Kafka partition is split into segments; each segment is a base-offset-named set of files: the .log (actual records), a .index, and a .timeindex. The .index maps a relative offset -> physical byte position in the .log so a consumer fetch at offset N can seek close to it instead of scanning from the start. The .timeindex maps a timestamp -> offset, enabling offset-by-timestamp lookups (e.g. consumer.offsetsForTimes, or time-based retention/rolling). Both are sparse: Kafka adds an entry only after roughly index.interval.bytes (default 4096) of new log data, not per record, so they stay small and mostly memory-resident. Lookups are a binary search over the sorted entries to find the nearest preceding entry, then a short linear scan of the .log.

go deeper

for a junior

Know there are two per-segment index files and what each maps; know they make lookups fast.

for a middle

Explain sparseness, relative offsets, and the binary-search-then-scan lookup.

for a senior

Tie indexes to offset/time fetches, retention, and segment rolling; discuss memory-mapping and sizing.

for a principal

Reason about index density vs lookup cost tradeoffs and operational impact at scale (page cache, recovery).

## Background A Kafka topic-partition is an append-only log. To keep it manageable, Kafka does not store one giant file — it breaks the partition's log into **segments**. Each segment is named by its **base offset** (the offset of its first record), e.g. `00000000000000368769.log`. Alongside that `.log` file, the same segment has companion index files with the same base-offset name: `00000000000000368769.index` and `00000000000000368769.timeindex` (and a `.txnindex` if transactions are used). ## What each file maps - **.index (offset index):** maps a **relative offset** (offset minus the segment base offset, stored as a 4-byte int) to a **physical byte position** within the `.log` file. So given a target offset, Kafka can find roughly where in the segment file that record lives and `seek()` there, instead of reading the segment from byte 0. - **.timeindex (time index):** maps a **timestamp** (8-byte long) to a **relative offset**. This answers "what is the first offset at or after time T?" — used by `KafkaConsumer.offsetsForTimes()`, by time-based retention, and by deciding when to roll a segment. ## Why they exist Without an index, finding offset N or a timestamp would require scanning records. The indexes turn that into a fast binary search plus a tiny linear scan. ## Sparse, not dense Kafka does NOT write an index entry for every record. It writes one roughly every `index.interval.bytes` of appended log (default 4 KB). That is why they are called **sparse** indexes — they trade a little extra scan work for far smaller index files that can stay in the OS page cache / memory-mapped. ## Edge cases - A lookup almost never lands exactly on an indexed offset; it finds the **largest entry whose offset/timestamp is <= the target**, then linearly scans forward in the `.log` for the exact record. - The active (newest) segment's index files are pre-allocated to `segment.index.bytes` (default 10 MB) and memory-mapped; they are trimmed to actual size when the segment is rolled and closed.

  • Why does the .index store a relative offset instead of the absolute offset?
    Relative offset (offset - base offset) fits in 4 bytes, halving each entry's size versus an 8-byte absolute offset, and the absolute value is recoverable by adding the segment base offset.
  • If a record isn't directly listed in the .index, how does Kafka find it?
    Binary search finds the index entry with the largest offset <= target, Kafka seeks to that byte position in the .log, then scans forward record-by-record until it reaches the exact offset.

saying these in an interview costs you the question

  • Saying the .index has an entry for every message (it is sparse, ~every index.interval.bytes).
  • Confusing the two: .index is offset->position, .timeindex is timestamp->offset.
  • Claiming the index stores the record data (it only stores positions/offsets; data is in .log).

context

open as a page

What is index.interval.bytes, and how does it govern the density of Kafka's sparse indexes? What are the tradeoffs of changing it?

level: middleimportance: should knowfreq 40%

basics

~20 s

index.interval.bytes (default 4096) is how many bytes of log Kafka appends before adding a new index entry. Smaller means denser indexes and faster lookups but larger files; larger means sparser indexes, smaller files, slightly slower lookups.

open as a page

Walk through exactly how Kafka resolves a fetch request for a specific offset using the sparse .index file.

level: seniorimportance: should knowfreq 35%

basics

~20 s

Kafka first picks the right segment (the one whose base offset is the largest <= target). It binary-searches that segment's .index for the entry with the largest offset <= target, seeks to that byte position in the .log, then scans record batches forward until it reaches the requested offset.

open as a page

How does Kafka answer an offset-by-timestamp query (e.g. consumer.offsetsForTimes), and what role does the .timeindex play and what are its caveats?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The consumer sends a ListOffsets request with the target timestamp. The broker uses each segment's .timeindex (timestamp -> offset) to binary-search for the first offset whose timestamp is >= the target, then refines via the .index. It returns that offset and its timestamp.

open as a page

What is the .txnindex file, what does it record, and how do consumers use it to honor read_committed isolation?

level: principalimportance: nice to knowfreq 18%

basics

~20 s

The .txnindex is a per-segment file listing aborted transactions as ranges (producer id, first offset, last stable offset). With read_committed isolation, the broker sends this list so consumers can filter out records from aborted transactions and not read past the last stable offset.

open as a page