How does Kafka answer an offset-by-timestamp query (e.g. consumer.offsetsForTimes), and what role does the .timeindex play and what are its caveats?
answer
- offsetsForTimes -> ListOffsets RPC
- .timeindex: timestamp -> offset, then .index for position
- find FIRST offset with ts >= target
- CreateTime can be out-of-order -> approximate
- LogAppendTime monotonic; -1 ts not found
basics
~20 sThe consumer sends a ListOffsets request with the target timestamp. The broker uses each segment's .timeindex (timestamp -> offset) to binary-search for the first offset whose timestamp is >= the target, then refines via the .index. It returns that offset and its timestamp.
solid answer
~40 soffsetsForTimes triggers a ListOffsets RPC carrying a per-partition target timestamp. The broker locates the segment whose largest timestamp covers the target (segments track maxTimestamp), then binary-searches that segment's .timeindex — a sorted array of (timestamp, relativeOffset) entries — for the floor entry <= target, then scans forward in the log to find the FIRST record with timestamp >= target, returning its offset and timestamp. Caveats: (1) the timestamp used is governed by message.timestamp.type — CreateTime (producer-set, can be out of order) vs LogAppendTime (broker-set, monotonic). With CreateTime, out-of-order timestamps can make results approximate. (2) The .timeindex is sparse like the offset index. (3) If no record has timestamp >= target, the partition returns null offset. (4) Records produced before message timestamps existed, or with -1 timestamps, won't be found by time.
go deeper
Know offsetsForTimes uses the .timeindex to map a time to an offset.
Describe binary search of the time index plus the hop through the offset index.
Explain CreateTime vs LogAppendTime effects and the 'first offset >= target' semantics.
Reason about correctness guarantees under clock skew and design implications for time-travel/replay.
## The query path `KafkaConsumer.offsetsForTimes(Map<TopicPartition, Long timestamp>)` asks: *for each partition, what is the earliest offset whose record timestamp is >= the given epoch-millis timestamp?* The client issues a **ListOffsets** request (with the special timestamp value being the target time, distinct from the sentinels -1 = LATEST and -2 = EARLIEST). The **leader broker** answers. ## Role of the .timeindex Each segment maintains, in memory and on disk, its **maxTimestamp** (largest timestamp seen). The broker walks segments to find the first one whose maxTimestamp >= target — that segment may contain the answer. It then **binary-searches that segment's `.timeindex`**, a sorted array of **(8-byte timestamp, 4-byte relative offset)** entries, for the floor entry whose timestamp <= target. That entry yields a starting offset; the broker resolves that offset's byte position via the `.index`, seeks into the `.log`, and **scans forward** to find the FIRST record whose timestamp >= target. It returns `{offset, timestamp}`. ## Why sparse + two-index hop The time index is sparse (entries added with the offset index, every ~`index.interval.bytes`), so a binary search plus short scan suffices. Because the time index stores an offset (not a byte position), Kafka does a second hop through the offset `.index` to get the physical position. ## Critical caveat: timestamp semantics `message.timestamp.type` decides which timestamp lives in records: - **CreateTime** (default): the **producer** sets it. Clock skew or late/out-of-order producers mean timestamps are NOT guaranteed monotonic across offsets. The time index stores the **max** timestamp per entry, and lookups assume non-decreasing timestamps — so with badly out-of-order CreateTime, `offsetsForTimes` results are approximate (it finds *an* offset near the time, not provably the global first). - **LogAppendTime**: the **broker** stamps append time, which is monotonic, giving precise time lookups. ## Other edge cases - Records with timestamp **-1** (no timestamp, e.g. very old clients) are not time-indexed and won't be found by time. - If the target is **after all records**, the partition returns a **null** offset (no record at/after that time). - Time-based **retention** (`retention.ms`) and segment rolling (`segment.ms`) also consult the time index / maxTimestamp.
- Why can offsetsForTimes give approximate results with the default CreateTime?CreateTime is producer-assigned, so timestamps aren't guaranteed to increase with offset (clock skew, late events). The time index assumes non-decreasing timestamps, so the returned offset may not be the provable global-first record at/after the target.
- After the .timeindex gives an offset, why does Kafka still consult the .index?The time index maps timestamp->offset, not to a byte position. Kafka resolves that offset to a physical position via the offset .index before seeking into the .log to scan for the exact record.
saying these in an interview costs you the question
- Saying the .timeindex maps timestamp directly to a byte position (it maps to an offset).
- Assuming time lookups are always exact (CreateTime out-of-order makes them approximate).
- Forgetting that records with -1 timestamps aren't time-indexed.
- Confusing the LATEST/EARLIEST sentinels with a real timestamp query.