skip to content

How does Kafka resolve an offset from a timestamp (offsetsForTimes), and what are its limitations?

level: seniorimportance: should knowfreq 45%

answer

  1. offsetsForTimes -> first offset with ts >= X
  2. backed by .timeindex (sparse), O(log n)
  3. null if time after last record
  4. CreateTime vs LogAppendTime matters
  5. per partition, seek() after

basics

~20 s

Kafka can look up, per partition, the offset of the earliest record whose timestamp is >= a given time, using the consumer's offsetsForTimes (ListOffsets API) and the per-segment .timeindex files. It returns null if no record has a timestamp at or after that time.

solid answer

~50 s

`KafkaConsumer.offsetsForTimes(Map<TopicPartition, Long>)` resolves, **per partition**, the offset of the **first record whose timestamp ≥ the requested epoch-millis**, returning an `OffsetAndTimestamp` (or `null` if no such record exists, e.g. the time is after the last record). It uses the **ListOffsets** broker API, backed by each segment's `.timeindex` file (sparse timestamp->offset entries), so the lookup is roughly O(log n) per segment. A critical subtlety: the timestamp it searches is the record's **message timestamp**, governed by the topic's `message.timestamp.type` — `CreateTime` (producer-set) or `LogAppendTime` (broker-set at append). With `CreateTime`, timestamps can be out of order (producers, clock skew, reprocessing), so the result is 'first offset with ts ≥ X' by index order, not a perfectly time-sorted guarantee. You then `seek()` to the returned offsets. Common uses: bootstrap a consumer from 'one hour ago', or replay from a point in time.

go deeper

for a junior

Know Kafka can find an offset from a timestamp so you can start reading from a point in time.

for a middle

Use offsetsForTimes + seek correctly, including the null-when-after-last-record case.

for a senior

Explain the .timeindex mechanism, O(log n) lookup, and CreateTime vs LogAppendTime correctness implications.

for a principal

Design time-travel/replay tooling accounting for clock skew, sparse index granularity, retention, and per-partition resolution.

## The problem Sometimes you want to start consuming from a **point in time** ('replay everything since 09:00') rather than an offset. Offsets are partition-local integers with no inherent time meaning, so Kafka provides a timestamp->offset lookup. ## The API `consumer.offsetsForTimes(Map<TopicPartition, Long> timestampsToSearch)` takes, per partition, an **epoch-millisecond** timestamp and returns a `Map<TopicPartition, OffsetAndTimestamp>` where each value is: - the **offset of the earliest record whose timestamp is ≥ the requested timestamp**, plus that record's actual timestamp, or - **`null`** if no record in that partition has a timestamp at or after the requested time (e.g. you asked for a time newer than the last record). You then typically call `consumer.seek(tp, offsetAndTimestamp.offset())` to position the consumer. The CLI equivalent is `kafka-get-offsets.sh --time <ms>` (or the older `kafka-run-class GetOffsetShell`), and special sentinel values `-1` (latest/LEO) and `-2` (earliest/log-start) are accepted. ## How it works under the hood The broker serves this via the **ListOffsets** request. Each log **segment** has a `.timeindex` file alongside its `.index`: a **sparse** mapping from a record timestamp to the offset/byte position of a record. Kafka: 1. Finds the segment that could contain the timestamp (segments store their max timestamp). 2. Binary-searches the `.timeindex` to a nearby entry, then scans forward to the first qualifying record. This makes the lookup approximately **O(log n)** rather than a full scan. ## The timestamp it actually uses The searched timestamp is the **record's message timestamp**, whose meaning depends on the topic config `message.timestamp.type`: - **`CreateTime`** (default): the timestamp the **producer** set (usually wall-clock at send). These can be **non-monotonic** across the log — late producers, clock skew, batching, or reprocessing can interleave timestamps. So 'first offset with ts ≥ X' is by the index's ordering, and the time semantics are only as good as producer clocks. - **`LogAppendTime`**: the **broker** stamps the time at append, so timestamps are monotonic with offsets and the lookup is clean and reliable. ## Limitations and gotchas - **Per partition, non-comparable:** each partition resolves independently; you get a different offset per partition for the same wall-clock time. There is no single global offset. - **Null results:** if the requested time is after the partition's newest record, you get `null` and must handle it (e.g. seek to end). - **Clock dependence with CreateTime:** unreliable producer clocks make time-based seeks fuzzy; for precise replay, `LogAppendTime` is safer. - **Granularity:** the `.timeindex` is sparse (`log.index.interval.bytes`), so resolution is to the nearest indexed entry plus a forward scan — fine for seeking, not a precise per-millisecond guarantee of which record is 'first'. - **Retention:** if the data for that time has aged out, the earliest available offset (log-start-offset) is the best you can get; older timestamps resolve to the log-start. - **Costs a round trip** per call; batch all partitions into one `offsetsForTimes` call.

  • Why might offset-by-timestamp give surprising results with CreateTime topics?
    With CreateTime the producer sets the timestamp, so clock skew, late/reordered producers, or reprocessing can make timestamps non-monotonic across offsets. The lookup returns the first record in index order with ts ≥ X, which may not perfectly match wall-clock ordering. LogAppendTime avoids this by stamping at append.
  • What does offsetsForTimes return if you ask for a timestamp newer than every record in a partition?
    It returns null for that partition, meaning no record has a timestamp at or after the requested time. Code must handle null, typically by seeking to the end (LEO) or skipping the partition.

saying these in an interview costs you the question

  • Saying it returns the last record before the timestamp (it returns the first record at or after it).
  • Assuming a single offset applies across all partitions for a given time.
  • Ignoring CreateTime vs LogAppendTime when reasoning about correctness.
  • Believing it scans the whole log instead of using the .timeindex.

context