skip to content

Offset Addressing Model

Offsets as partition-local addresses: log-start offset, LEO, high watermark, and committed position. Interviewers use it to check you know offsets are not comparable across partitions and are not a global sequence.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

6

What is an offset in Kafka, and why are offsets meaningful only within a single partition?

level: juniorimportance: must knowfreq 78%

answer

  1. per-partition counter from 0
  2. monotonic, immutable, leader-assigned
  3. no global offset across topic
  4. offset 5 in P0 != offset 5 in P1
  5. ordering only within a partition

basics

~20 s

An offset is a number that marks a record's position inside one partition. It starts at 0 and increases by 1 for each new record. Offsets are per-partition, so offset 5 in partition 0 and offset 5 in partition 1 are unrelated records.

solid answer

~40 s

Each Kafka partition is an append-only log, and every record written to it gets a monotonically increasing 64-bit offset: the first record is offset 0, the next is 1, and so on, with no gaps in a clean log (compaction can later create gaps). The offset is assigned by the partition leader at append time and never changes. Because each partition numbers its records independently, offsets are partition-local: there is no ordering or comparability of offsets across partitions. Consumers track progress per (topic, partition) — committing 'partition 0 -> offset 100' says nothing about partition 1. This is why Kafka only guarantees ordering within a partition, not across a topic, and why the partition is the unit of parallelism and offset bookkeeping.

go deeper

for a junior

Know: offset = position in a partition, starts at 0, goes up by 1, and is per-partition.

for a middle

Explain leader-assigns-at-append, immutability, and why ordering is only per-partition.

for a senior

Discuss gaps from compaction/transactions, the absence of a global offset by design, and offset as the unit of consumer bookkeeping.

for a principal

Frame the partition-local offset model as the design choice that enables horizontal scale and per-partition parallelism; teach the ordering/scaling tradeoff.

## What a partition is A Kafka **topic** is split into one or more **partitions**. Each partition is an ordered, append-only sequence of records — conceptually a log file you can only add to at the end. Producers append records; the data is immutable once written. ## What an offset is An **offset** is the position of a record within its partition. The very first record in a partition has offset `0`; the next has `1`, then `2`, and so on. The offset is a 64-bit integer (`long`). It is: - **Monotonically increasing** — each new append gets a strictly larger offset than the previous one. - **Assigned by the partition leader** at the moment the record is appended to the log. The producer does not choose it. - **Immutable** — once a record sits at offset 42, it is offset 42 forever; offsets are never reused or renumbered. - **Dense in a normal log** — consecutive, no gaps. (Exception: **log compaction** removes superseded records, leaving offset gaps; transactional topics also consume offsets for commit/abort control markers, so you can see small gaps even in non-compacted topics.) ## Why offsets are partition-local Each partition maintains its **own** offset counter, starting from 0 independently. So: - Offset `5` in partition `0` and offset `5` in partition `1` are **completely unrelated records**. - You **cannot compare** offsets across partitions: 'offset 1000 in P0' is not 'later' or 'earlier' than 'offset 30 in P1'. They live in different number lines. - There is **no global offset** for a topic. Kafka deliberately avoids a global sequence number because that would require cross-partition coordination and destroy horizontal scalability. ## Consequences - **Ordering guarantee:** Kafka guarantees ordering only *within* a partition. Records that must be ordered relative to each other must land in the same partition (typically by sharing a key). - **Consumer progress** is tracked per `(topic, partition)`. A consumer's committed offset for partition 0 says nothing about partition 1. - **Parallelism:** because each partition is an independent log with its own offsets, partitions can be consumed in parallel by different consumers in a group. ## Edge cases - A brand-new partition has no records; its first append will be offset 0. - After retention deletes old segments, the lowest available offset (the **log-start-offset**) moves forward, but the offsets of remaining records do not change. - Compaction and transaction markers can make offsets non-contiguous, so consumers must never assume `nextOffset == lastOffset + 1`.

  • Can offsets ever have gaps in a partition?
    Yes. Log compaction removes superseded records leaving gaps, and transaction control markers (commit/abort) consume offsets. So consumers must not assume offsets are strictly contiguous; they should rely on the next offset returned by Kafka, not lastOffset + 1.
  • If I need records A and B processed in order, what must I ensure?
    They must go to the same partition, since ordering is only guaranteed within a partition. Typically you give them the same record key so the default partitioner hashes them to the same partition.

saying these in an interview costs you the question

  • Saying there is a single global offset that orders the whole topic.
  • Claiming offsets are comparable across partitions (e.g., 'P0 offset 100 is after P1 offset 30').
  • Saying offsets always increment by exactly 1 with no gaps ever (ignores compaction and transaction markers).
  • Thinking the producer chooses the offset.

context

open as a page

Distinguish a consumer's committed offset from its current position. How do poll, position, and commit interact?

level: middleimportance: must knowfreq 70%

basics

~20 s

The current position is the offset of the next record the consumer will fetch — it lives in memory and moves forward as you poll. The committed offset is the position durably saved to Kafka, used to resume after a restart or rebalance. They can differ.

open as a page

What is the high watermark, how does it relate to LEO, and why can a consumer not read up to the LEO?

level: seniorimportance: must knowfreq 60%

basics

~20 s

The high watermark is the offset up to which records are replicated to all in-sync replicas, so they are 'committed' and safe. Consumers can only read records below the high watermark, never up to the LEO, because un-replicated records could be lost on leader failure.

open as a page

Define base offset, log-start-offset, and log-end-offset (LEO) for a Kafka partition. How do they move over time?

level: middleimportance: should knowfreq 55%

basics

~20 s

Log-start-offset is the offset of the earliest record still retained. LEO (log-end-offset) is the offset that will be assigned to the next record — one past the last written record. Base offset is the first offset of a particular log segment file.

open as a page

How does Kafka resolve an offset from a timestamp (offsetsForTimes), and what are its limitations?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Kafka can look up, per partition, the offset of the earliest record whose timestamp is >= a given time, using the consumer's offsetsForTimes (ListOffsets API) and the per-segment .timeindex files. It returns null if no record has a timestamp at or after that time.

open as a page

A teammate wants a single monotonic sequence number across all partitions of a topic to order events globally. Why is this impossible with offsets, and what are the alternatives?

level: principalimportance: should knowfreq 35%

basics

~20 s

Offsets are independent per-partition counters, so they can't order events across partitions. Kafka gives no global ordering by design. Alternatives: use a single partition, key-partition only what must be ordered, or add an application-level timestamp/sequence and reorder downstream.

open as a page