What is an offset in Kafka, and why are offsets meaningful only within a single partition?
answer
- per-partition counter from 0
- monotonic, immutable, leader-assigned
- no global offset across topic
- offset 5 in P0 != offset 5 in P1
- ordering only within a partition
basics
~20 sAn offset is a number that marks a record's position inside one partition. It starts at 0 and increases by 1 for each new record. Offsets are per-partition, so offset 5 in partition 0 and offset 5 in partition 1 are unrelated records.
solid answer
~40 sEach Kafka partition is an append-only log, and every record written to it gets a monotonically increasing 64-bit offset: the first record is offset 0, the next is 1, and so on, with no gaps in a clean log (compaction can later create gaps). The offset is assigned by the partition leader at append time and never changes. Because each partition numbers its records independently, offsets are partition-local: there is no ordering or comparability of offsets across partitions. Consumers track progress per (topic, partition) — committing 'partition 0 -> offset 100' says nothing about partition 1. This is why Kafka only guarantees ordering within a partition, not across a topic, and why the partition is the unit of parallelism and offset bookkeeping.
go deeper
Know: offset = position in a partition, starts at 0, goes up by 1, and is per-partition.
Explain leader-assigns-at-append, immutability, and why ordering is only per-partition.
Discuss gaps from compaction/transactions, the absence of a global offset by design, and offset as the unit of consumer bookkeeping.
Frame the partition-local offset model as the design choice that enables horizontal scale and per-partition parallelism; teach the ordering/scaling tradeoff.
## What a partition is A Kafka **topic** is split into one or more **partitions**. Each partition is an ordered, append-only sequence of records — conceptually a log file you can only add to at the end. Producers append records; the data is immutable once written. ## What an offset is An **offset** is the position of a record within its partition. The very first record in a partition has offset `0`; the next has `1`, then `2`, and so on. The offset is a 64-bit integer (`long`). It is: - **Monotonically increasing** — each new append gets a strictly larger offset than the previous one. - **Assigned by the partition leader** at the moment the record is appended to the log. The producer does not choose it. - **Immutable** — once a record sits at offset 42, it is offset 42 forever; offsets are never reused or renumbered. - **Dense in a normal log** — consecutive, no gaps. (Exception: **log compaction** removes superseded records, leaving offset gaps; transactional topics also consume offsets for commit/abort control markers, so you can see small gaps even in non-compacted topics.) ## Why offsets are partition-local Each partition maintains its **own** offset counter, starting from 0 independently. So: - Offset `5` in partition `0` and offset `5` in partition `1` are **completely unrelated records**. - You **cannot compare** offsets across partitions: 'offset 1000 in P0' is not 'later' or 'earlier' than 'offset 30 in P1'. They live in different number lines. - There is **no global offset** for a topic. Kafka deliberately avoids a global sequence number because that would require cross-partition coordination and destroy horizontal scalability. ## Consequences - **Ordering guarantee:** Kafka guarantees ordering only *within* a partition. Records that must be ordered relative to each other must land in the same partition (typically by sharing a key). - **Consumer progress** is tracked per `(topic, partition)`. A consumer's committed offset for partition 0 says nothing about partition 1. - **Parallelism:** because each partition is an independent log with its own offsets, partitions can be consumed in parallel by different consumers in a group. ## Edge cases - A brand-new partition has no records; its first append will be offset 0. - After retention deletes old segments, the lowest available offset (the **log-start-offset**) moves forward, but the offsets of remaining records do not change. - Compaction and transaction markers can make offsets non-contiguous, so consumers must never assume `nextOffset == lastOffset + 1`.
- Can offsets ever have gaps in a partition?Yes. Log compaction removes superseded records leaving gaps, and transaction control markers (commit/abort) consume offsets. So consumers must not assume offsets are strictly contiguous; they should rely on the next offset returned by Kafka, not lastOffset + 1.
- If I need records A and B processed in order, what must I ensure?They must go to the same partition, since ordering is only guaranteed within a partition. Typically you give them the same record key so the default partitioner hashes them to the same partition.
saying these in an interview costs you the question
- Saying there is a single global offset that orders the whole topic.
- Claiming offsets are comparable across partitions (e.g., 'P0 offset 100 is after P1 offset 30').
- Saying offsets always increment by exactly 1 with no gaps ever (ignores compaction and transaction markers).
- Thinking the producer chooses the offset.