What guarantees does Kafka provide about offset ordering within a partition, and how do segments preserve it?
answer
- per-partition order only
- offset = append position, immutable
- next segment base = prev last + 1
- increasing but NOT gapless
- LEO vs high water mark
basics
~20 sWithin a single partition, offsets are strictly increasing and assigned in append order — record N+1 always has a higher offset than record N. Across partitions there is no ordering. Segments preserve this because each new segment's base offset continues where the previous one ended.
solid answer
~50 sKafka guarantees that offsets within one partition are monotonically increasing in append order: each appended record (or batch) gets the next offset, never a lower one, and records are immutable once written. Order is only guaranteed per partition, not across a topic. Segments preserve monotonicity because rolls are seamless: when a segment closes at offset M, the next segment's base offset is the offset right after M, so the global sequence is continuous. Note offsets are not necessarily gapless — log compaction can remove intermediate records, transaction markers consume offsets, and aborted transactions leave offsets that consumers skip — so offsets always increase but may have holes. The high water mark (last committed/replicated offset) and log end offset (next offset to be assigned) track the tail. Consumers rely on monotonicity to resume from a committed offset deterministically.
go deeper
Know offsets increase in append order within a partition.
Add per-partition-only scope and seamless continuation across segment boundaries.
Explain increasing-but-not-gapless, LEO vs HWM, and keyed routing for ordering.
Reason about ordering under unclean leader election, truncation, and exactly-once/transaction semantics.
## The core guarantee An **offset** is the position of a record within its partition. Kafka assigns offsets on the leader at append time, strictly in the order records are appended. The guarantee is **per-partition total order**: if record A is appended before record B in the same partition, then offset(A) < offset(B), and that ordering is permanent because records are immutable. There is **no ordering guarantee across partitions** of the same topic — that's the price of partition-level parallelism. To get ordering for related records, you route them to the same partition (usually via a shared key, since partition = hash(key) % numPartitions by default). ## How segments keep it continuous The partition log is split into segments, but the offset space is one continuous sequence across them. When the active segment rolls at last offset M, the new segment's **base offset** is M+1 (the offset of its first record). So reading segment files in base-offset order reproduces the exact append order with no gaps at segment boundaries. This is why locating a record is a two-step floor lookup: find the segment whose base offset is the largest <= target, then use that segment's .index. ## Increasing != gapless A subtle senior point: offsets are **monotonically increasing** but **not guaranteed contiguous**. Holes can appear because: - **Log compaction** deletes superseded records by key, leaving gaps where they used to be. - **Transactions** write control records (commit/abort markers) that consume offsets but are invisible to consumers. - **Aborted transactions** leave records whose offsets a read-committed consumer skips. So a consumer may see offsets 5, 7, 8, 12 — increasing, with gaps. Code must never assume offset+1 is the next visible record. ## Tail markers - **Log End Offset (LEO):** the offset that will be assigned to the next record — i.e., one past the last record. - **High Water Mark (HWM):** the highest offset that has been replicated to all in-sync replicas and is therefore visible to consumers. Consumers cannot read past the HWM. The active segment holds records between HWM and LEO that are written but not yet fully committed. ## Why it matters Monotonicity is what makes committed offsets meaningful: a consumer that commits offset 100 can resume at 100 and know it will see every record after it, in order, exactly once relative to position. It also underpins replication (followers fetch by offset) and exactly-once semantics. ## Edge cases - After unclean leader election, a new leader may truncate uncommitted records, so the LEO can move backward on a replica — but committed offsets (<= HWM) never reorder. - Producer retries with idempotence enabled prevent duplicate offsets for the same record; without idempotence, duplicates get distinct increasing offsets.
- Are partition offsets always contiguous (no gaps)?No. They strictly increase, but compaction, transaction control markers, and aborted transactions create gaps, so a consumer can see offsets like 5, 7, 12. Never assume the next visible record is offset+1.
- How do you get ordering across what would otherwise be different partitions?You can't order across partitions. Route records that must stay ordered to the same partition by giving them the same key, since the default partitioner maps key -> partition deterministically.
saying these in an interview costs you the question
- Claiming Kafka guarantees ordering across an entire topic — it's per-partition only.
- Assuming offsets are gapless — compaction and transactions create holes.
- Confusing offset (logical, per-partition) with a global timestamp or sequence across the topic.
- Saying offsets can be reused or rewritten — appended records are immutable.