skip to content

A teammate wants a single monotonic sequence number across all partitions of a topic to order events globally. Why is this impossible with offsets, and what are the alternatives?

level: principalimportance: should knowfreq 35%

answer

  1. offsets per-partition, no global counter
  2. global order = global serialization = no scale
  3. single partition: total order, no parallelism
  4. key-partition = per-entity order (idiomatic)
  5. app timestamp/seq + downstream reorder

basics

~20 s

Offsets are independent per-partition counters, so they can't order events across partitions. Kafka gives no global ordering by design. Alternatives: use a single partition, key-partition only what must be ordered, or add an application-level timestamp/sequence and reorder downstream.

solid answer

~50 s

Offsets are **partition-local**: each partition increments its own counter from 0, with no coordination between partitions, so two records in different partitions have offsets on separate number lines that can't be compared. A truly global, gap-free monotonic sequence would require cross-partition consensus on every write, which destroys the horizontal scalability that partitioning exists to provide — so Kafka deliberately offers **ordering only within a partition**. Practical alternatives: (1) **single partition** — perfect total order but no parallelism and capped throughput; (2) **partition by key** so only records that must be ordered share a partition (per-entity ordering, the common pattern); (3) carry an **application-level monotonic value** (event-time timestamp, a producer-side sequence, or an external sequence like a DB sequence/snowflake ID) and **reorder downstream** by buffering/windowing. Each trades latency, throughput, or complexity. The right answer is almost always per-key ordering, not global.

go deeper

for a junior

Know that offsets don't order events across partitions and Kafka only orders within a partition.

for a middle

Recommend partition-by-key for per-entity ordering and explain why a single partition limits throughput.

for a senior

Lay out the single-partition vs key-partition vs app-sequence tradeoffs and the idempotent-producer caveat.

for a principal

Interrogate the real ordering requirement, explain why global order conflicts with scale, and design idempotent/event-time pipelines accordingly.

## Why the request is impossible as stated Kafka offsets are **per-partition counters**. Partition 0 numbers its records 0,1,2…; partition 1 independently numbers its records 0,1,2…. There is **no shared counter** and **no cross-partition coordination** at append time. Therefore: - Offsets across partitions are **not comparable** — 'P0 offset 900' is neither before nor after 'P1 offset 12'. - There is **no topic-wide monotonic sequence**. This is a **design choice, not a limitation to fix**. A global gap-free sequence would force every producer to obtain the next number from a single coordinating authority before each append — a global serialization point. That reintroduces exactly the bottleneck partitioning exists to remove, capping write throughput at what one coordinator can do and adding a single point of contention/failure. Kafka instead guarantees **ordering and monotonic offsets only within a single partition**, which is what makes it horizontally scalable. ## What ordering Kafka does guarantee - **Within a partition:** total order; offsets strictly increasing; a consumer reads records in offset order. - **Across partitions:** none. Consumers may interleave partitions arbitrarily. ## Alternatives (with tradeoffs) ### 1. Single partition Put everything in one partition. You get a perfect total order (offset == global sequence). Cost: **no parallelism** — one consumer in a group can read it, throughput is capped by one broker/partition, and you lose Kafka's main scaling lever. Acceptable only for low-volume streams that genuinely need total order. ### 2. Partition by key (per-entity ordering) Give records a **key** (e.g. accountId, orderId). The default partitioner hashes the key so all records for a key land in the **same** partition and are totally ordered **relative to each other**. Different keys may interleave, but you usually only need ordering *per entity*, not globally. This is the idiomatic answer and preserves parallelism (up to the partition count). ### 3. Application-level ordering value + downstream reorder Embed a monotonic value in the payload and reorder after consumption: - **Event-time timestamp** (and use Kafka Streams / Flink event-time windows with watermarks to handle out-of-order arrival). - **Producer-side sequence number** per logical stream. - **External global ID**: a database sequence, or a **Snowflake-style ID** (time + node + counter) that is globally roughly-monotonic without a central bottleneck on every write. Consumers buffer across partitions and emit in sorted order within a bounded window. Cost: **latency** (you must wait for laggards), **memory** (buffering), and **complexity** (handling late/duplicate events). ## Decision guidance (principal lens) - Ask **what ordering is actually required**. 'Global order' is usually over-specified; per-entity order suffices -> key-partition. - If true total order is non-negotiable and volume is small -> single partition. - If you need global order *and* scale -> accept eventual/approximate order via timestamps + downstream windowing, and design idempotent/commutative processing so exact global order matters less. - Beware false 'solutions': a global counter service, distributed locks, or `acks=all` do **not** create cross-partition order. ## Edge cases - Even within a partition, **retries with `max.in.flight.requests > 1` and no idempotence** can reorder; enable the idempotent producer (`enable.idempotence=true`) to preserve per-partition order under retries. - Compaction and transaction markers create offset gaps, so even a single-partition 'sequence' isn't perfectly contiguous.

  • If a stream truly needs total order but also high throughput, what would you tell the team?
    Those goals conflict: total order needs a single serialization point (one partition), which caps throughput. I'd push to relax the requirement to per-key ordering (partition by key) which scales, and only fall back to a single partition for genuinely low-volume must-be-totally-ordered streams. For high volume needing global order, accept approximate ordering via event-time timestamps and downstream windowing with idempotent processing.
  • Does enabling the idempotent producer help with cross-partition ordering?
    No. enable.idempotence=true preserves ordering and dedup within a single partition under retries (and is needed when max.in.flight > 1). It does nothing for cross-partition ordering, which remains undefined by design.

saying these in an interview costs you the question

  • Proposing a global counter/lock service as if it scales like Kafka.
  • Claiming offsets can be compared across partitions.
  • Treating 'global order' as a default requirement instead of interrogating whether per-key order suffices.
  • Believing acks=all or replication provides cross-partition ordering.

context