How does message keying preserve per-key ordering, and what is the exact mechanism?
answer
- partition = murmur2(key) % numPartitions
- Same key -> same partition -> ordered
- null key -> sticky partitioner, no per-key order
- Repartition breaks key mapping
- Explicit partition / custom partitioner override
basics
~20 sRecords with the same key are hashed to the same partition by the producer's default partitioner. Since one partition is strictly ordered, all records for a given key are processed in the order they were produced.
solid answer
~40 sTo keep events for the same entity ordered, you set the record key (e.g. userId). The default partitioner computes partition = murmur2(key) % numPartitions, so identical keys deterministically map to the same partition. Because that partition is a single ordered log, all records sharing a key are totally ordered relative to each other, while different keys spread across partitions for parallelism. Caveats: (1) the mapping depends on numPartitions, so increasing partition count remaps keys and breaks ordering for in-flight/new records relative to old ones; (2) null-key records are distributed (round-robin / sticky partitioner) with no per-key guarantee; (3) a custom partitioner or explicit partition argument overrides this; (4) Kafka 2.4+ uses the sticky partitioner for null keys to batch better. Same-key routing is the standard pattern for ordered-by-entity streams.
go deeper
Knows same key -> same partition -> ordered.
Explains murmur2 % numPartitions and the null-key/sticky-partitioner case.
Calls out repartitioning as the silent ordering breaker and overrides.
Designs key schemes and partition counts up front to avoid ever needing to repartition ordered topics.
## The problem Kafka only orders within a partition. So to order all events for one logical entity (a user, an order, an account), you must funnel them all into the *same* partition. Keying is how you do that. ## The key Every Kafka record is `(key, value, headers, ...)`. The **key** is optional bytes. When you set a key (e.g. the account ID), the producer uses it to choose a partition. ## The default partitioner mechanism With a non-null key, the default partitioner computes: ``` partition = murmur2(serialize(key)) % numPartitions ``` `murmur2` is a fast non-cryptographic hash. Because it's deterministic, the *same key always yields the same partition* (as long as `numPartitions` is unchanged). Therefore every record with key `K` lands in one fixed partition, and since that partition is a strictly ordered log, all `K` records are totally ordered — even though records for other keys are spread across other partitions and processed in parallel. ## Null keys If the key is `null`, there is nothing to hash. Old clients used round-robin; **Kafka 2.4+** uses the **sticky partitioner** (and KIP-794's strictly-uniform partitioner in 3.3+) to fill a batch for one partition before moving on, improving batching/throughput. Either way, **null-key records get no per-key ordering**. ## Critical caveat: changing partition count The modulo `% numPartitions` means the key→partition mapping depends on the partition count. If you grow a topic from 6 to 12 partitions, key `K` may now hash to a different partition. Records for `K` produced *after* the change can land in a different partition than older `K` records, so the global per-key order across the change boundary is **not** preserved. This is why partition count is treated as nearly immutable for keyed/ordered topics. ## Other overrides - Passing an explicit `partition` to `ProducerRecord` bypasses the partitioner entirely. - A **custom partitioner** (`partitioner.class`) can implement different routing (e.g. tenant-aware). ## Putting it together Keying gives you *ordering where it matters* (per entity) while keeping *parallelism where it's safe* (across entities). It's the canonical Kafka ordering design.
- What breaks per-key ordering even when you key correctly?Increasing the partition count (modulo remaps keys), null keys, an explicit partition override, or consumer-side parallel processing that reorders records after delivery.
- Why is changing partition count dangerous for ordered topics?Because partition selection is hash(key) % numPartitions; changing numPartitions sends a key to a new partition, so new records for that key are no longer co-located/ordered with its older records.
saying these in an interview costs you the question
- Claiming keys guarantee ordering regardless of partition count
- Thinking null-key records keep per-key order
- Believing the partitioner uses the value, not the key
- Assuming you can freely add partitions to a keyed topic with no ordering impact