How does a Kafka producer decide which partition a record goes to when the record has a non-null key?
answer
- murmur2(keyBytes) % numPartitions
- Utils.toPositive
- same key -> same partition -> per-key order
- ordering only within a partition
- adding partitions breaks the mapping
basics
~20 sFor a keyed record, Kafka hashes the key and maps it to a partition. The same key always lands on the same partition (as long as the partition count is unchanged), which keeps records with that key in order.
solid answer
~40 sWhen a record has a non-null key, the default partitioner computes murmur2(serializedKey) and takes it modulo the number of partitions: partition = Utils.toPositive(murmur2(keyBytes)) % numPartitions. Because the hash is deterministic, every record with the same key consistently maps to the same partition. That is what gives you per-key ordering: Kafka only guarantees ordering within a partition, so co-locating a key's records on one partition means they are consumed in produce order. The hash is computed on the serialized key bytes, so the serializer matters. One caveat: the modulo uses the current partition count, so adding partitions later breaks the mapping for existing keys (a key may move to a different partition), which can break ordering assumptions during repartitioning.
go deeper
Know that a key is hashed and the same key always goes to the same partition, giving per-key ordering.
Name murmur2 and the modulo-by-partition-count formula, and explain that ordering is only within a partition.
Discuss the serialized-bytes detail, the repartitioning hazard, and how key distribution drives skew.
Reason about partition-count provisioning, co-partitioning contracts across producers/streams, and ordering guarantees across topic resizes.
## The problem partitioning solves A Kafka topic is split into **partitions** — independent, append-only logs. Kafka only guarantees message ordering *within a single partition*, never across partitions. So if you need records that share something (e.g. all events for `user-42`) to be processed in order, they must all go to the same partition. ## How a key picks a partition When you send a `ProducerRecord` with a **non-null key**, the producer's partitioner does this: 1. Serialize the key to bytes using the configured key serializer. 2. Hash those bytes with **murmur2** (Kafka's chosen non-cryptographic hash, implemented in `org.apache.kafka.common.utils.Utils.murmur2`). 3. Force the hash positive and take it modulo the partition count: `partition = Utils.toPositive(Utils.murmur2(keyBytes)) % numPartitions`. Because murmur2 is deterministic, the **same key bytes always produce the same partition** (for a fixed partition count). That is the mechanism behind *key-based ordering*: all records for `user-42` land on, say, partition 3, and a consumer reading partition 3 sees them in the order they were produced. ## Why the serializer matters The hash is over the *serialized* bytes, not the logical value. If you change key serializers (e.g. String vs Avro), the byte representation — and therefore the partition — can change. Two producers must use the same serialization to co-partition the same logical key. ## The repartitioning caveat The `% numPartitions` step ties the mapping to the **current** partition count. If you grow a topic from 6 to 12 partitions, `hash % 6` and `hash % 12` generally differ, so existing keys may relocate. New records for `user-42` may now go to a different partition than the historical ones, breaking the global ordering guarantee for that key across the resize boundary. This is why teams over-provision partitions up front or use keyed compaction carefully — Kafka does not rehash/move existing data when you add partitions. ## Skew Key hashing distributes keys evenly *only if keys are diverse*. A few hot keys (e.g. one tenant generating most traffic) concentrate on their partitions, causing **partition skew** — some partitions/brokers/consumers do far more work. The hash itself is uniform; skew comes from the key distribution.
- Why can adding partitions to a topic break ordering for existing keys?The partition is hash % numPartitions. Changing numPartitions changes the modulo result, so a key that mapped to partition 3 under 6 partitions may map elsewhere under 12. Old and new records for the same key then live on different partitions, and Kafka only orders within a partition.
- Does the key value or its serialized bytes get hashed?The serialized bytes. murmur2 runs on the output of the key serializer, so two producers must use the same serializer to co-partition the same logical key.
saying these in an interview costs you the question
- Saying Kafka guarantees global ordering across the whole topic (it only orders within a partition).
- Claiming a hash like SHA-256 or Java's hashCode is used — it is murmur2.
- Believing adding partitions automatically moves existing keys to keep them consistent.
- Thinking the logical key value is hashed rather than its serialized bytes.