When you send a Kafka record with a non-null key, how does the producer decide which partition it goes to, and why does that matter?
answer
- murmur2(key) % numPartitions
- deterministic same-key-same-partition
- ordering only within a partition
- key skew = partition skew
- not Object.hashCode
basics
~10 sThe producer hashes the key and takes hash % numberOfPartitions. The same key always lands in the same partition, so all records for that key stay ordered together on one partition.
solid answer
~40 sWith a non-null key the default partitioner computes a hash of the serialized key bytes and maps it to a partition with hash % numPartitions. This is deterministic: the same key always routes to the same partition (as long as the partition count is unchanged), which is the mechanism that guarantees per-key ordering, because Kafka only orders records within a single partition. Kafka uses the murmur2 hash over the key bytes (not Java's Object.hashCode), so routing is consistent across producer instances and languages. This matters when you need related events — say all events for one user or order id — to be processed in order; you key by that id so they share a partition. The trade-off is that key skew (a few hot keys) creates partition skew and uneven consumer load.
go deeper
Know: non-null key => hash(key) % partitions => same key always same partition => keeps that key's records ordered.
Add the murmur2 detail, the modulo-on-partition-count caveat, and that ordering is only within a partition.
Discuss key skew, repartitioning consequences, and the link between keying, ordering and compaction.
Frame keying as the lever for ordering vs. parallelism trade-offs and capacity planning (partition count chosen up front).
**The problem.** A Kafka topic is split into N **partitions** — append-only logs. Kafka guarantees ordering only *within* a partition, never across partitions. So if you need related records processed in order, they must land on the same partition. The **partition key** is how you control that. **The mechanism.** When you call `producer.send(record)` and the record has a non-null key, the producer serializes the key to bytes, then the **partitioner** maps those bytes to a partition number. The default algorithm is `Utils.murmur2(serializedKey)` then `toPositive(hash) % numPartitions`. murmur2 is a fast non-cryptographic hash; Kafka uses it (not `Object.hashCode()`) precisely so that the same key bytes hash identically regardless of which producer, JVM, or client language emits them. **Determinism.** Given the same key bytes and the same partition count, the result is always the same partition. That is what delivers **per-key ordering**: every record with key `user-42` goes to, say, partition 3, and within partition 3 they are appended in send order. **Why it matters.** Ordering, co-location (compaction keeps the latest value per key in one place), and even consumer parallelism all hinge on keying. You key by the entity whose events must stay ordered (order id, account id, device id). **Edge cases.** - **Changing partition count breaks the mapping.** `hash % N` changes if N changes, so old keys may route to new partitions — historical ordering for a key is not preserved across a repartition. This is why teams over-provision partitions up front. - **Key skew → partition skew.** If 90% of traffic has the same key, one partition (and one consumer) gets hammered. - **null key** is a different path entirely (sticky partitioning), covered separately.
- What happens to existing keys' routing if you add partitions to a topic?The modulo divisor changes, so keys re-map: a key that was on partition 3 may now hash to a different partition. Past records stay where they were, so per-key ordering is not preserved across the partition-count change.
- Why does Kafka use murmur2 instead of the key's Java hashCode?hashCode is JVM/implementation-specific and not portable across languages or even JVM versions; murmur2 over the raw serialized bytes gives a stable, language-agnostic mapping so producers in any client compute the same partition.
saying these in an interview costs you the question
- Saying Kafka guarantees global ordering across all partitions
- Claiming the same key can land on different partitions run-to-run (it can't, given fixed partition count)
- Thinking it uses Java Object.hashCode()
- Believing adding partitions preserves existing key routing