skip to content

Partition Keys and Hash Routing

How a record key becomes a partition: murmur2 hashing, null-key sticky behavior, and custom partitioners. Interviewers ask because key choice decides both ordering and skew.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

When you send a Kafka record with a non-null key, how does the producer decide which partition it goes to, and why does that matter?

level: juniorimportance: must knowfreq 78%

answer

  1. murmur2(key) % numPartitions
  2. deterministic same-key-same-partition
  3. ordering only within a partition
  4. key skew = partition skew
  5. not Object.hashCode

basics

~10 s

The producer hashes the key and takes hash % numberOfPartitions. The same key always lands in the same partition, so all records for that key stay ordered together on one partition.

solid answer

~40 s

With a non-null key the default partitioner computes a hash of the serialized key bytes and maps it to a partition with hash % numPartitions. This is deterministic: the same key always routes to the same partition (as long as the partition count is unchanged), which is the mechanism that guarantees per-key ordering, because Kafka only orders records within a single partition. Kafka uses the murmur2 hash over the key bytes (not Java's Object.hashCode), so routing is consistent across producer instances and languages. This matters when you need related events — say all events for one user or order id — to be processed in order; you key by that id so they share a partition. The trade-off is that key skew (a few hot keys) creates partition skew and uneven consumer load.

go deeper

for a junior

Know: non-null key => hash(key) % partitions => same key always same partition => keeps that key's records ordered.

for a middle

Add the murmur2 detail, the modulo-on-partition-count caveat, and that ordering is only within a partition.

for a senior

Discuss key skew, repartitioning consequences, and the link between keying, ordering and compaction.

for a principal

Frame keying as the lever for ordering vs. parallelism trade-offs and capacity planning (partition count chosen up front).

**The problem.** A Kafka topic is split into N **partitions** — append-only logs. Kafka guarantees ordering only *within* a partition, never across partitions. So if you need related records processed in order, they must land on the same partition. The **partition key** is how you control that. **The mechanism.** When you call `producer.send(record)` and the record has a non-null key, the producer serializes the key to bytes, then the **partitioner** maps those bytes to a partition number. The default algorithm is `Utils.murmur2(serializedKey)` then `toPositive(hash) % numPartitions`. murmur2 is a fast non-cryptographic hash; Kafka uses it (not `Object.hashCode()`) precisely so that the same key bytes hash identically regardless of which producer, JVM, or client language emits them. **Determinism.** Given the same key bytes and the same partition count, the result is always the same partition. That is what delivers **per-key ordering**: every record with key `user-42` goes to, say, partition 3, and within partition 3 they are appended in send order. **Why it matters.** Ordering, co-location (compaction keeps the latest value per key in one place), and even consumer parallelism all hinge on keying. You key by the entity whose events must stay ordered (order id, account id, device id). **Edge cases.** - **Changing partition count breaks the mapping.** `hash % N` changes if N changes, so old keys may route to new partitions — historical ordering for a key is not preserved across a repartition. This is why teams over-provision partitions up front. - **Key skew → partition skew.** If 90% of traffic has the same key, one partition (and one consumer) gets hammered. - **null key** is a different path entirely (sticky partitioning), covered separately.

  • What happens to existing keys' routing if you add partitions to a topic?
    The modulo divisor changes, so keys re-map: a key that was on partition 3 may now hash to a different partition. Past records stay where they were, so per-key ordering is not preserved across the partition-count change.
  • Why does Kafka use murmur2 instead of the key's Java hashCode?
    hashCode is JVM/implementation-specific and not portable across languages or even JVM versions; murmur2 over the raw serialized bytes gives a stable, language-agnostic mapping so producers in any client compute the same partition.

saying these in an interview costs you the question

  • Saying Kafka guarantees global ordering across all partitions
  • Claiming the same key can land on different partitions run-to-run (it can't, given fixed partition count)
  • Thinking it uses Java Object.hashCode()
  • Believing adding partitions preserves existing key routing

context

open as a page

What happens to partition selection when a record's key is null, and how has that behavior changed across Kafka versions (sticky partitioning)?

level: middleimportance: must knowfreq 62%

basics

~20 s

With a null key there's no hash to route on, so the producer spreads records across partitions. Modern Kafka uses sticky partitioning: it fills one partition's batch, then switches to another, instead of strict round-robin per record.

open as a page

How can a producer bypass the partitioner and write to a specific partition explicitly, and what are the consequences of doing so?

level: middleimportance: should knowfreq 38%

basics

~20 s

Use a ProducerRecord constructor that takes a partition number, e.g. new ProducerRecord(topic, partition, key, value). Then the partitioner is skipped entirely and the record goes to exactly that partition. You become responsible for balancing and for valid partition numbers.

open as a page

When would you implement a custom Kafka Partitioner, and what does the interface require you to do?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Implement org.apache.kafka.clients.producer.Partitioner when the default key-hash routing isn't what you need — e.g. routing hot keys specially or grouping by a field of the value. You override partition(...) to return the partition number and wire it with partitioner.class.

open as a page

A team keys Kafka records by customer_id and a few large customers cause severe partition skew. Walk through the trade-offs and options for fixing it.

level: principalimportance: should knowfreq 34%

basics

~20 s

Keying by customer_id keeps each customer ordered but routes all of a hot customer's traffic to one partition, overloading it and its consumer. Options: composite keys to spread, a custom partitioner, more partitions, or accepting weaker per-key ordering — each trades ordering against balance.

open as a page