skip to content

Partitioner Strategies

How records land on partitions through keyed hashing, sticky batching for null keys, and custom partitioners. A frequent question because a bad partitioner produces hot partitions and consumer skew.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

How does a Kafka producer decide which partition a record goes to when the record has a non-null key?

level: juniorimportance: must knowfreq 78%

answer

  1. murmur2(keyBytes) % numPartitions
  2. Utils.toPositive
  3. same key -> same partition -> per-key order
  4. ordering only within a partition
  5. adding partitions breaks the mapping

basics

~20 s

For a keyed record, Kafka hashes the key and maps it to a partition. The same key always lands on the same partition (as long as the partition count is unchanged), which keeps records with that key in order.

solid answer

~40 s

When a record has a non-null key, the default partitioner computes murmur2(serializedKey) and takes it modulo the number of partitions: partition = Utils.toPositive(murmur2(keyBytes)) % numPartitions. Because the hash is deterministic, every record with the same key consistently maps to the same partition. That is what gives you per-key ordering: Kafka only guarantees ordering within a partition, so co-locating a key's records on one partition means they are consumed in produce order. The hash is computed on the serialized key bytes, so the serializer matters. One caveat: the modulo uses the current partition count, so adding partitions later breaks the mapping for existing keys (a key may move to a different partition), which can break ordering assumptions during repartitioning.

go deeper

for a junior

Know that a key is hashed and the same key always goes to the same partition, giving per-key ordering.

for a middle

Name murmur2 and the modulo-by-partition-count formula, and explain that ordering is only within a partition.

for a senior

Discuss the serialized-bytes detail, the repartitioning hazard, and how key distribution drives skew.

for a principal

Reason about partition-count provisioning, co-partitioning contracts across producers/streams, and ordering guarantees across topic resizes.

## The problem partitioning solves A Kafka topic is split into **partitions** — independent, append-only logs. Kafka only guarantees message ordering *within a single partition*, never across partitions. So if you need records that share something (e.g. all events for `user-42`) to be processed in order, they must all go to the same partition. ## How a key picks a partition When you send a `ProducerRecord` with a **non-null key**, the producer's partitioner does this: 1. Serialize the key to bytes using the configured key serializer. 2. Hash those bytes with **murmur2** (Kafka's chosen non-cryptographic hash, implemented in `org.apache.kafka.common.utils.Utils.murmur2`). 3. Force the hash positive and take it modulo the partition count: `partition = Utils.toPositive(Utils.murmur2(keyBytes)) % numPartitions`. Because murmur2 is deterministic, the **same key bytes always produce the same partition** (for a fixed partition count). That is the mechanism behind *key-based ordering*: all records for `user-42` land on, say, partition 3, and a consumer reading partition 3 sees them in the order they were produced. ## Why the serializer matters The hash is over the *serialized* bytes, not the logical value. If you change key serializers (e.g. String vs Avro), the byte representation — and therefore the partition — can change. Two producers must use the same serialization to co-partition the same logical key. ## The repartitioning caveat The `% numPartitions` step ties the mapping to the **current** partition count. If you grow a topic from 6 to 12 partitions, `hash % 6` and `hash % 12` generally differ, so existing keys may relocate. New records for `user-42` may now go to a different partition than the historical ones, breaking the global ordering guarantee for that key across the resize boundary. This is why teams over-provision partitions up front or use keyed compaction carefully — Kafka does not rehash/move existing data when you add partitions. ## Skew Key hashing distributes keys evenly *only if keys are diverse*. A few hot keys (e.g. one tenant generating most traffic) concentrate on their partitions, causing **partition skew** — some partitions/brokers/consumers do far more work. The hash itself is uniform; skew comes from the key distribution.

  • Why can adding partitions to a topic break ordering for existing keys?
    The partition is hash % numPartitions. Changing numPartitions changes the modulo result, so a key that mapped to partition 3 under 6 partitions may map elsewhere under 12. Old and new records for the same key then live on different partitions, and Kafka only orders within a partition.
  • Does the key value or its serialized bytes get hashed?
    The serialized bytes. murmur2 runs on the output of the key serializer, so two producers must use the same serializer to co-partition the same logical key.

saying these in an interview costs you the question

  • Saying Kafka guarantees global ordering across the whole topic (it only orders within a partition).
  • Claiming a hash like SHA-256 or Java's hashCode is used — it is murmur2.
  • Believing adding partitions automatically moves existing keys to keep them consistent.
  • Thinking the logical key value is hashed rather than its serialized bytes.

context

open as a page

What happens to partition selection when a record has a null key, and why was the sticky partitioner (KIP-480) introduced?

level: middleimportance: must knowfreq 70%

basics

~20 s

With a null key there is nothing to hash, so Kafka picks a partition without a key. The old DefaultPartitioner round-robined per record; KIP-480's sticky partitioner instead fills one partition's batch, then switches, producing fuller batches and lower latency.

open as a page

What is the difference between DefaultPartitioner and UniformStickyPartitioner, and when would you choose one over the other?

level: middleimportance: should knowfreq 48%

basics

~10 s

Both use sticky batching for null keys. DefaultPartitioner hashes non-null keys with murmur2 to preserve per-key ordering; UniformStickyPartitioner ignores the key entirely and sticks for every record, so it never gives key-based ordering.

open as a page

How do you implement a custom Partitioner in Kafka, and what is a legitimate use case for one?

level: seniorimportance: should knowfreq 42%

basics

~10 s

Implement org.apache.kafka.clients.producer.Partitioner (the partition() method returns an int partition), then set partitioner.class to your class. A common reason is routing hot or special keys to dedicated partitions to control skew or isolate VIP traffic.

open as a page

Your keyed topic shows severe partition skew (a few partitions hold most of the load). What are your options, and what is the fundamental tension you must navigate?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Skew comes from uneven key distribution (hot keys), not the hash. Options: salt or compound the key, route hot keys with a custom partitioner, or increase partitions. The tension: spreading a key for balance destroys the per-key ordering the key was meant to guarantee.

open as a page