How does the message key passed to KafkaTemplate.send() affect partitioning, and when would you write a custom Partitioner?
answer
- key hash (murmur2) % partitions
- same key -> same partition -> ordering
- null key -> sticky/round-robin spread
- custom Partitioner via partitioner.class
- adding partitions breaks hash affinity
basics
~20 sThe key decides the partition: Kafka hashes the key and maps it to a partition, so all records with the same key go to the same partition and stay ordered. With a null key, records are spread across partitions. A custom Partitioner overrides this mapping.
solid answer
~40 sWhen you call send(topic, key, value), the key is serialized and, if no explicit partition is given, the producer's Partitioner chooses the partition. The default logic hashes the key (murmur2) modulo the partition count, so equal keys always land on the same partition — that's how Kafka gives you per-key ordering (e.g. all events for one orderId in sequence). A null key uses a sticky/round-robin strategy that spreads load across partitions with no ordering guarantee across records. You write a custom Partitioner (implementing org.apache.kafka.clients.producer.Partitioner, set via partitioner.class) when default hashing isn't enough — e.g. routing hot keys to dedicated partitions, geography-based routing, or keeping a group of related keys co-located. Remember: partition count is fixed at send time; adding partitions later changes the hash mapping and breaks existing key→partition affinity.
code
java · 16 linespublic class RegionPartitioner implements Partitioner {
@Override
public int partition(String topic, Object key, byte[] keyBytes,
Object value, byte[] valueBytes, Cluster cluster) {
int numPartitions = cluster.partitionCountForTopic(topic);
String k = (String) key;
// route EU keys to partition 0, everything else hashed across the rest
if (k != null && k.startsWith("EU-")) return 0;
return 1 + Math.floorMod(Objects.hashCode(k), Math.max(1, numPartitions - 1));
}
@Override public void close() {}
@Override public void configure(Map<String, ?> configs) {}
}
// Register it on the ProducerFactory:
// props.put(ProducerConfig.PARTITIONER_CLASS_CONFIG, RegionPartitioner.class.getName());go deeper
Know same key = same partition = ordered; null key spreads records.
Explain hash % partitions, per-key ordering, and null-key round-robin/sticky behavior.
Discuss custom Partitioner use cases, skew/hot partitions, and the repartitioning hazard.
Weigh ordering vs parallelism, capacity-plan partition counts, and design keys to avoid skew across the whole system.
**Why partitions matter.** A Kafka *topic* is split into *partitions*; each partition is an ordered, append-only log. Ordering is guaranteed **only within a partition**, never across partitions. Consumers in a group each own some partitions. So *which partition a record goes to* determines both ordering and load distribution. **The key's role.** `send(topic, key, value)` sets the `ProducerRecord`'s key. When you don't specify an explicit partition, the producer's **Partitioner** picks one: - **Non-null key:** the default partitioning hashes the *serialized key bytes* with **murmur2** and takes modulo the number of partitions. Deterministic ⇒ the same key always maps to the same partition (as long as partition count is unchanged). This is the mechanism behind **per-entity ordering**: put `orderId` as the key and every event for that order is appended in order to one partition. - **Null key:** since Kafka 2.4 the built-in logic uses a **sticky partitioner** — it fills one partition's batch, then switches, giving good batching while spreading records roughly evenly. No cross-record ordering is implied. (Older clients did plain round-robin.) **Explicit partition.** `send(topic, partition, key, value)` bypasses the partitioner entirely and pins the record to that partition number. **Custom Partitioner.** Implement `org.apache.kafka.clients.producer.Partitioner` and register it via the producer property `partitioner.class` (in Spring, put it in the ProducerFactory config map / `spring.kafka.producer.properties.partitioner.class`). Its `partition(topic, key, keyBytes, value, valueBytes, cluster)` returns an int partition. Use cases: - **Hot-key isolation** — route a known high-volume key to its own partition so it doesn't skew others. - **Semantic routing** — send by region/tenant to specific partitions consumed by dedicated processors. - **Co-location** — force several distinct keys to share a partition so a consumer sees them together. - **Custom balancing** for skewed key distributions where murmur2 clusters unevenly. **Gotchas.** 1. **Partition count is baked into the hash.** `hash % N` — if you increase partitions from N to N+1, most keys remap to different partitions, so ordering/affinity for existing keys breaks and previously co-located history is now split. Plan partition counts up front. 2. **Skew.** A poorly chosen key (low cardinality, or one dominant value like a single big customer) creates a *hot partition* — one consumer overloaded while others idle. 3. **Ordering vs parallelism trade-off.** Strong per-key ordering needs same-partition, which caps parallelism at the partition count. More partitions = more parallelism but more overhead. 4. **Serialization of the key** happens before partitioning; the partitioner sees `keyBytes`. A different key serializer changes the bytes and thus the mapping. 5. **Retries + ordering:** with `max.in.flight.requests.per.connection>1` and retries, same-partition records can reorder unless `enable.idempotence=true`. **When to just use defaults.** For most workloads the default keyed hashing is exactly right — pick a good business key and don't write a custom partitioner unless you have a concrete routing or skew requirement.
- Why can increasing a topic's partition count break ordering guarantees for existing keys?Default partitioning is hash(key) % partitionCount. Changing the count changes the modulus, so most keys map to a different partition than before. New records for a key may land on a new partition while its history sits on the old one — the single-partition ordering assumption no longer holds.
- You have a null key and need related records ordered together. What are your options?A null key won't give ordering. Options: assign a meaningful key so they hash to the same partition, pass an explicit partition, or write a custom Partitioner that maps those records to one partition based on the value.
saying these in an interview costs you the question
- Believing Kafka guarantees ordering across an entire topic rather than per partition
- Thinking a null key still keeps related records ordered together
- Assuming you can freely add partitions without affecting key→partition mapping
- Choosing a low-cardinality key and not anticipating a hot partition