When would you implement a custom Kafka Partitioner, and what does the interface require you to do?
answer
- implements Partitioner.partition(...)
- partitioner.class config
- cluster.availablePartitionsForTopic
- runs on send hot path
- you own null-key + determinism
basics
~20 sImplement org.apache.kafka.clients.producer.Partitioner when the default key-hash routing isn't what you need — e.g. routing hot keys specially or grouping by a field of the value. You override partition(...) to return the partition number and wire it with partitioner.class.
solid answer
~40 sYou write a custom `Partitioner` when default `murmur2(key) % N` routing doesn't fit: examples include isolating a few hot/whale keys onto dedicated partitions, routing on something other than the literal key (a field inside the value, a tenant id), geo/affinity routing, or pinning low-cardinality keys to avoid skew. You implement `org.apache.kafka.clients.producer.Partitioner` — `partition(topic, key, keyBytes, value, valueBytes, cluster)` returns an int in `[0, numPartitions)`, plus `configure(Map)` and `close()`. Get the partition list via `cluster.partitionsForTopic(topic)`, and you can read `cluster.availablePartitionsForTopic` to avoid offline leaders. You register it with `partitioner.class=com.acme.MyPartitioner`. Caveats: you must preserve determinism if you rely on ordering, handle null keys yourself, keep it cheap (it runs on the send hot path), and remember that adding partitions still re-maps modulo-style logic unless you design around it.
go deeper
Know that you can plug in custom routing via partitioner.class but the default is usually fine.
Name the Partitioner interface and that partition() returns the partition int, registered via partitioner.class.
Give concrete use cases (hot-key isolation, value-based routing), the full method signature, and the determinism/null-key/hot-path caveats.
Weigh custom routing against operability — repartitioning, skew monitoring, and whether built-in partitioner.ignore.keys/consistent hashing is a better fit than bespoke code.
**Why customize.** The default partitioner routes non-null keys by `murmur2(keyBytes) % numPartitions` and load-balances null keys. That is right most of the time, but real systems sometimes need different routing: - **Hot-key / whale isolation:** a handful of keys carry most traffic; pin them to dedicated partitions so they don't starve everyone sharing a partition. - **Routing on the value, not the key:** e.g. partition by a `region` field inside the payload while keeping a different key for compaction. - **Multi-tenant affinity:** keep one tenant's data on a known partition subset for isolation or locality. - **Reducing skew for low-cardinality keys:** spread a few distinct keys deliberately. **The interface.** `org.apache.kafka.clients.producer.Partitioner` (extends `Configurable`, `Closeable`): - `int partition(String topic, Object key, byte[] keyBytes, Object value, byte[] valueBytes, Cluster cluster)` — return the target partition. You get both the deserialized objects and their byte forms, plus the `Cluster` metadata. - `void configure(Map<String,?> configs)` — read your own props (passed through producer config). - `void close()` — release resources. **Using cluster metadata.** `cluster.partitionsForTopic(topic)` gives all partitions; `cluster.availablePartitionsForTopic(topic)` gives only those with a live leader. Routing to an unavailable partition risks send failures, so for load-balancing logic you often prefer the available list (the built-in sticky logic does this). **Registration.** `props.put(ProducerConfig.PARTITIONER_CLASS_CONFIG, "com.acme.MyPartitioner")` (or `partitioner.class=...`). The producer instantiates it per producer. **Critical caveats.** 1. **Hot path:** `partition(...)` is called for every record — keep it allocation-light and fast. 2. **Determinism vs. ordering:** if downstream relies on per-key ordering, the same key must always return the same partition; non-deterministic logic breaks ordering. 3. **Null keys:** the default null-key handling (sticky) is *not* applied automatically inside your custom logic — you decide what null keys do. 4. **Partition-count changes:** any `% N` scheme re-maps if N changes — historic keys move. Design explicitly (e.g. a fixed lookup, consistent-hashing) if that matters. 5. **Don't reinvent balancing badly:** if you only want uniform balance, prefer the built-in partitioner / `partitioner.ignore.keys` over a hand-rolled round-robin that fights batching.
- Your custom partitioner routes on a value field. What ordering risk does that create versus default key routing?Default key routing keeps all records for one key on one partition (ordered). Routing on a value field means records that share a key but differ in that field scatter across partitions, so per-key ordering is lost; you only preserve ordering for whatever you route on.
- Why prefer cluster.availablePartitionsForTopic over partitionsForTopic in balancing logic?availablePartitionsForTopic excludes partitions whose leader is currently offline; routing to those would fail or block. For uniform balancing you want only partitions that can accept writes; for strict key-determinism you may still need the full list to keep mapping stable.
saying these in an interview costs you the question
- Claiming you must subclass DefaultPartitioner (you implement the Partitioner interface)
- Forgetting that the custom partitioner must handle null keys itself
- Putting heavy I/O or blocking calls inside partition() on the send hot path
- Assuming a custom modulo scheme survives a partition-count increase