skip to content

When would you implement a custom Kafka Partitioner, and what does the interface require you to do?

level: seniorimportance: should knowfreq 45%

answer

  1. implements Partitioner.partition(...)
  2. partitioner.class config
  3. cluster.availablePartitionsForTopic
  4. runs on send hot path
  5. you own null-key + determinism

basics

~20 s

Implement org.apache.kafka.clients.producer.Partitioner when the default key-hash routing isn't what you need — e.g. routing hot keys specially or grouping by a field of the value. You override partition(...) to return the partition number and wire it with partitioner.class.

solid answer

~40 s

You write a custom `Partitioner` when default `murmur2(key) % N` routing doesn't fit: examples include isolating a few hot/whale keys onto dedicated partitions, routing on something other than the literal key (a field inside the value, a tenant id), geo/affinity routing, or pinning low-cardinality keys to avoid skew. You implement `org.apache.kafka.clients.producer.Partitioner` — `partition(topic, key, keyBytes, value, valueBytes, cluster)` returns an int in `[0, numPartitions)`, plus `configure(Map)` and `close()`. Get the partition list via `cluster.partitionsForTopic(topic)`, and you can read `cluster.availablePartitionsForTopic` to avoid offline leaders. You register it with `partitioner.class=com.acme.MyPartitioner`. Caveats: you must preserve determinism if you rely on ordering, handle null keys yourself, keep it cheap (it runs on the send hot path), and remember that adding partitions still re-maps modulo-style logic unless you design around it.

go deeper

for a junior

Know that you can plug in custom routing via partitioner.class but the default is usually fine.

for a middle

Name the Partitioner interface and that partition() returns the partition int, registered via partitioner.class.

for a senior

Give concrete use cases (hot-key isolation, value-based routing), the full method signature, and the determinism/null-key/hot-path caveats.

for a principal

Weigh custom routing against operability — repartitioning, skew monitoring, and whether built-in partitioner.ignore.keys/consistent hashing is a better fit than bespoke code.

**Why customize.** The default partitioner routes non-null keys by `murmur2(keyBytes) % numPartitions` and load-balances null keys. That is right most of the time, but real systems sometimes need different routing: - **Hot-key / whale isolation:** a handful of keys carry most traffic; pin them to dedicated partitions so they don't starve everyone sharing a partition. - **Routing on the value, not the key:** e.g. partition by a `region` field inside the payload while keeping a different key for compaction. - **Multi-tenant affinity:** keep one tenant's data on a known partition subset for isolation or locality. - **Reducing skew for low-cardinality keys:** spread a few distinct keys deliberately. **The interface.** `org.apache.kafka.clients.producer.Partitioner` (extends `Configurable`, `Closeable`): - `int partition(String topic, Object key, byte[] keyBytes, Object value, byte[] valueBytes, Cluster cluster)` — return the target partition. You get both the deserialized objects and their byte forms, plus the `Cluster` metadata. - `void configure(Map<String,?> configs)` — read your own props (passed through producer config). - `void close()` — release resources. **Using cluster metadata.** `cluster.partitionsForTopic(topic)` gives all partitions; `cluster.availablePartitionsForTopic(topic)` gives only those with a live leader. Routing to an unavailable partition risks send failures, so for load-balancing logic you often prefer the available list (the built-in sticky logic does this). **Registration.** `props.put(ProducerConfig.PARTITIONER_CLASS_CONFIG, "com.acme.MyPartitioner")` (or `partitioner.class=...`). The producer instantiates it per producer. **Critical caveats.** 1. **Hot path:** `partition(...)` is called for every record — keep it allocation-light and fast. 2. **Determinism vs. ordering:** if downstream relies on per-key ordering, the same key must always return the same partition; non-deterministic logic breaks ordering. 3. **Null keys:** the default null-key handling (sticky) is *not* applied automatically inside your custom logic — you decide what null keys do. 4. **Partition-count changes:** any `% N` scheme re-maps if N changes — historic keys move. Design explicitly (e.g. a fixed lookup, consistent-hashing) if that matters. 5. **Don't reinvent balancing badly:** if you only want uniform balance, prefer the built-in partitioner / `partitioner.ignore.keys` over a hand-rolled round-robin that fights batching.

  • Your custom partitioner routes on a value field. What ordering risk does that create versus default key routing?
    Default key routing keeps all records for one key on one partition (ordered). Routing on a value field means records that share a key but differ in that field scatter across partitions, so per-key ordering is lost; you only preserve ordering for whatever you route on.
  • Why prefer cluster.availablePartitionsForTopic over partitionsForTopic in balancing logic?
    availablePartitionsForTopic excludes partitions whose leader is currently offline; routing to those would fail or block. For uniform balancing you want only partitions that can accept writes; for strict key-determinism you may still need the full list to keep mapping stable.

saying these in an interview costs you the question

  • Claiming you must subclass DefaultPartitioner (you implement the Partitioner interface)
  • Forgetting that the custom partitioner must handle null keys itself
  • Putting heavy I/O or blocking calls inside partition() on the send hot path
  • Assuming a custom modulo scheme survives a partition-count increase

context