skip to content

What happens to partition selection when a record's key is null, and how has that behavior changed across Kafka versions (sticky partitioning)?

level: middleimportance: must knowfreq 62%

answer

  1. null key => load balance, no hash
  2. KIP-480 sticky = per-batch not per-record
  3. KIP-794 uniform sticky / BuiltInPartitioner
  4. partitioner.ignore.keys=true
  5. sticky != fixed forever

basics

~20 s

With a null key there's no hash to route on, so the producer spreads records across partitions. Modern Kafka uses sticky partitioning: it fills one partition's batch, then switches to another, instead of strict round-robin per record.

solid answer

~40 s

A null key means the partitioner has nothing to hash, so it load-balances. Old behavior (pre-2.4) was per-record round-robin, which scattered records into many small batches and hurt throughput. KIP-480 introduced the **sticky partitioner** (DefaultPartitioner, 2.4+): it 'sticks' to one partition until the current batch fills or `linger.ms` elapses, then picks a new random partition — so a whole batch lands together, improving batching and latency without losing overall balance over time. KIP-794 (3.3+) replaced this with **strictly uniform sticky partitioning** in the BuiltInPartitioner, which also accounts for broker load/queue size so faster brokers get proportionally more records. There is no ordering guarantee tied to a null key. You can also force the murmur2-style behavior to ignore keys entirely via `partitioner.ignore.keys=true`.

go deeper

for a junior

Know that a null key means the producer spreads records to balance load (no key to hash).

for a middle

Explain sticky partitioning (KIP-480): batch-at-a-time instead of per-record round-robin, and why it helps throughput.

for a senior

Contrast KIP-480 vs KIP-794 uniform sticky, BuiltInPartitioner defaults, and partitioner.ignore.keys.

for a principal

Reason about adaptive partitioning under broker skew/backpressure and the throughput-vs-balance trade-offs when designing high-volume keyless streams.

**Setup.** A Kafka record carries an optional key. With a *non-null* key, routing is `murmur2(key) % numPartitions`. With a **null key**, there is nothing to hash, so the producer must instead spread records to balance load. How it spreads has evolved. **Old behavior (before 2.4): round-robin per record.** Each null-key record went to the next partition in rotation. Problem: the producer batches records *per partition* before sending. Round-robin meant each partition's batch held very few records, so you sent many tiny requests — poor compression, more requests, higher latency. **KIP-480 — Sticky Partitioner (2.4, `DefaultPartitioner`).** Instead of rotating every record, the producer **sticks** to one randomly chosen partition and keeps appending null-key records to *that* partition's batch until the batch is full (`batch.size`) or `linger.ms` fires. Then it picks a new random partition. Net effect: full batches, far better throughput and lower latency, and over many batches the distribution across partitions is still roughly uniform. **KIP-794 — Uniform Sticky / `BuiltInPartitioner` (3.3+).** The classic stickiness could over-send to a slow partition while its batch filled. KIP-794 introduced strictly uniform stickiness that switches partitions based on the **amount of data** produced (`partitioner.availability.timeout.ms`, adaptive by broker readiness) so slower/backlogged brokers receive proportionally fewer records. In 3.3+ the explicit `partitioner.class` defaults to `null`, which selects this built-in logic rather than the old `DefaultPartitioner` class. **`partitioner.ignore.keys` (3.3+).** Set `partitioner.ignore.keys=true` to make the built-in partitioner ignore the key even when one is present and apply uniform sticky balancing anyway — useful when you keep keys for compaction/value but do not want key-based routing. **Key takeaways / pitfalls.** - Null key => **no ordering guarantee** for those records relative to each other across partitions. - 'Sticky' does NOT mean a fixed partition forever — it means per-batch, then it moves. - Don't confuse sticky *producer* partitioning with the consumer's *sticky/cooperative-sticky assignor* — unrelated mechanisms that share the word 'sticky'.

  • Why was per-record round-robin replaced — what was the throughput problem?
    It scattered records across many partition batches, so each batch was nearly empty when linger.ms/batch.size triggered a send. That meant many small requests, weak compression, and higher per-record latency. Sticky partitioning fills one batch fully before moving on.
  • If I want sticky load balancing even though my records have keys, how?
    Set partitioner.ignore.keys=true (3.3+) so the built-in partitioner skips the key hash and applies uniform sticky balancing, while the key is still stored for compaction or downstream use.

saying these in an interview costs you the question

  • Saying null-key records guarantee ordering
  • Describing sticky partitioning as pinning to one partition permanently
  • Confusing the producer sticky partitioner with the consumer CooperativeStickyAssignor
  • Claiming null key throws an error or always goes to partition 0

context