What happens to partition selection when a record has a null key, and why was the sticky partitioner (KIP-480) introduced?
answer
- null key -> nothing to hash
- old: per-record round-robin -> tiny batches
- KIP-480 sticky: fill one batch, then switch
- fuller batches = lower latency, higher throughput
- KIP-794 made it load-aware, deprecated DefaultPartitioner
basics
~20 sWith a null key there is nothing to hash, so Kafka picks a partition without a key. The old DefaultPartitioner round-robined per record; KIP-480's sticky partitioner instead fills one partition's batch, then switches, producing fuller batches and lower latency.
solid answer
~50 sA null-keyed record has no key to hash, so the partitioner must choose freely. Before Kafka 2.4, DefaultPartitioner did per-record round-robin among available partitions. That spread records thinly: with many partitions, each batch held few records, so batches lingered until linger.ms elapsed and throughput/latency suffered. KIP-480 introduced the **sticky partitioner**: it 'sticks' to one partition until that partition's batch is full (or linger.ms fires), then picks a new random partition for the next batch. This produces larger, fuller batches, fewer requests, and lower latency under load, while still distributing roughly evenly over time because the sticky choice rotates per batch. In Kafka 2.4 the sticky behavior became part of DefaultPartitioner (and was exposed as UniformStickyPartitioner). It only applies to null keys — keyed records still hash. Over short windows distribution is uneven (one partition at a time), but it evens out across many batches.
go deeper
Know null keys don't get hashed and that newer Kafka batches them onto one partition at a time.
Explain round-robin vs sticky and why sticky gives fuller batches and lower latency (KIP-480, 2.4).
Discuss batch.size/linger.ms interplay, short-window skew, and the KIP-794 load-aware successor.
Reason about end-to-end latency/throughput tradeoffs, consumer burstiness from stickiness, and when to tune or replace the built-in partitioner.
## Why null keys need special handling When a `ProducerRecord` has **no key** (key is null), there is nothing to hash, so the producer is free to put the record on any partition. The goal becomes balancing load evenly while batching efficiently. ## The old behavior: per-record round-robin Before Kafka 2.4, `DefaultPartitioner` assigned null-key records by **round-robin**: record 1 → partition 0, record 2 → partition 1, and so on. This is evenly distributed but has a hidden cost tied to **batching**. The producer groups records into per-partition **batches** before sending. A batch is sent when it reaches `batch.size` bytes OR `linger.ms` milliseconds elapse, whichever first. If you round-robin every record across, say, 30 partitions, each batch accumulates only ~1/30th of the records, so batches rarely fill — they almost always wait the full `linger.ms`. Result: many small requests, more overhead, and *higher* latency because records sit waiting. ## KIP-480: the sticky partitioner KIP-480 (Kafka 2.4) changed the strategy for null keys to **stickiness**: - Pick one partition and keep sending null-key records to **that same partition** until its batch is completed (full at `batch.size`, or `linger.ms` fires). - When that batch closes, choose a **new random partition** for the next batch and stick to it. Because records concentrate on one partition at a time, batches fill quickly → **fuller batches, fewer requests, lower latency, higher throughput** under load. Counterintuitively, sticking to one partition *reduces* latency versus spreading, because you stop waiting on half-empty batches. Over many batches the random rotation spreads load evenly across partitions, so long-term distribution stays balanced. ## Where it lives In 2.4 the sticky logic became the default inside `DefaultPartitioner`, and `UniformStickyPartitioner` was added as an explicit class with the same behavior for callers who want it named. Both apply **only to null keys**; keyed records still go through murmur2 hashing. ## Edge cases and the later story - **Short-window skew:** within a single batch window all null-key records land on one partition, so instantaneous distribution is lopsided. Downstream consumers may see bursts. It balances out across batches, not within one. - **Idle/low-rate producers:** with low throughput, batches close on `linger.ms` rather than size, so the partition switches frequently and the benefit shrinks. - **KIP-794 (Kafka 3.3):** the original sticky partitioner could actually send *more* to slower partitions in some cases. KIP-794 added a **strictly uniform / partition-load-aware** sticky partitioner and deprecated `DefaultPartitioner`/`UniformStickyPartitioner` in favor of `partitioner.class=null` with `partitioner.adaptive.partitioning.enable` and `partitioner.availability.timeout.ms`. Mentioning this shows current knowledge.
- Why does sticking to one partition reduce latency instead of increasing it?Latency comes from records waiting in half-empty batches until linger.ms expires. Concentrating null-key records on one partition fills that batch to batch.size quickly, so it ships sooner and you make fewer, larger requests. Spreading per-record kept every batch nearly empty, forcing the full linger wait.
- Does the sticky partitioner affect records with a non-null key?No. Stickiness only governs null-key records. Keyed records always go through murmur2 hashing so per-key ordering is preserved.
saying these in an interview costs you the question
- Saying the sticky partitioner randomizes per record (it sticks per batch, not per record).
- Claiming stickiness increases latency because it 'overloads one partition' — it lowers latency by filling batches.
- Thinking the sticky partitioner changes how keyed records are placed.
- Believing distribution is perfectly even within any short window — it's even only over many batches.