How can a producer bypass the partitioner and write to a specific partition explicitly, and what are the consequences of doing so?
answer
- ProducerRecord(topic, partition, key, value)
- explicit partition skips the partitioner
- key still stored, not used to route
- must be in [0, numPartitions)
- you own balancing + repartition gaps
basics
~20 sUse a ProducerRecord constructor that takes a partition number, e.g. new ProducerRecord(topic, partition, key, value). Then the partitioner is skipped entirely and the record goes to exactly that partition. You become responsible for balancing and for valid partition numbers.
solid answer
~40 sIf you build a `ProducerRecord` with an explicit partition argument (`new ProducerRecord<>(topic, partition, key, value)`), Kafka uses that partition directly and never calls the configured `Partitioner` — the key, if present, is still stored but not used for routing. This is useful for precise control: replaying into a known partition, custom sharding decided by your app, or co-locating data you've already routed upstream. The consequences: you own load balancing (easy to create skew), you must pass a partition in `[0, numPartitions)` or you get an error, ordering is whatever you arrange, and if the topic later gains partitions your hardcoded numbers won't use them. Compared to a custom Partitioner, explicit partitions move routing logic into the call site rather than a reusable plugin — fine for one-off control, worse for consistency across many producers.
go deeper
Know that you can pass a partition number in ProducerRecord to send to a specific partition.
Explain that the explicit partition skips the partitioner, the key is still stored, and you now own balancing and bounds.
Compare explicit assignment vs custom Partitioner (call-site vs reusable plugin) and the repartition/skew consequences.
Decide when app-level explicit routing is justified vs centralizing routing in a partitioner for operability and consistency across producers.
**The two routing paths.** Normally the producer picks a partition for you: with a key it hashes (`murmur2(key) % N`), without a key it load-balances (sticky). But you can override both by **specifying the partition yourself**. **How.** `ProducerRecord` has overloaded constructors. The ones with an `Integer partition` argument — `new ProducerRecord<>(topic, partition, key, value)` (and variants with timestamp/headers) — tell the producer to send to *exactly* that partition. When `partition` is non-null, the producer **does not invoke the partitioner at all**. The key is still serialized and stored on the record (so compaction and consumer-side key access still work), it just isn't used to choose the partition. **When it's useful.** - **Deterministic replay / repair:** re-emit a record into the same partition it originally came from. - **App-level sharding:** your service already computed a shard and wants direct control without writing a Partitioner plugin. - **Testing / tooling:** force records onto a specific partition. **Consequences and pitfalls.** 1. **You own balancing.** No automatic spreading — bad math creates partition skew and a hot consumer. 2. **Bounds matter.** The partition must be in `[0, numPartitions)`. An out-of-range partition fails the send. (The producer/broker validates against current metadata.) 3. **Ordering is your problem.** Same-key-same-partition is only guaranteed if you keep sending that key to the same number. 4. **Repartitioning blind spot.** If the topic grows from 6 to 12 partitions, hardcoded `0..5` never touches the new partitions, leaving them empty and unbalanced. 5. **Scattered logic.** Routing now lives at every call site instead of one `Partitioner` — harder to keep consistent across services. For reusable, system-wide routing prefer a custom `Partitioner`; reserve explicit partitions for targeted, local decisions. **Precedence.** Explicit partition > custom/default partitioner. If you pass a partition, the partitioner is irrelevant for that record.
- If you pass both a key and an explicit partition, what does the key do?The key is still serialized and stored on the record (so it's available for log compaction and to consumers), but it is NOT used for routing — the explicit partition wins and the partitioner is never called.
- What breaks if the topic's partition count later increases and you hardcoded partitions 0-5?Records keep going only to 0-5; the new partitions stay empty, creating imbalance and wasting the added capacity. Explicit assignment doesn't adapt to partition-count changes the way a modulo-based or built-in partitioner can.
saying these in an interview costs you the question
- Saying the partitioner still runs when a partition is specified (it doesn't)
- Thinking the key overrides the explicit partition
- Forgetting bounds-checking — an invalid partition number fails the send
- Claiming explicit partition is the recommended default for load balancing