Walk through how you would repartition a high-traffic keyed topic — for example to change the partition count or partitioning key — given Kafka cannot do it in place. What ordering and cutover concerns matter?
answer
- new topic, don't alter in place
- bridge consumer -> producer re-keys
- backfill then catch up to tail
- producers cut over, then consumers
- dual-write window reorders keys
basics
~20 sCreate a new topic with the desired partition count/key scheme, then republish: a bridge consumer reads the old topic and a producer writes into the new one with the new partitioning. Migrate consumers, then producers, to the new topic, drain the old, and retire it.
solid answer
~50 sSince Kafka can't shrink partitions or re-key in place, you do a repartition-by-republish into a fresh topic. Steps: (1) create the new topic with the target partition count and any new partitioner/key. (2) Stand up a bridge — a consumer reading the old topic producing into the new one, re-keying so every key is consistently placed under the new scheme. (3) Let the bridge backfill and catch up to near-tail. (4) Cut producers over to write to the new topic. (5) Once the bridge has drained the remaining tail of the old topic, switch consumers to the new topic and decommission the old one. Ordering is the hard part: during the dual-write window, in-flight records for a key may exist in both topics; quiesce or fence per-key writes at cutover, or design consumers to be idempotent/order-tolerant. Use the same key on the new topic so a key still lands in one partition. Tools: a custom job, Kafka Streams, MirrorMaker 2, or Connect.
go deeper
Know the headline: you can't repartition in place; you copy into a new topic.
Lay out the basic steps: new topic, bridge, switch producers then consumers, delete old.
Address ordering during dual-write, offset translation, and idempotency/dedup.
Own the full cutover plan, tooling choice, capacity, and the upstream design decisions (sizing, key strategy) that avoid needing this at all.
## Why republish at all Kafka offers no in-place way to (a) reduce partitions, (b) change the partitioning key, or (c) change partition count without disturbing key locality. The only correct mechanism is **repartition-by-republish**: build a new topic laid out the way you want and copy the data through a re-partitioning step. ## The procedure 1. **Create the target topic.** `kafka-topics.sh --create --topic orders.v2 --partitions <target> ...` with the desired partition count and (in code) the desired key/partitioner. 2. **Build the bridge.** A consumer subscribes to the old topic; for each record it computes the new key (if re-keying) and produces to the new topic. Because this single pass re-routes *all* records against the new partition count, every key's full history lands in one partition of the new topic — the consistency an in-place increase can't give you. Options: a bespoke consumer→producer app, **Kafka Streams** (`through`/repartition), **MirrorMaker 2**, or **Kafka Connect**. 3. **Backfill + catch up.** Run the bridge from the old topic's earliest offset to replicate history, then let it approach the tail. 4. **Producer cutover.** Switch producers to write to the new topic. From here new data only enters via the new topic; the bridge keeps draining whatever producers wrote to the old topic before they switched. 5. **Consumer cutover.** When the bridge has drained the old topic's remaining tail, repoint consumer groups at the new topic. Because the new topic has its own offsets, plan how consumers start (earliest vs a translated offset). 6. **Retire the old topic** once nothing reads or writes it. ## Ordering and correctness concerns - **Dual-write window.** Between producer cutover and consumer cutover, a key's records may exist partly in the old topic (still draining through the bridge) and partly written directly to the new topic. Without care this reorders the key. Mitigations: route a key's writes through only one path at a time (fence/quiesce per key), or make consumers idempotent and order-tolerant (e.g. carry a monotonic sequence/version and drop stale). - **Keep the key.** Use the same logical key on the new topic so per-key locality and ordering are preserved within the new topic. - **Offset translation.** Old and new topics have independent offsets; committed positions don't carry over. Decide whether consumers reprocess from earliest or use a tool (MM2 offset sync) to translate. - **Exactly-once / dedup.** Backfill + live traffic can double-deliver; use idempotent producers and consumer-side dedup keys to stay correct. - **Capacity.** During migration you temporarily store the data twice and run extra I/O for the bridge — size for it. ## When you can avoid all this Provision the right partition count and a stable key-routing strategy up front. Republishing is the escape hatch precisely because the partition decision is expensive to reverse — that's the design lesson.
- During the migration, why is the window between producer cutover and consumer cutover risky for ordering?In that window a key can have older records still draining from the old topic through the bridge while newer records were written directly to the new topic. The consumer can then see the new records before the bridged-in old ones, reordering the key. You mitigate by fencing each key to one write path at a time or making consumers idempotent and version-aware.
- Which tools can act as the republish bridge, and what do they buy you?A bespoke consumer→producer app (full control of re-keying), Kafka Streams (built-in repartition + state handling), Kafka Connect (config-driven), or MirrorMaker 2 (cross/intra-cluster replication with offset translation). Streams and MM2 add offset/state handling you'd otherwise build yourself.
- Why must you keep the same logical key on the new topic?So each key still maps to a single partition under the new partitioner, preserving per-key locality and ordering. Changing the key would re-scatter related records and defeat the purpose unless that re-keying is the explicit goal.
saying these in an interview costs you the question
- Proposing an in-place --alter to reduce partitions or change the key (impossible).
- Ignoring the dual-write window and assuming ordering is automatically preserved during cutover.
- Assuming consumer offsets carry over between the old and new topic.
- Forgetting the temporary double storage / double I/O cost of running the bridge.