skip to content

Choosing and Changing Partition Count

Picking a partition count and living with it: the consumer parallelism ceiling, per-partition overhead, and why you can add partitions but never remove them. Interviewers like it because adding partitions silently breaks key-to-partition stability.

part ofApache Kafkaoverview, primer and where to startread it →
on this pageshow

questions

5

What does `kafka-topics.sh --alter --partitions` allow you to do, and what is the key directionality constraint?

level: juniorimportance: must knowfreq 70%

answer

  1. --alter --partitions = increase only
  2. no in-place shrink ever
  3. createPartitions admin API
  4. shrink = new topic + republish
  5. increase still breaks key routing

basics

~10 s

It can only INCREASE the partition count of an existing topic, never decrease it. To shrink, you must create a new topic with fewer partitions and republish the data; Kafka has no in-place shrink.

solid answer

~40 s

`kafka-topics.sh --alter --topic <t> --partitions <N>` changes a topic's partition count, but only upward: N must be greater than the current count. Kafka rejects any attempt to lower it because shrinking would require deciding what to do with the data in the partitions being removed and would break offset/log semantics. There is no `--partitions` value below current that the broker will accept. If you truly need fewer partitions, the only path is to create a brand-new topic with the desired count and republish (mirror) the data into it, then cut consumers over. Also be aware that increasing partitions changes key-to-partition routing for new records, so keyed ordering guarantees are disrupted from the moment of the increase.

go deeper

for a junior

Memorize: --alter --partitions only goes up, never down.

for a middle

Explain WHY shrink is impossible (offsets, independent logs) and the republish workaround.

for a senior

Connect the increase to broken key routing and the operational steps to repartition safely.

for a principal

Frame partition count as a hard-to-reverse architectural decision and design topics with headroom and key-aware sizing up front.

## What the command does ``` kafka-topics.sh --bootstrap-server localhost:9092 \ --alter --topic orders --partitions 12 ``` This tells the broker to expand topic `orders` to 12 partitions. The admin API behind it is `AdminClient.createPartitions(...)`. ## Increase-only constraint Kafka **only supports increasing** the partition count. If the topic currently has 12 partitions and you pass `--partitions 6`, the broker rejects it with an error (`InvalidPartitionsException` / "Topic currently has N partitions, which is higher than the requested N"). The same value or a lower value is invalid; the new count must be strictly greater. ### Why no shrink? - Each partition is an independent log with its own offsets and committed consumer positions. Removing a partition would orphan its data and its committed offsets. - Producers and consumers cache partition metadata; silently dropping partitions would create correctness hazards. - The log files, segments, and replicas for a partition are physical; collapsing them into others has no well-defined, safe semantics. So Kafka simply forbids it rather than offering a lossy operation. ## The real cost of increasing Increasing is allowed but not free of consequences: - **Key routing changes.** The default partitioner maps a key by `hash(key) % numPartitions`. When `numPartitions` changes, the same key can now map to a different partition. Records for a key written before and after the change can live in different partitions, breaking per-key ordering and any consumer assuming a key always lands in one partition. - **No rebalancing of existing data.** Old records stay where they were; only new writes use the new partition count. ## Shrinking: republish pattern The supported way to reduce partitions is **repartition-by-republish**: 1. Create a new topic with the target (smaller) partition count. 2. Run a consumer/producer (or MirrorMaker/Streams/Connect) that reads the old topic and writes to the new one. 3. Migrate consumers, then producers, to the new topic. 4. Retire the old topic. This is also how you'd change the *partitioning key* or fix a too-high partition count.

  • You ran --alter to go from 6 to 6 partitions by mistake. What happens?
    It is rejected — the new count must be strictly greater than the current count. Equal or lower values are invalid (InvalidPartitionsException).
  • Your team needs to go from 50 partitions down to 10 because per-partition overhead is hurting the cluster. What's the procedure?
    Create a new topic with 10 partitions and republish all data into it (consumer→producer bridge, MirrorMaker, or Streams), migrate consumers and producers over, then delete the old 50-partition topic. There is no in-place shrink.

saying these in an interview costs you the question

  • Believing `--alter` can lower the partition count.
  • Thinking increasing partitions redistributes existing records across the new partitions.
  • Assuming an increase preserves key-to-partition mapping.

context

open as a page

How does the partition count of a Kafka topic relate to consumer parallelism, and what is the maximum number of consumers in a single consumer group that can actively consume from it?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Partition count is the parallelism ceiling for one consumer group. Each partition is assigned to at most one consumer in the group, so if a topic has 6 partitions, at most 6 consumers do useful work; extra consumers sit idle.

open as a page

Why does increasing the partition count of a keyed topic break ordering and key-locality guarantees, and how do you avoid that disruption?

level: seniorimportance: must knowfreq 65%

basics

~20 s

The default partitioner routes a key via hash(key) % numPartitions. Changing numPartitions changes the modulo result, so a key that always landed in partition 2 may now land elsewhere — splitting that key's records across partitions and breaking per-key ordering.

open as a page

What are the per-partition costs that make over-partitioning harmful, and what cluster limits should you weigh when choosing a high partition count?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Each partition costs open file handles (log segments + index files), broker memory, replication threads, and controller/metadata load. More partitions also slow leader-election failover and lengthen rebalances. Open file descriptor limits and end-to-end latency cap how many partitions a cluster can hold.

open as a page

Walk through how you would repartition a high-traffic keyed topic — for example to change the partition count or partitioning key — given Kafka cannot do it in place. What ordering and cutover concerns matter?

level: principalimportance: should knowfreq 45%

basics

~20 s

Create a new topic with the desired partition count/key scheme, then republish: a bridge consumer reads the old topic and a producer writes into the new one with the new partitioning. Migrate consumers, then producers, to the new topic, drain the old, and retire it.

open as a page