Why is a stream's partition count treated as a one-way door, when raising it later is routine and lowering it is not?
answer
- the two directions are not symmetric
- new parts start empty
- removing a part orphans stored records
- down means a replacement stream
- raise changes later key placement
basics
~20 sThe change is asymmetric. Raising the partition count on a stream is a supported, routine operation; lowering it generally is not offered at all, so coming down means creating a second stream at the smaller count and moving writers and readers across.
solid answer
~60 sRaising and lowering are not two directions of one operation. Platforms in this class will generally let you raise a stream's partition count in place, because new parts can simply be added empty and start taking records. They generally will not let you lower it, because the records already stored in the parts that would disappear have nowhere to go without merging stored data and rewriting what every reader is tracking. So the only real way down is to create a new stream with the smaller count, point writers at it, let readers finish the old one, then retire it - a migration and a cut-over, not a setting. That is what makes the number a one-way door: you walk through it at design time, and the cost of having gone too far is a project rather than an edit. It is also why a raise is scheduled rather than casual - where records are placed by hashing a key, a raise changes which part a given key's later records land in.
go deeper
Recall the direction that works: the number of parts on a stream can be increased later, and it cannot be decreased. That single fact explains why the first choice gets attention.
Explain why the two directions differ - a new part starts empty, while a removed part's stored records have nowhere sensible to go - and describe the replacement-stream path as the real way down.
Show that you plan around the raise rather than fearing it: schedule it, account for later records of a key landing elsewhere, and refuse the over-provisioning reflex.
Treat the number as a commitment the organisation lives with: what it forecloses, what reversing it would cost across every team reading the stream, and what default you would publish because of that.
Most capacity numbers on a cluster can be nudged in both directions. The parallelism ceiling of a split stream is not one of them, and the asymmetry is the whole point of the question. ## Why up is easy Adding parts to a stream is cheap because a new part starts **empty**. There is no stored data to move, nothing to merge, and no existing part is disturbed. The cluster records the larger count in its metadata, the new parts are placed on record-serving nodes, and records begin to land in them. From the operator's side this is a structural change with a blast radius small enough to be routine. The thing that makes a raise something you schedule rather than do casually is what it does to **placement**. Where a platform decides a record's part by hashing a key, the arithmetic that maps key to part depends on how many parts there are - so after a raise, later records for a given key can land in a different part from their predecessors. Teams that care about that consequence plan the raise around it. (What that means for ordering is its own subject, and this question does not need it.) ## Why down is not on offer Removing a part would mean deciding what happens to the records already stored in it. Every option is bad: - **Discard them** - silent data loss, which no platform will do on your behalf. - **Merge them into a surviving part** - the merged records interleave with records that were written independently, and every reader's understanding of where it had got to in the vanished part is meaningless. - **Block until everything is drained** - only workable for a stream nobody is writing to, which is not the stream you are trying to shrink. Because none of these is acceptable as a generic operation, platforms in this class generally do not offer a decrease at all. The operator-visible fact is simpler than the reasoning: the number goes up, and it does not come down. ## What coming down actually costs The supported path is a **replacement stream**: 1. Create a second stream at the count you actually want. 2. Move writers to it - either at a clean cut-over, or by writing to both for a period. 3. Let readers finish the remaining records on the old stream. 4. Retire the old stream once it has been drained. That is a cross-team migration with a cut-over, a period of ambiguity about where the truth lives, and a name change that every consumer of the stream has to absorb. It is not difficult; it is simply disproportionate to what was, at design time, a single number in a creation request. ## How the asymmetry should change your first choice | pressure | consequence of getting it wrong | |---|---| | count too low | readers hit the ceiling; you raise it, a supported change, planned around key placement | | count too high | no supported way down; you live with the per-part overhead, or you run a migration | The temptation is to read "you cannot come down" as "so start very high". That is the wrong lesson, because a very high count has its own standing costs - fixed overhead per part on every node holding a copy, longer recovery, and a larger metadata set the coordination membership must agree on. The right lesson is narrower: **choose a count you can defend from the reader parallelism you expect within your planning horizon, with modest room, and treat the raise as the supported answer to being wrong on the low side.** ## The queue comparison On a queue-shaped broker this door does not exist, because there is no structural count to walk through. The number of competing readers is an operational dial you turn in either direction by starting and stopping processes. That contrast is worth stating in an interview: the irreversibility is a property of the split-stream model, not of messaging in general. ## What good sounds like A strong answer names the asymmetry, explains *why* the two directions differ rather than reciting that they do, says what the replacement-stream path costs, and resists the over-provisioning conclusion. A weak answer treats the count as a setting like any other, or claims it can be lowered with a brief pause.
- Given a raise is supported, why not create every stream with a very large count so you never need one?Because a part is not free. Each one carries fixed overhead on every node holding a copy of it, multiplied by the copy count and by every stream in the estate, and a very large total lengthens recovery and enlarges the metadata the coordination membership must agree on. Headroom is sensible; an order of magnitude of it is not.
- Is the asymmetry the same on a queue-shaped broker?No. A shared queue has no structural count to raise or lower, so there is no one-way door: you add and remove competing reader processes freely in either direction. The irreversibility described here is a property of platforms that split a stream into parts, not of brokers in general.
- What should be written down when the count is first chosen?The reader parallelism it was sized for, the planning horizon it was sized over, and the fact that going lower means a replacement stream. Without that note the number looks arbitrary to whoever inherits it, and the usual outcome is that nobody dares touch it in either direction.
saying these in an interview costs you the question
- Thinks the count can be lowered with a brief pause
- Says a decrease just merges the parts automatically
- Concludes the fix is to start enormously high
- Treats the count as an ordinary editable setting
- Believes a raise leaves key placement untouched