Why does adding a reader to a group pause work on some brokers and cost nothing on others?
answer
- assigned versus contested
- one holder per part
- nothing owned, nothing reassigned
- the cap is the number of parts
- no pause is not no cost
basics
~20 sBecause only one of the two shapes assigns anything. Where a stream is split into parts and each part has one holder, a new member forces the division to be recomputed. Where readers compete for records from a shared queue, nothing is held, so nothing is reassigned.
solid answer
~50 sIt comes down to whether responsibility is assigned or contested. On a split-stream design the stream is divided into parts, each part is held by exactly one member of the reader group, and adding a member means taking parts from existing holders — a recomputation that stops the affected members, sometimes the whole group, until it settles. On a competing-reader design the readers all pull from the same waiting work; nothing is owned, so a newcomer just starts pulling and no reassignment or pause exists to be triggered. That is not the same as being free of every effect: the newcomer contends for the same records and the same downstream dependencies. The practical consequence is that safe operating procedures differ. On the first shape, churning membership during an incident has a real cost; on the second, adding readers is the cheapest lever available.
go deeper
Hold on to the distinction: some platforms assign each part of a stream to one reader, others let readers compete for whatever is waiting, and only the first has anything to reassign.
Explain why coupling membership to a division forces a handover and therefore a pause, and why contested work has no equivalent event at all.
Show the incident consequence: on assigned shares, know the cap and add members in one batch; on contested work, adding readers is the cheap first lever.
Make the shape an explicit input to platform choice and to runbooks, so nobody carries the reflexes of one shape into an estate that runs the other.
## Two ways to divide work among readers Platforms in this class solve the same problem — several processes reading one stream without doing the same record twice — in two structurally different ways, and almost every operational difference in reassignment behaviour follows from which one you are on. **Assigned shares (split-stream).** The stream is divided into parts. Each part is held by exactly one member of the **reader group** at a time, and each part carries its own recorded **read position**. Membership and division are coupled: change the membership and the division must be recomputed. Because a part may have only one holder, the recomputation requires a handover, and a handover requires a stop. **Contested work (competing readers).** There is one body of waiting work and every reader pulls from it. No reader owns anything; a record is handed to whichever reader asks next, and it stops being available to the others while it is held. A reader that dies mid-record simply never acknowledges it, and the record becomes available again once the window before redelivery expires. ## What a join costs in each | Aspect | Assigned shares | Contested work | |---|---|---| | Effect of a new member | Division recomputed; shares taken from existing holders | Newcomer starts pulling; nothing is recomputed | | Pause on join | Yes — the affected members, and on many designs the whole group | None of this kind | | Effect of a member vanishing | Its share is unread until a timer fires and the division moves | Its held records return to the pool after the redelivery window | | Useful parallelism | Capped by the number of parts; extra members sit idle | No structural cap of this kind | | Cheapest reaction to a backlog | Add members **up to the cap**, accepting the pause | Add readers freely | The row that catches people out is the fourth. On the assigned shape, a member beyond the number of parts receives nothing at all — it joins, it forces a reassignment, and then it idles. So a panicked scale-out during an incident can buy a pause and zero extra throughput. On the contested shape, there is no equivalent structural cap, and a newcomer starts contributing immediately. ## Which shape am I on? Three questions settle it without naming anything: 1. **Is there a recorded position you could move backwards?** A per-part read position that a reader resumes from implies assigned shares. Work that disappears once acknowledged implies contested work. 2. **Does a membership chart show pauses?** If starting an instance produces a visible dip in processed records across other instances, the division is being recomputed. 3. **Does adding readers past a certain number stop helping, sharply rather than gradually?** A hard ceiling is the number of parts; gradual diminishing returns is ordinary contention downstream. ## Why this is not the same as 'free' A join that triggers no reassignment still has effects, and claiming otherwise is the overcorrection to avoid: - The newcomer contends for the same records and the same downstream systems, so the bottleneck may simply move. - More readers holding records at once means more work in flight, and more of it to redo if several readers fail together. - Connection and request budgets on the cluster are consumed per reader whatever the shape. What is genuinely absent is the **reassignment pause** and everything downstream of it: no handover, no stop, no members idling while the division settles. ## Why it matters in an incident The two shapes justify different reflexes, and the reflex from the wrong shape is actively harmful: - On **contested work**, adding readers is close to free and is usually the first thing to try against a growing backlog of unread records. - On **assigned shares**, every add is a membership change. Adding them one at a time produces one pause each; adding them together produces one pause total. Knowing the cap before you start is the difference between buying throughput and buying pauses. - On either shape, removing readers in an announced way is cheaper than killing them, because an inferred departure costs a timeout of stalled work that an announced one does not. An engineer who has only operated one of the two shapes will state its behaviour as though it were the law of the class. It is not, and interviewers on this subject are frequently probing for exactly that assumption.
- On a design with assigned shares, why can adding members during an incident buy a pause and no extra throughput?Useful parallelism is capped by the number of parts the stream is divided into. A member beyond that cap receives no share, so it contributes nothing — but joining still forced the division to be recomputed, which stopped the members that were working. You paid the pause and got no reader.
- If you must add several readers to a group with assigned shares, how do you minimise the damage?Start them together rather than one at a time, so the membership settles in a single recomputation instead of one per instance, and know the cap first so you do not add members that will idle. Removing readers later should be announced rather than by kill.
- How do you tell quickly which of the two shapes a platform gives you?Ask whether a reader has a recorded position it resumes from, and whether starting an instance visibly dips the other instances' processing. A resumable per-part position plus a dip on join means assigned shares; work that vanishes on acknowledgement and no dip means readers are competing.
saying these in an interview costs you the question
- States one shape's behaviour as the law of all brokers
- Thinks every platform pauses when a reader joins
- Believes extra readers always add throughput
- Calls a join with no reassignment completely free
- Adds readers one at a time on an assigned-share design
- Expects to rewind a position on a destructive-read design