Why does a reader group stop making progress while its shares are being reassigned, and how much of it stops?
answer
- one holder per share
- stop, record, hand over
- group-wide or only the movers
- members stopped times seconds
- writers never pause
basics
~20 sA share may be held by only one member at a time, so the old holder must stop before the new one starts. How much of the reader group stops depends on the design: many stop every member until the division settles, others stop only the shares that move.
solid answer
~50 sThe rule that makes a reassignment pause anything is that a share has exactly one holder. Before the division can be recomputed safely, whoever holds a moving share has to stop reading it and record its read position, so the next holder can resume from a known point rather than duplicating or skipping work. The question is how widely that stop is applied. On many designs it is group-wide: every member releases everything, the division is recomputed, and everyone picks up a new share — so a member whose share did not change still stopped. Other designs hand over only the shares that actually move and leave the rest reading throughout. The cost of a pause is therefore the number of members stopped times its duration, plus the backlog of unread records that arrived meanwhile, which the group must then work off.
go deeper
Remember the rule that makes any of this necessary: one holder per share, so the old holder must stop and record its place before a new one can start reading it.
Explain both behaviours — stopping every member versus handing over only the shares that move — and be able to say which one your platform has and why that changes the cost.
Price a pause in members times seconds plus the backlog it leaves, and show that recovery is gated by the margin between read rate and arrival rate, not by the pause length.
Decide what churn the estate tolerates before it counts as an incident, and make the group-wide-versus-partial behaviour an explicit input when choosing how instances are deployed and scaled.
## Why anything has to stop at all The constraint behind the whole behaviour is simple: on a design that assigns shares, **a share has exactly one holder at a time**. If a share were being read by its old holder and its new holder simultaneously, the two would process the same records and record conflicting read positions. So the system enforces a handover: the old holder stops reading the moving share and records where it got to, and only then does the new holder begin, resuming from that recorded **read position**. That handover is not instant. It involves the members noticing the membership change, the division being recomputed, every affected member being told its new share, and each new holder fetching the recorded position it should resume from. During that interval, the affected shares are being read by nobody, and records keep arriving. The **reassignment pause** is that interval. ## How much stops: two behaviours, and you must know which you have Platforms in this class differ, and this is the single fact most worth establishing about the one in front of you. | Behaviour | What members do | Who pays | What it costs | |---|---|---|---| | Group-wide stop | Every member releases every share; the division is recomputed; everyone picks up again | All members, including those whose share is unchanged | Duration times the whole group's throughput | | Partial handover | Only the shares that move are released; other shares keep being read | Only the members losing or gaining a share | Duration times the moved shares' throughput | | No reassignment | Readers compete for records; nothing is held, nothing is handed over | Nobody | No pause of this kind exists | The group-wide behaviour is what surprises people. A group of twelve members where one instance restarts can stop all twelve — the eleven that keep exactly the shares they already had stopped anyway, because the division was recomputed from scratch and the safe way to do that is for everyone to let go first. On designs that hand over incrementally, the same event stops only the share that actually moved. ## What the pause actually costs The visible cost is not the pause itself but what it leaves behind: - **Stalled throughput.** Members stopped, times the seconds they were stopped, times their normal rate. Twelve members stopped for eight seconds is ninety-six member-seconds of lost reading, whatever the dashboard's per-member view suggests. - **A backlog that must be worked off.** Writers do not pause. Whatever arrived during the pause is added to the backlog of unread records, and the group must run above its arrival rate afterwards to get back to where it was — so a short pause on a stream near its capacity ceiling takes far longer than the pause itself to recover from. - **Repeated work at the seam.** The new holder resumes from the last recorded read position, so anything the old holder had processed but not yet recorded gets done again. How much that is depends on how often positions are recorded, which is a property of the reading side, not of the reassignment. - **Nothing about the records themselves.** No record is lost, deleted or reordered by a reassignment; the stream is untouched. Only responsibility moved. ## Where the time in a pause goes A pause that lasts a fraction of a second and one that lasts tens of seconds have different causes, and the difference is usually not the broker: 1. **Noticing.** If the trigger was a silent disappearance rather than an announced leave, the timer has to expire first. That part is often the bulk of the elapsed time. 2. **Agreeing.** Members must be told the new division and confirm it. Every member must reach this point, so the **slowest** member sets the pace — a member busy in a long unit of work does not respond until it surfaces. 3. **Resuming.** New holders read their starting positions and begin. Usually the cheapest step, unless positions are held somewhere slow. That second point is why a pause and a slow handler feed each other: the member everyone is waiting on is the one that is behind, so the group's recovery is gated by its worst member. ## The operator's takeaway Know which behaviour your platform has before you plan any change that churns membership. If the stop is group-wide, an instance flapping every few minutes is not a local problem — it is a tax on the entire group's throughput, and the arithmetic of members times seconds is the number to put in front of whoever owns the flapping instance.
- A reader group of twelve pauses for eight seconds when one instance restarts, though eleven keep the same shares. Why did they stop?Because the division was recomputed from scratch, and on designs with a group-wide stop the safe way to do that is for every member to release everything first and pick up afterwards. The eleven unchanged members stopped for correctness of the recomputation, not because their own work moved.
- Why does a short pause on a busy stream take much longer than the pause itself to recover from?Writers keep writing throughout, so the pause adds a backlog of unread records. Clearing it requires reading faster than the arrival rate, and the spare capacity above that rate is usually small — so recovery time is the pause's backlog divided by that narrow margin, not the pause's duration.
- Does a reassignment risk losing records?No. The stream is untouched; only responsibility for reading it moves. The realistic seam is the opposite one: the new holder resumes from the last recorded read position, so work done but not yet recorded by the previous holder is done again.
A shift handover where the rule is that only one person may hold the keys. On some sites every key on the board is surrendered and reissued, so staff whose assignment did not change still stand idle through the handover; on others only the one contested key changes hands and everyone else keeps working. The work arriving at the door does not pause for either.
saying these in an interview costs you the question
- Assumes only the members losing a share stop
- Thinks a reassignment can lose or delete records
- Measures the cost as the pause duration alone
- Believes writers stop while the group settles
- Claims every platform pauses the whole group
- Ignores that the slowest member paces the handover