A reader group is reassigned every ninety seconds and pauses about twenty each time, with no deploys running. What is happening?
answer
- a rhythm means a timer
- alive but past the deadline
- rejoining triggers another reassignment
- same records attempted repeatedly
- shorten the cycle or lengthen the deadline
basics
~20 sHandling is slower than the progress deadline, so a live member is declared gone; its share is reassigned, the group pauses, the member rejoins and starts the same slow work, and the cycle repeats. The group loses most of its time to the loop, not to the work.
solid answer
~50 sThis is a self-feeding reassignment loop, and the giveaway is the rhythm: a cadence nobody caused is a timer, not a crash. A member takes longer over one unit of work than the progress deadline allows, so it is declared gone even though the process is alive and its liveness signal is still arriving. Its share is reassigned, part or all of the reader group pauses while shares settle, the member rejoins and is handed work again — and, being no faster than before, it blows the deadline again. Meanwhile every restarted unit of work resumes from the last recorded read position, so the same records are attempted repeatedly and little net progress is made. The two honest levers are the handler's cycle time and the deadline itself; restarting the members does not help, because the work is exactly as slow when it comes back.
go deeper
Recall that a member can be declared gone while it is still running, because a separate timer limits how long it may hold work without showing progress.
Walk the loop end to end — deadline expires, share moves, group pauses, member rejoins, reassigned again — and say why the interval matching the deadline is the tell.
Diagnose from the rhythm and the staircase in the unread count, then choose between shortening the cycle and raising the deadline, naming the detection cost of the second.
Set the estate rule: deadlines chosen against a measured worst case rather than a default, and an explicit position on how long an undetected dead member is tolerable.
## The loop, step by step The pattern has a fixed shape, and once you recognise it the diagnosis is quick: 1. A member is handed a share and begins working through records. 2. One unit of work takes longer than the **progress deadline** — the timer that bounds how long a member may hold work without showing that it is still moving. 3. The deadline expires. The member is declared gone, even though its process is alive and its **liveness signal** is still arriving on time. 4. Its share is redistributed. Part or all of the reader group pauses while the division settles. 5. The member finishes, discovers it no longer holds the share, and rejoins. 6. Rejoining is itself a membership change, so the group is reassigned again. 7. The member is handed work, is exactly as slow as before, and step 2 repeats. The cadence is the fingerprint. Crashes and deploys are irregular; a timer produces a rhythm. A group reassigning on an interval that nobody scheduled is reporting a timer expiring over and over. ``` t+0s member holds share, starts a slow unit of work t+60s progress deadline expires -> member declared gone t+60s reassignment begins -> group pauses t+80s shares settled -> reading resumes t+85s slow unit finishes; member finds its share gone; rejoins t+85s reassignment begins again -> group pauses again t+90s member handed work, equally slow -> back to t+0 ``` ## Why a healthy member gets declared gone Most designs run **two independent timers**, and conflating them is why this looks impossible at first: | Timer | What it proves | What its expiry means | |---|---|---| | Liveness signal | The process exists and can still talk to the cluster | Crash, host loss, or a broken network path | | Progress deadline | The member is still working through what it holds, not wedged | The member is alive but has held work too long without progressing | A handler that makes one slow call per record, or that occasionally hits a pathological record, satisfies the first timer perfectly while failing the second. The system's belief that the member is gone is not wrong by its own definition — the member really has held a share for longer than the group is willing to wait — but it is wrong about the cause, and the remedy it applies makes things worse rather than better. ## Reading the numbers you have - **Reassignment interval roughly equal to the deadline.** Strong evidence that the deadline is what is firing, not the liveness timer, which is shorter. - **Unread count rising in a staircase.** It climbs through every pause and barely falls between them, because each cycle leaves less productive time than the last. - **The same records attempted repeatedly.** Every restarted unit of work resumes from the last recorded read position, so work done but not recorded before the deadline expired is done again by whoever inherits the share. - **Members alive throughout.** No restarts, no crash traces, no host events — which is what separates this from a genuinely crashing instance. - **Adding members makes it worse, not better.** Each newcomer is another membership change and another pause, and if the slowness is per record the new members are equally slow. ## Breaking the loop There are only two honest levers, and choosing between them is the judgment being tested: 1. **Make the cycle finish inside the deadline.** Reduce how much a member takes on before it next reports progress, or move the slow step out of the path that the deadline measures. This is right when the slowness is incidental — a chatty dependency, an unbatched lookup, an accidental serial loop. 2. **Raise the deadline to exceed the real worst case.** This is right when the work genuinely takes that long and nothing about it is wasteful. The price is honest and should be stated: a member that truly dies while holding a share is now undetected for longer, so its share sits unread for the new, longer timeout. What does **not** work: restarting the members, which hands the same slow work back to a fresh process; or adding capacity, which adds membership churn to a group already spending its time churning. Stabilise the loop first, then scale. ## What this is not Two neighbouring incidents look similar on a chart and are not this one. A member that **keeps** its share and simply makes no progress is a different diagnosis — nothing is being reassigned there, and the rhythm is absent. And a **broker node** leaving or returning moves data between cluster nodes; that is a cluster-membership event, not a reader-group one, and no amount of tuning a reader's deadline will address it. Establish which of the three you are looking at before you change anything, because the fixes are mutually useless.
- Why does the group's liveness signal arriving normally not rule out this diagnosis?Because two independent timers are running. The liveness signal only proves the process exists and can talk to the cluster; the progress deadline separately bounds how long it may hold work without progressing. A healthy, responsive member that is slow satisfies the first and fails the second.
- What is the real cost of simply raising the progress deadline until the loop stops?Detection of a genuinely dead member gets slower by exactly the same amount. Its share then sits unread for the longer timeout before anyone inherits it. That is an acceptable trade when the work honestly takes that long, and a bad one when the deadline is masking waste in the handler.
- Why does adding more members during a reassignment loop usually make it worse?Each newcomer is itself a membership change, so it buys another pause before it does any work. If the slowness is per record, the new members hit the same deadline as the existing ones, adding churn to a group that is already spending most of its time settling rather than reading.
- How would you separate this from a member that is simply wedged?Look for the rhythm and for who holds what. A looping group reassigns on a cadence close to the deadline and its membership keeps changing. A wedged member keeps its share throughout and nothing is being redistributed — the shares are stable and one of them is simply not advancing.
A relay team where a runner is scratched if he has not reached the handover line by a fixed time. He is not injured and not missing — he is simply slower than the rule allows. The team restarts the leg with a fresh runner who is no faster, scratches him too, and the baton never advances. Either the leg is shortened or the rule is relaxed; swapping runners changes nothing.
saying these in an interview costs you the question
- Assumes the members must be crashing
- Restarts the readers to clear the loop
- Adds members to a group already churning
- Thinks a live liveness signal rules out being declared gone
- Raises the deadline without naming the detection cost
- Confuses it with a broker node leaving the cluster