skip to content

Why is every member of a reader group stopped before its recorded read position is moved, and what breaks if one keeps running?

level: middleimportance: should knowfreq 52%

answer

  1. two copies of one number
  2. in-memory place versus stored place
  3. a live member writes its own value back
  4. capture the prior position per share first

basics

~20 s

A running member holds its own place in memory and writes it back, so it overwrites the move or carries on past it. Stopping the whole group first makes the stored position the only writer of record.

solid answer

~50 s

While a reader group is live, the stored position is not the only copy of the truth: each member also holds its current place in memory and periodically writes it back, and again when it shuts down. Change the stored value underneath a running member and one of two things happens — the member's next write-back replaces your value with its own, or the member simply keeps reading from where it already was and the move never takes effect for its shares. Platforms differ in how much they protect you: some refuse the change outright while the group has live members, and some accept it silently, which is worse, because the group looks moved and then drifts back. Stopping every member first removes the competing writer. Before touching anything, record the current position of each share, the time and the unread count: that is the only way back and the only measurement of what changed.

code

json · 12 lines
json
{
  "stream": "orders-events",
  "readerGroup": "billing-writer",
  "capturedAt": "2026-09-20T02:41:07Z",
  "reason": "rewind to reprocess after bad deploy",
  "unreadCountAtCapture": 18422,
  "shares": [
    { "share": 0, "priorPosition": 84420331, "newPosition": 84102998 },
    { "share": 1, "priorPosition": 84418907, "newPosition": 84101554 },
    { "share": 2, "priorPosition": 84421756, "newPosition": 84104012 }
  ]
}

go deeper

for a junior

Remember the rule and the reason in one line: stop every member of the reader group first, because a running member keeps its own place in memory and will write it back over whatever you stored.

for a middle

Explain the two copies of the position and the three failure shapes they produce — the move overwritten, the move ignored, and the group left split across two places — and say why capturing the prior position is part of the move, not paperwork after it.

for a senior

Demonstrate the procedure as a runbook with a verification step: capture per share, stop, move, read back, restart and watch the curve's direction, and state that platforms differ in whether they refuse the change while the group is live.

for a principal

Decide where this discipline is enforced rather than remembered: who may perform a move, what must be captured before it, and whether the estate's positions live somewhere an operator can edit with no interlock at all.

## Two copies of one number A reader group's **read position** looks like a single stored value, but while the group is running there are really two copies: the one in the **position store** (broker-side, a side table, or wherever the design keeps it) and the one each member holds **in memory** as it works. The in-memory copy is authoritative for that member's behaviour; the stored copy is authoritative only for a member that is *starting*. Members reconcile the two by writing their in-memory value back — periodically, at a reassignment, and usually on a clean shutdown. Every hazard in this operation follows from those two copies existing at once. ## What a live member does to your move If an operator writes a new position while a member still holds the affected shares: - **The move is overwritten.** The member's next write-back stores the place it has actually reached, and the administrative value disappears. The group looks moved for a few seconds and then reverts — a symptom teams often misdiagnose as "the change did not apply". - **The move is ignored.** The member never re-reads the stored value while it is running, so it keeps reading from its in-memory place regardless. This is the more dangerous of the two, because a rewind that was meant to reprocess simply does not, and nobody notices until the missing output is reported. - **The move is applied to part of the group.** Shares held by stopped members take the new value; shares held by live members do not. The group is left split across two places in the stream, and the resulting output is neither the old behaviour nor the intended one. - **A reassignment lands in the middle of it.** If members are stopped one at a time, their shares are handed to the survivors as they go, and whichever member picks a share up next reads whatever value happens to be stored at that instant. The result depends on timing. Platforms vary in how much of this they let you do. **Some refuse the change while the group has live members**, which turns a silent data problem into a loud error message. **Others accept it and let the running member overwrite it.** Never assume you are on the protective kind; the procedure below is safe on both. ## Record the prior position before anything is touched The move overwrites the old value. There is generally no undo history, so once it is gone the answers to two separate questions are gone with it: 1. **How do we go back?** A rewind to the wrong place, or a forward move to a point that turns out to be too far, is recoverable only if the original number exists somewhere outside the position store. 2. **What exactly changed?** The span between the prior position and the new one is the span that was skipped or that will be reprocessed. That is the number the rest of the organisation needs — for the team whose output has a hole in it, for the people who will see duplicate effects, for the incident record. Record it **per share**, not as one number: on designs that split a stream into parts, each part has its own position and they are rarely at the same place. Capture the wall-clock time and the unread count too, because they let you say how far behind the group was, not only where it was. ## The procedure 1. **Capture** the current position of every share, plus the time and the unread count. 2. **Stop** every member of the group and confirm no member still holds a share — not "scale to one", not "pause the handler": stopped. 3. **Apply** the move, forward or backward, to every affected share. 4. **Read back** the stored positions and confirm they are the values you asked for. 5. **Restart** the readers and watch the gap move in the intended direction for the first minutes. Step 4 exists because steps 2 and 3 are the two that fail quietly. Step 5 exists because a group that comes back at the old place looks identical to one that came back at the new place, until you look at the direction the curve moves. ## Where the guard is weaker Where the position is not kept by the broker at all — held in a side table the application owns, or by the reader itself — nothing validates the move and nothing refuses it. The same discipline applies with more force: the operator is editing live state with no interlock, and "stop the readers first" is the only protection that exists. And where nothing stores a position, this procedure has no meaning: there is no value to capture and none to move, and the only forward-like action is discarding the waiting records.

  • Is pausing the handler enough, or must the members be stopped?
    Stopped. A paused member is still a live member: it still holds its shares, still proves it is present, and will still write its in-memory place back when it resumes or exits. Only a member that has released its shares has stopped competing for the stored value.
  • Why record the position per share rather than one number for the group?
    Where a stream is split into parts, each part carries its own position and they sit at different places, because the parts are read at different rates. One aggregate number cannot be moved back to, and it hides which part was furthest behind.
  • How do you confirm the move actually took effect?
    Read the stored positions back before restarting anything, and compare them with the values you asked for. Then watch the gap for the first minutes after restart: a group that silently kept its old place produces the old curve, which is the only visible difference.

saying these in an interview costs you the question

  • Thinks the stored position is the only copy while readers run
  • Believes every platform refuses the move while members are live
  • Pauses the handler instead of stopping the members
  • Skips recording the prior position because the move is 'obvious'
  • Captures one position for the group instead of one per share
  • Restarts the readers without reading the stored values back