Before rewinding a reader group's read position by six hours to reprocess, what side effects must an operator account for first?
answer
- everything downstream fires again
- converging effects against escaping ones
- a handler cannot unsend a notification
- republished records reach other groups too
- disable the outbound leg before restarting
basics
~10 sEvery effect downstream of the new position fires again — including the ones that leave the system, such as customer email, payments, third-party calls and records republished onward, which no reader can take back.
solid answer
~50 sA backward move buys reprocessing of a span, and it buys every consequence of that span a second time. The effects that matter are the ones that escape the service boundary: notifications already sent again, payments or refunds re-attempted, calls into a partner's system, and records republished onto another stream where other teams' reader groups will see them as fresh work. A handler that absorbs a repeat safely does not help here — it can make its own database converge, but it cannot unsend an email. So the assessment before the move is an inventory: what sits downstream of the new position, which of it escapes, and which of those can be disabled, throttled or quarantined for the duration. It is also a load question, because the rewound span is read as fast as the readers can go, so every downstream dependency sees the spike at once.
go deeper
Hold on to the headline: rewinding a reader group's read position makes the work happen again, so anything the service already sent outward for that span goes out a second time.
Explain the split between effects that converge on the same state when repeated and effects that leave the system, and why only the first class is rescued by writing the handler to tolerate a repeat.
Show the pre-move sequence: inventory what is downstream of the new position, bound the span, disable the outbound leg, throttle the re-run, and warn the owners of anything fed by records this group republishes.
Frame it as a trade the organisation signs: the value of the corrected output against the duplicate external effects, who carries the customer-visible residue, and whether the estate should carry a standing way to reprocess with the escaping legs disabled.
## The move is cheap; the consequences are not Moving a reader group's recorded read position backward is a single administrative write. What it triggers is a full re-run of everything the group does for the span between the new position and where it was — at whatever rate the readers can manage, which is usually far faster than the rate the records originally arrived at. The question an operator has to answer before the move is therefore not "can we rewind" but **"what fires again, and which of it escapes?"** ## The inventory: three classes of effect | class | examples | what a rewind costs | |---|---|---| | **Convergent, inside the service** | a row written by key, a value set to the latest seen, a cache entry rewritten | usually nothing but CPU: the second run lands on the same state | | **Accumulating, inside the service** | counters that add, rows appended to a history, totals summed | silent corruption: the span is counted twice unless the accumulation is keyed against a repeat | | **Escaping the service** | email, SMS and push, payment capture or refund, calls into a partner, physical fulfilment, records published onward | not recoverable by the readers at all: the effect has already left | The third row is the one that turns a rewind into an incident, and it is the class that a well-written handler cannot rescue. Making the handler absorb a repeat is a real and separate discipline, and it operates on state the service owns. It has no reach over a message that has already been delivered to a customer or an instruction already accepted by a payment provider. ## The effect that is most often forgotten A reader that **publishes records onward** — enriching, splitting, or routing them to another stream — makes every downstream reader group a participant in your rewind without being asked. They receive the re-published span as ordinary new work, and their own escaping effects fire too. One team's six-hour rewind becomes four teams' six-hour rewind, discovered by the others from their own incident channels. This is why the inventory has to follow the records, not the service: enumerate what is downstream of the new position across the boundary, and tell the owners before the readers restart, not after. ## Containment, in the order it is decided 1. **Bound the span.** Rewind to the narrowest range that fixes the problem, not to a comfortable round number. Every extra hour is extra duplicate effect for no benefit. 2. **Disable the escaping leg.** Turn off the outbound path — the notification sender, the payment call, the partner client — for the duration, or point it at a sink. This is an operational change made before the readers restart, and it is the single most effective control available. 3. **Throttle the re-run.** The rewound span arrives at the downstream dependencies as a burst. A partner's rate limit that is comfortable at the natural arrival rate will be exceeded by a replay running an order of magnitude faster, and rejections during a recovery are their own incident. 4. **Warn the neighbours.** Anyone whose reader group sits downstream of a republished stream, and anyone who owns an accumulating total, needs to know the span and the time window. 5. **Agree who owns the residue.** Duplicate customer-visible effects that escape containment need an owner and a story before the move, not an apology afterwards. ## Making the decision honestly A rewind is worth its cost when the missing or wrong output is worth more than the duplicate effects it will produce — a derived total that is wrong for six hours, a batch of records dropped by a broken handler. It is a poor trade when the only convergent part is cheap and the escaping part is large: re-sending a day of notifications to fix a reporting number is a worse outcome than the number. Where the effects are heavily escaping, the alternative is to reprocess **without the live path**: give the work a separate reader group with the escaping legs disabled, and reconcile the difference by hand. That keeps the live group's position untouched and turns a broadcast into a measurement. ## Where the move is not available All of this assumes a design that keeps records after they are read and keeps a position that can be written. Where records are deleted once acknowledged, there is nothing to rewind into: the only source of a second pass is a copy written elsewhere when the records were first produced, which is a decision made long before the incident. ## What an interviewer is listening for Not "make the consumer idempotent" — that answers a different question and is assumed here. The signal is whether the candidate separates effects that converge from effects that escape, remembers the republished stream, and treats disabling the outbound leg and bounding the span as steps taken **before** the position is written rather than as mitigations discovered afterwards.
- If the handler is written to absorb a repeat safely, is a rewind then free?No. That discipline governs state the service owns, so it makes the internal result converge. It has no reach over effects that already left — a notification delivered, a payment instruction accepted, a record republished to another team's reader group. Those fire again regardless.
- Why is the speed of the re-run a risk in its own right?The rewound span is read as fast as the readers can go, not at the rate it originally arrived. Downstream dependencies and rate-limited partners see hours of traffic compressed into minutes, so a recovery can trigger rejections, throttling or a dependency outage that the original load never would.
- How do you reprocess when the escaping effects cannot be disabled at all?Do not move the live group's position. Run the span through a separate reader group with the escaping legs routed to a sink, compare its output against the live state, and correct the difference deliberately. That converts an uncontrollable broadcast into a measurement plus a small, reviewed repair.
- Does rewinding one reader group move any other group's position?No — each reader group keeps its own recorded position, so the others stay where they are. They are affected only indirectly: if the rewound group republishes records onward, the downstream groups receive that span again as new work.
saying these in an interview costs you the question
- Assumes a safely repeatable handler makes a rewind harmless
- Forgets that republished records reach other teams' reader groups
- Treats counters that add as safe to reprocess
- Rewinds to a round number instead of the narrowest span
- Ignores that the re-run arrives far faster than the original traffic
- Believes rewinding one group moves the other groups' positions