A reader is frozen on one record it cannot handle: what are the exits, and what does each one cost in correctness?
answer
- four exits, four different bills
- speed is paid in correctness
- capture before you discard
- who finds out, and when
- capacity is not an exit here
basics
~20 sFour exits: fix the handler, set the record aside, discard it, or wait for a transient cause to clear. Fixing loses nothing but is slowest; discarding is fastest and loses the record silently. Adding capacity is not an exit.
solid answer
~50 sRank them by what they cost, not by how fast they are. **Fixing the handler and redeploying** loses nothing, but it is the slowest, and while you write it the backlog behind the record ages towards the edge of the retained window. **Setting the record aside** to a side destination restores flow in minutes and keeps the record, at the price of taking it out of order and of needing a named owner, or it is simply a slower way of losing it. **Discarding it** — advancing past it, or acknowledging it away where there is no position — restores flow in seconds and loses that record silently, so capture its identity and content first. **Waiting** is legitimate when the cause is a dependency that will return, and it costs only the growing backlog. What is *not* an exit is adding capacity: a frozen share has no rate to raise, and the extra readers merely make the other shares drain faster, which flatters the aggregate chart.
go deeper
Recall that unblocking a frozen reader means the record is either fixed, kept elsewhere, discarded or waited out, and that the fast options cost correctness rather than being free.
Explain each exit's mechanism and immediate consequence: what advancing past a record actually records, why a set-aside record is out of order, and why a fixed handler sees that record again.
Show the sequencing you would actually run — capture, restore flow with the cheapest reversible option, assign an owner to the leftovers, fix the handler — and reject adding capacity for a share with no completion rate.
Argue about who pays: the loss lands on a downstream consumer at a later reconciliation, so the acceptable exit is a property of what the stream carries and should be settled before an incident, not during one.
## There is no free exit A member that is holding an assignment and completing nothing puts an operator in front of a choice where every option is paid for in a different currency: time, ordering, correctness, or somebody else's future confusion. The purpose of knowing the menu in advance is that the currency is chosen deliberately rather than by whoever is awake. The backlog behind the frozen record is also a clock. Records have a finite **retained window**, so "decide tomorrow" is itself a decision that can lose far more than the one record. ## The four exits 1. **Fix the handler and redeploy.** The only exit that loses nothing. It is also the slowest: diagnosis, a change, a review, a deployment. Two caveats are worth volunteering — the backlog grows the whole time, and once the fix is live the record is processed again, so the handler must tolerate a second copy of the same record being delivered. 2. **Set the record aside.** Move it to a side destination so the rest of the share proceeds. Flow is restored in minutes and the record still exists. It costs ordering — that record is now out of sequence relative to everything around it — and it costs an owner, because a side destination nobody reads is a slower discard with better paperwork. 3. **Discard it.** Advance the recorded read position past it, or acknowledge it away where there is no position to move. Fastest restoration available and an unconditional loss of that record's work. The loss is silent: nothing downstream is told that a record is missing. 4. **Wait.** When the cause is a dependency that is down or a lock that will clear, the member will resume on its own. This costs only the growing backlog — as long as the assumption that the cause is transient is actually tested rather than hoped for. ## The bill for each | exit | time to restore flow | what it costs | who finds out, and when | |---|---|---|---| | fix the handler | hours | the backlog grows meanwhile; the record is reprocessed | nobody, if the handler tolerates a repeat delivery | | set aside | minutes | ordering for that record; needs a named owner | whoever replays it, applying an old record over newer state | | discard | seconds | the record's work, irrecoverably | a downstream consumer of the derived state, at a later reconciliation | | wait | unbounded | backlog ageing towards the retained window | everyone, if the cause was not transient after all | ## The action that is not an exit Adding readers, enlarging batches or making the handler faster all raise a completion rate. A frozen share has a rate of zero, and none of those multiply it into something else. Worse, the intervention appears to help: the other shares of the stream drain faster, the group's aggregate gap bends downwards, and the frozen share is untouched. This is the single most common wrong move on this failure, and an interviewer is often specifically listening for whether the candidate reaches for it. ## Sequencing under pressure 1. **Capture before you discard.** Whatever you are about to destroy, write its identity and content somewhere durable first. A discard you can describe afterwards is an incident; one you cannot is a mystery for the rest of the system's life. 2. **Restore flow with the cheapest reversible option available.** Setting aside beats discarding whenever the platform and the time allow it, because it preserves the ability to change your mind. 3. **Give the leftovers an owner before the incident closes.** The record set aside is not handled; it is deferred, and deferral without a name is loss. 4. **Fix the handler anyway.** Every fast exit treats the symptom. If one record could do this, another will. 5. **Decide about replay explicitly.** Reapplying an old record over state that has since moved on can be worse than the gap it fills. That decision belongs to whoever owns the downstream meaning, not to the operator at 03:00. ## What varies between platforms - Where the reader owns a recorded position, discarding means moving that point forward past the record, and the skip is visible as a gap if anyone looks for it. - Where records are removed on acknowledgement, discarding means acknowledging without doing the work, and there is no gap left behind to notice — capture is therefore even more important. - Where a design offers no way to move a record aside at all, the menu collapses to fix, discard, or wait, which is worth knowing before the incident rather than during it. - Whether the record is reprocessed after a fix depends on which exit was taken: fixing and waiting both reprocess it, discarding does not, and setting it aside reprocesses it only if somebody replays it.
- Why is capturing the record before discarding it more than good manners?Because the loss is otherwise undescribable. A discarded record leaves no downstream notification and, on designs that remove records on acknowledgement, no gap either. Without a captured copy nobody can answer the question that arrives weeks later — which entity is missing, and how much work must be reconstructed. The capture converts an unbounded correctness problem into a bounded, repairable one.
- When is waiting the right exit rather than the lazy one?When the cause is demonstrably external and transient — a dependency that is down and being recovered, a lock held by a process that is being restarted — and when the oldest unread record on the share is far from the edge of the retained window. State both conditions explicitly. Waiting on an untested assumption converts one blocked record into the loss of everything behind it.
- Does replaying what was set aside always restore correctness?No. An old record applied over state that has since advanced can be worse than the gap it fills, particularly where the record carries an absolute value or a position in a sequence. Replay is a decision for whoever owns the downstream meaning, and it should check whether the record is still meaningful before resending it, not merely whether it can be resent.
saying these in an interview costs you the question
- Discards the record without capturing what was thrown away
- Adds readers to a share that is completing nothing at all
- Calls setting a record aside a fix with nobody owning the destination
- Believes a discarded record can be recovered after the retained window passes
- Forgets that a fixed handler will meet that record a second time
- Presents the four exits as equally safe choices