A pipeline consuming a live shared feed of send-requests fails and retries by resubscribing; why does the request that failed not come back?
answer
- a new subscription is not a new run
- the source never stopped
- retry buys the future, not the past
- the gap loses more than one value
- a per-subscription test source hides it
basics
~20 sResubscribing attaches a new subscription to a source that never stopped running, so it delivers only what arrives from that moment on. Whatever the feed emitted before the new subscription attached, including the failed request, is gone.
solid answer
~50 sRetry replays work only when the source **re-runs its work for each subscription**. A live shared feed is already producing, independently of who is attached; a new subscription is a new listener, not a new run. So resubscribing hands the pipeline the future of the feed, never its past. Worse than losing the one request that failed, there is a **gap**: between the failure tearing the subscription down and the new one attaching, nothing is listening, and every request that arrives in that window is lost silently — no failure, no signal, just absent work. The fix is not a larger attempt limit. Keep the failure away from the feed subscription by retrying the per-request work inside the pipeline, so the outer subscription is never released; and if genuine replay is required, the record of unfinished requests has to live somewhere outside the subscription.
code
pseudocode · 12 lines// outer retry: the feed subscription is released and re-established
live_requests // already running, shared
.flat_map(deliver) // fails
.retry(max_attempts = 3)
// requests arriving while nothing is attached are never seen
// inner retry: the feed subscription is never released
live_requests
.flat_map(function(request) {
return deliver(request).retry(max_attempts = 3)
})
// the failure is contained to one request; the feed keeps flowinggo deeper
Hold on to the core fact: subscribing again to a source that never stopped gives you what it emits from now on, not what it emitted before. Retry cannot rewind a live feed.
Explain why the value that failed is not redelivered — delivery already happened and nothing retains it — and why the retry mechanism has no way to identify the new subscription with the old consumer.
Bring the gap and the silence: every value emitted while nothing was attached is lost without a signal, so the size of the loss scales with how long the retry waits. Say how you would detect it.
Rule on where recovery state is allowed to live. A retry stage is not durable, so decide what the platform requires of a pipeline over a live feed before it may claim at-least-once handling, and how that claim is reconciled.
## What resubscription can replay, and what it cannot A retry has exactly one move: subscribe again. Whether that recovers anything depends entirely on what a new subscription means to the source it attaches to. | | A source that runs its work per subscription | A live shared source | |---|---|---| | What a new subscription triggers | the work runs again, from the start | nothing; the source is already running | | What the new subscription receives | the full sequence, produced fresh | only what is emitted from now on | | What retry recovers | the failed work, genuinely re-attempted | nothing that was already emitted | | What the failure costs | time and repeated effects | every value emitted while detached | A feed of send-requests arriving from the outside world sits firmly in the right-hand column. It was producing before this pipeline attached and it keeps producing while the pipeline is away. The subscription is a tap on a flow, and closing and reopening the tap does not rewind the flow. ## The gap is wider than one request The intuition that loses the least money is this: the failed request is not the only casualty. A failure at the outer level tears the subscription down. The retry then re-establishes it — possibly after a deliberate pause between attempts. For the whole interval between those two moments **nothing is attached**, and everything the feed emits in that interval is delivered to nobody. One failed request becomes an unbounded number of unprocessed ones, scaling with how long the retry waits and how fast the feed runs. And the loss is silent. There is no failure signal for a value nobody was there to receive; the pipeline comes back up and looks healthy. The only evidence is a count that does not reconcile later. ## Why the failed request in particular is gone There is a specific misconception worth naming: that the failed value is held somewhere and redelivered to the next subscription. Nothing holds it. Delivery already happened — the value was handed downstream, and the pipeline failed while working on it. Retrying does not ask the source to re-send that value; it asks for a new subscription, and the source has no notion that this new subscriber is the same logical consumer as the old one, let alone which value that consumer was holding when it failed. ## Why this survives review The same pipeline, tested against a source that produces per subscription — a fixed list of requests, a source built for the test — genuinely does replay. The retry looks like it works. Every assertion passes. The behaviour changes only when the source is live and shared, and the change is not a failure but an absence, which is the hardest thing to write a test for. When someone says `retry is covered`, the question to ask is what the source under test does when a second subscriber attaches. ## What actually recovers a live feed 1. **Keep the failure away from the feed subscription.** Contain it to the per-request work, so the outer subscription to the feed is never released and no gap opens. This is the first and largest win, and it costs only a change of where the retry sits. 2. **Bound what an inner retry can hold up.** Containing the failure per request means a request can sit in its retry window for a while; decide whether the pipeline holds it, drops it, or hands it somewhere else after a deadline. 3. **Put the record of unfinished work outside the subscription.** If a request genuinely must survive a process restart or a torn-down subscription, the fact that it is unfinished has to be recorded somewhere that outlives the subscription. A retry stage is not durable state and was never claiming to be. 4. **Reconcile.** Because loss is silent, the only reliable detection is comparing what the feed produced against what the pipeline completed. ## The shape of the answer The mechanism is small: retry means resubscribe, and resubscribing to a source that is already running gives you the future rather than the past. Everything else follows — the gap, the silence, the false confidence from tests, and the fact that the fix is about **placement and durability**, not about attempt counts.
- How would you keep a retry over a live feed from losing anything?Stop the failure reaching the feed subscription: retry the per-request work inside, so the outer subscription is never released and no gap opens. That removes the loss the resubscription itself causes; anything that must survive a restart still needs its unfinished state recorded outside the subscription.
- Why does this defect usually get through review and testing?Because a test source normally re-runs per subscription, so the retry really does replay and every assertion passes. Against a live shared feed the loss is an absence rather than a failure — no signal, no error, just values nobody processed.
- Does adding a buffer between the feed and the pipeline solve it?Only if the buffer is on the far side of the subscription that gets torn down. A buffering stage inside the retried scope is discarded with everything else on resubscription, so it protects nothing; the holding has to survive the subscription's lifetime.
Reattaching to a live feed is like switching a radio back on: the broadcast carried on without you, and the minutes you missed are not waiting.
saying these in an interview costs you the question
- Expects resubscribing to replay values the source already emitted.
- Thinks the failed value is held and redelivered to the new subscription.
- Assumes only the failing value is lost, not the whole gap.
- Concludes from a passing test that retry replays a live feed.
- Treats a larger attempt limit as a fix for lost values.
- Calls a buffer inside the retried scope durable storage.