skip to content

When a notification pipeline that loads a recipient, records an attempt and calls a carrier is retried after the carrier fails, which steps run again?

level: juniorimportance: must knowfreq 55%

answer

  1. retry is not resume
  2. the subscription is remade
  3. everything above the retry point re-runs
  4. effects repeat, not just the call
  5. two attempt records for one send

basics

~20 s

Retrying resubscribes to the source, so every stage above the retry point runs again from the start: the recipient is loaded a second time, a second attempt record is written, and the carrier is called again.

solid answer

~40 s

A retry does not resume the sequence at the stage that failed. It drops the failed subscription and subscribes again to everything above it, so the whole retried scope re-runs from its source. In this dispatcher that means the recipient load, the message composition, the attempt record and the carrier call all happen a second time. Two of those are harmless — a read and a pure computation — but the attempt record is an effect on the outside world, so one logical send now leaves two records behind. The pipeline has no idea any stage was meant to happen once. Before attaching a retry, the question to answer is not `how many attempts` but `what exactly sits above this point`.

code

pseudocode · 12 lines
pseudocode
// assembled once; nothing runs until a subscriber attaches
pipeline = recipient_for(message_id)
    .map(compose_message)
    .effect(record_attempt)        // writes one attempt record
    .flat_map(call_carrier)        // fails on attempt 1
    .retry(max_attempts = 3)

// attempt 2 subscribes to recipient_for(message_id) again:
//   recipient loaded again
//   message composed again
//   a SECOND attempt record written
//   carrier called again

go deeper

for a junior

Remember the one-line fact: retrying a stream resubscribes to its source, so the work starts over rather than resuming. Be able to point at which stages of a small pipeline would therefore run a second time.

for a middle

Explain why no finer resumption is possible: a pipeline is a description with no checkpointed per-stage state, so the only recovery available is to run the description again. Classify each upstream stage as pure, read or effect.

for a senior

Show that you look at the retried scope before you add a retry, and name the duplicate the scope would produce in a real dispatcher. Mention that already-delivered values are not withdrawn when a multi-value sequence restarts.

for a principal

Frame it as a standard: where retries are allowed to sit in a pipeline, what a stage must guarantee to be inside a retried scope, and who owns the duplicate that escapes when that guarantee is only assumed.

## Retry is a resubscription, not a resumption A stream pipeline is a **description** of work, not the work itself. Nothing runs until something subscribes, and the subscription is what turns that description into a running sequence of stages. A retry stage lives under the same rule. When a failure signal reaches it, it cannot reach back into the stage that failed and try that stage alone — it holds no handle on that stage's half-finished state. What it holds is the description above it. So it releases the failed subscription and **subscribes again**, and everything above the retry point, inside the stream it is attached to, runs from the beginning. That one sentence accounts for nearly every surprise this mechanism produces: | What people expect | What resubscription actually does | |---|---| | execution resumes just after the failing stage | a fresh subscription starts at the top of the retried scope | | values already computed are reused | every stage above the retry point computes again | | only the remote call is repeated | every effect above the retry point happens again | | downstream sees one more value | downstream may see values it has already seen | ## Walking the dispatcher Take an outbound notification dispatcher with four stages: load the recipient, compose the message, record an attempt, hand the message to a carrier. The retry sits at the end, and the carrier fails. - **Load the recipient** — a read. It runs again. That is a second round trip, and what it returns is not guaranteed identical: something may have changed between attempts, so the second run is not necessarily a replay of the first. - **Compose the message** — a pure computation. It runs again and costs only time. - **Record the attempt** — an effect that writes. It runs again, so one logical send has now produced two attempt records. - **Call the carrier** — the stage the retry was actually for. It runs again, which is the point. Three of the four stages re-ran because they sat inside the retried scope; only one of them was meant to. ## Pure stages, reads and effects The useful move before attaching a retry is to classify every stage above the retry point: 1. **Pure stages** — computation over the value in hand. Repeating them costs latency and processor time, nothing more. 2. **Reads** — they change nothing outside, so repeating them is safe in that sense, but they add load and may answer differently on each attempt. 3. **Effects** — writes, sends, publishes, counters. Repeating one of these repeats it in the world, and this is where duplicates come from. The duplicate attempt record is not a defect in the retry mechanism. It is the mechanism doing exactly what it promises, over a scope nobody examined. ## Multi-value sequences repeat too A dispatcher with one message in flight hides a second consequence. If the retried scope produces several values and fails partway through, the values already delivered downstream stay delivered — resubscribing cannot un-emit them. The new subscription then starts the sequence from the top, so a consumer that had already seen the first three values sees them a second time. Any downstream stage that itself has an effect must therefore tolerate duplicates, or the retry has to sit somewhere that cannot produce them. ## Why the contract is built this way One could imagine a mechanism that checkpoints each stage and resumes after the failed one. Stream pipelines deliberately do not offer that: a stage keeps no addressable, resumable state, and the only thing guaranteed reproducible is the description itself. Resubscription is the one recovery a pipeline can always perform, precisely because it needs nothing except that description. The generality is paid for by the author, who has to know what running the description again means in the world. ## The practical rule Name every stage above the retry point. Mark each one pure, read or effect. For each effect, decide whether it is safe to repeat, safe once it is keyed so repeats converge on a single record, or whether it has to move out of the retried scope entirely. Then bound the attempts, so a failure that will never clear does not multiply that work without limit. A retry attached without that walk is not a resilience measure; it is a duplicate generator with a limit on it.

  • If every stage above the retry point is a pure computation, is there still an argument against retrying that whole scope?
    Yes, two of them. Repeating the work costs latency and processor time on every attempt, which matters when the scope is large. And a stage that only looks pure but actually reads something outside can answer differently on the second run, so the retried attempt is not the same attempt with a second chance.
  • What happens to values a multi-value sequence already delivered downstream when the retry restarts it?
    They stay delivered. Resubscribing starts a fresh sequence; it does not withdraw what the previous subscription already emitted. A consumer that saw the first three values sees them again after the restart, so any downstream effect has to tolerate duplicates.
  • Does the failed stage get any chance to clean up before the resubscription?
    Only whatever cleanup it registered for cancellation or failure. The retry releases the failed subscription, which is what triggers that cleanup; it does not roll back effects the stage already completed. A record already written stays written.

It is closer to restarting a recipe from the shopping trip than to picking up the pan where it burned: the ingredients are bought again whether or not you still had them.

saying these in an interview costs you the question

  • Thinks retry resumes at the failed stage, leaving earlier stages untouched.
  • Assumes only the remote call repeats, never the writes before it.
  • Believes the pipeline caches upstream values and replays them on retry.
  • Says a duplicate record is impossible because the attempt failed.
  • Attaches a retry without listing which stages sit above it.