skip to content

Error Handling

Errors travel down a stream and end it unless you supply a fallback value, a fallback source or a retry. Interviewers probe where you place recovery, since the wrong place kills the subscription.

on this pageshow

explore

questions

17

When a notification pipeline that loads a recipient, records an attempt and calls a carrier is retried after the carrier fails, which steps run again?

level: juniorimportance: must knowfreq 55%

answer

  1. retry is not resume
  2. the subscription is remade
  3. everything above the retry point re-runs
  4. effects repeat, not just the call
  5. two attempt records for one send

basics

~20 s

Retrying resubscribes to the source, so every stage above the retry point runs again from the start: the recipient is loaded a second time, a second attempt record is written, and the carrier is called again.

solid answer

~40 s

A retry does not resume the sequence at the stage that failed. It drops the failed subscription and subscribes again to everything above it, so the whole retried scope re-runs from its source. In this dispatcher that means the recipient load, the message composition, the attempt record and the carrier call all happen a second time. Two of those are harmless — a read and a pure computation — but the attempt record is an effect on the outside world, so one logical send now leaves two records behind. The pipeline has no idea any stage was meant to happen once. Before attaching a retry, the question to answer is not `how many attempts` but `what exactly sits above this point`.

code

pseudocode · 12 lines
pseudocode
// assembled once; nothing runs until a subscriber attaches
pipeline = recipient_for(message_id)
    .map(compose_message)
    .effect(record_attempt)        // writes one attempt record
    .flat_map(call_carrier)        // fails on attempt 1
    .retry(max_attempts = 3)

// attempt 2 subscribes to recipient_for(message_id) again:
//   recipient loaded again
//   message composed again
//   a SECOND attempt record written
//   carrier called again

go deeper

for a junior

Remember the one-line fact: retrying a stream resubscribes to its source, so the work starts over rather than resuming. Be able to point at which stages of a small pipeline would therefore run a second time.

for a middle

Explain why no finer resumption is possible: a pipeline is a description with no checkpointed per-stage state, so the only recovery available is to run the description again. Classify each upstream stage as pure, read or effect.

for a senior

Show that you look at the retried scope before you add a retry, and name the duplicate the scope would produce in a real dispatcher. Mention that already-delivered values are not withdrawn when a multi-value sequence restarts.

for a principal

Frame it as a standard: where retries are allowed to sit in a pipeline, what a stage must guarantee to be inside a retried scope, and who owns the duplicate that escapes when that guarantee is only assumed.

## Retry is a resubscription, not a resumption A stream pipeline is a **description** of work, not the work itself. Nothing runs until something subscribes, and the subscription is what turns that description into a running sequence of stages. A retry stage lives under the same rule. When a failure signal reaches it, it cannot reach back into the stage that failed and try that stage alone — it holds no handle on that stage's half-finished state. What it holds is the description above it. So it releases the failed subscription and **subscribes again**, and everything above the retry point, inside the stream it is attached to, runs from the beginning. That one sentence accounts for nearly every surprise this mechanism produces: | What people expect | What resubscription actually does | |---|---| | execution resumes just after the failing stage | a fresh subscription starts at the top of the retried scope | | values already computed are reused | every stage above the retry point computes again | | only the remote call is repeated | every effect above the retry point happens again | | downstream sees one more value | downstream may see values it has already seen | ## Walking the dispatcher Take an outbound notification dispatcher with four stages: load the recipient, compose the message, record an attempt, hand the message to a carrier. The retry sits at the end, and the carrier fails. - **Load the recipient** — a read. It runs again. That is a second round trip, and what it returns is not guaranteed identical: something may have changed between attempts, so the second run is not necessarily a replay of the first. - **Compose the message** — a pure computation. It runs again and costs only time. - **Record the attempt** — an effect that writes. It runs again, so one logical send has now produced two attempt records. - **Call the carrier** — the stage the retry was actually for. It runs again, which is the point. Three of the four stages re-ran because they sat inside the retried scope; only one of them was meant to. ## Pure stages, reads and effects The useful move before attaching a retry is to classify every stage above the retry point: 1. **Pure stages** — computation over the value in hand. Repeating them costs latency and processor time, nothing more. 2. **Reads** — they change nothing outside, so repeating them is safe in that sense, but they add load and may answer differently on each attempt. 3. **Effects** — writes, sends, publishes, counters. Repeating one of these repeats it in the world, and this is where duplicates come from. The duplicate attempt record is not a defect in the retry mechanism. It is the mechanism doing exactly what it promises, over a scope nobody examined. ## Multi-value sequences repeat too A dispatcher with one message in flight hides a second consequence. If the retried scope produces several values and fails partway through, the values already delivered downstream stay delivered — resubscribing cannot un-emit them. The new subscription then starts the sequence from the top, so a consumer that had already seen the first three values sees them a second time. Any downstream stage that itself has an effect must therefore tolerate duplicates, or the retry has to sit somewhere that cannot produce them. ## Why the contract is built this way One could imagine a mechanism that checkpoints each stage and resumes after the failed one. Stream pipelines deliberately do not offer that: a stage keeps no addressable, resumable state, and the only thing guaranteed reproducible is the description itself. Resubscription is the one recovery a pipeline can always perform, precisely because it needs nothing except that description. The generality is paid for by the author, who has to know what running the description again means in the world. ## The practical rule Name every stage above the retry point. Mark each one pure, read or effect. For each effect, decide whether it is safe to repeat, safe once it is keyed so repeats converge on a single record, or whether it has to move out of the retried scope entirely. Then bound the attempts, so a failure that will never clear does not multiply that work without limit. A retry attached without that walk is not a resilience measure; it is a duplicate generator with a limit on it.

  • If every stage above the retry point is a pure computation, is there still an argument against retrying that whole scope?
    Yes, two of them. Repeating the work costs latency and processor time on every attempt, which matters when the scope is large. And a stage that only looks pure but actually reads something outside can answer differently on the second run, so the retried attempt is not the same attempt with a second chance.
  • What happens to values a multi-value sequence already delivered downstream when the retry restarts it?
    They stay delivered. Resubscribing starts a fresh sequence; it does not withdraw what the previous subscription already emitted. A consumer that saw the first three values sees them again after the restart, so any downstream effect has to tolerate duplicates.
  • Does the failed stage get any chance to clean up before the resubscription?
    Only whatever cleanup it registered for cancellation or failure. The retry releases the failed subscription, which is what triggers that cleanup; it does not roll back effects the stage already completed. A record already written stays written.

It is closer to restarting a recipe from the shopping trip than to picking up the pan where it burned: the ingredients are bought again whether or not you still had them.

saying these in an interview costs you the question

  • Thinks retry resumes at the failed stage, leaving earlier stages untouched.
  • Assumes only the remote call repeats, never the writes before it.
  • Believes the pipeline caches upstream values and replays them on retry.
  • Says a duplicate record is impossible because the attempt failed.
  • Attaches a retry without listing which stages sit above it.
open as a page

An observe-only hook in a stream logs every failure signal that passes it — what does the hook change about that signal?

level: middleimportance: must knowfreq 58%

basics

~20 s

An observe-only hook changes nothing about the signal. It is a tap: it sees the failure, records it, and lets the same failure continue downstream, so the sequence still ends and every later stage still sees it.

open as a page

In a product-page pipeline, what happens to the formatting stages between a failing price lookup and a recovery step placed last?

level: middleimportance: must knowfreq 62%

basics

~20 s

Nothing runs in them. A failure signal travels past every stage that only handles values, so the formatting stages are skipped and the substituted value enters the sequence below them — it must therefore already be in the shape the subscriber expects.

open as a page

Your dispatcher resubscribes on every carrier failure without limit; what must a retry decision take into account before it resubscribes again?

level: middleimportance: must knowfreq 58%

basics

~20 s

A retry decision needs three inputs: the kind of failure, since some can never clear; a bound on attempts, as a count or a deadline; and some space between attempts. When the bound is reached, the failure must reach the subscriber.

open as a page

A stream-based import validates ten thousand address rows and row twelve signals a failure — what happens to rows thirteen onward?

level: middleimportance: must knowfreq 72%

basics

~10 s

Nothing processes them. A failure signal is terminal: it ends the sequence at row twelve and releases the subscription, so the source is never asked for rows thirteen onward and nothing downstream sees them.

open as a page

A per-row validator rejects an address for a malformed postcode — why carry that rejection as a value rather than a failure signal?

level: middleimportance: must knowfreq 58%

basics

~20 s

A rejected row is an expected outcome, not a broken stream. Signalled as a failure it ends the run at that row; carried as a value in the sequence it flows on, so the remaining rows are still processed and both counts are reported.

open as a page

A background indexing job stopped a day ago and no failure was logged anywhere — how can an asynchronous failure disappear entirely?

level: seniorimportance: must knowfreq 54%

basics

~20 s

A failure signal must be delivered to something. If the terminal subscription registered only a value handler, it has no destination, so it is raised on an anonymous worker or routed to a process-wide sink nobody collects.

open as a page

In a product-page stream that flattens in a failing recommendation source, what does moving recovery outside the flattening cost you?

level: seniorimportance: must knowfreq 54%

basics

~20 s

Everything except the fallback. The inner failure escapes into the outer sequence and terminates it, so one recovery step outside the flattening replaces the whole remaining page — the price and stock work already done is abandoned and sources still in flight are cancelled.

open as a page

When a stock-check stream fails and a recovery step substitutes an empty result, what does the subscriber then receive?

level: juniorimportance: should knowfreq 55%

basics

~20 s

The subscriber receives the substituted empty result as an ordinary value, followed by a normal completion. Recovery does not re-run the stock check; it replaces the unfinished remainder of the sequence, so values delivered before the failure still stand.

open as a page

An indexing failure's stack trace names only worker-pool frames and none of your code — what happened to the caller?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The trace was captured on the pool worker that raised the failure, and that worker's stack begins at its task loop. The code that built the chain and subscribed ran elsewhere and returned long ago.

open as a page

Retrying your dispatcher pipeline duplicates the attempt rows written before the carrier call, so how do you restructure it so only the carrier call repeats?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Make the carrier call its own per-item source and attach the retry to that inner source, so resubscription re-enters only the call. The recording stays in the outer pipeline, outside the retried scope, and runs once.

open as a page

A nightly address import writes validated rows in one batch at the end, and a malformed row at position 9,998 signalled a failure — why were the earlier good rows lost?

level: seniorimportance: should knowfreq 46%

basics

~10 s

The write was keyed to the sequence ending normally. A failure is a different ending, so the stage that had accumulated 9,997 validated rows was torn down with the subscription and never wrote anything.

open as a page

Your teams keep losing asynchronous failures in production streams — what visibility standard would you set, and how would you make it stick?

level: principalimportance: should knowfreq 38%

basics

~20 s

Set a small outcome contract every stream job must meet, then make the compliant path the easiest one to write rather than a rule to remember. Standardise what must be observable, never which library produces it.

open as a page

Why should a stream run that ends by cancellation be counted as an outcome distinct from one that ends by failure?

level: middleimportance: nice to knowfreq 31%

basics

~10 s

Cancellation means the consumer stopped wanting values, not that anything broke. Folding it into the failure counter inflates the error rate during ordinary disconnects; folding it into success hides work that stopped half done.

open as a page

What does a recovery step that substitutes an empty recommendation list for every failure hide from you?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Every failure that is not the recommendation dependency being unavailable: defects raised by stages above the recovery step, malformed responses, configuration mistakes. All of them render as a healthy-looking empty block, so the page never breaks and nobody learns it is broken.

open as a page

A pipeline consuming a live shared feed of send-requests fails and retries by resubscribing; why does the request that failed not come back?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Resubscribing attaches a new subscription to a source that never stopped running, so it delivers only what arrives from that moment on. Whatever the feed emitted before the new subscription attached, including the failed request, is gone.

open as a page

Each bulk import pipeline on your platform decides for itself whether a rejected record ends the run — what standard would you set?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Define terminality by whether the remainder of the run still has value. Environment-level faults end the run; per-record rejections travel as data with a required destination, a count on every run report, and a rate threshold that stops a run when the input contract has clearly changed.

open as a page