skip to content

Failure Visibility

An asynchronous failure can vanish: nothing subscribes to handle it and the trace names only framework frames. Interviewers ask how you would find its origin in production.

on this pageshow

questions

5

An observe-only hook in a stream logs every failure signal that passes it — what does the hook change about that signal?

level: middleimportance: must knowfreq 58%

answer

  1. a tap, not a valve
  2. signal continues unchanged
  3. position decides what it sees
  4. upstream failures only
  5. fires once per subscription

basics

~20 s

An observe-only hook changes nothing about the signal. It is a tap: it sees the failure, records it, and lets the same failure continue downstream, so the sequence still ends and every later stage still sees it.

solid answer

~40 s

An observe-only hook is a stage that watches a signal go past and then forwards it untouched. That is exactly what you want for logging and metrics: the record is written, but the failure still terminates the sequence and still reaches whatever is downstream, so the hook cannot accidentally become recovery. Two properties decide what it is worth in production. First, position: a hook sees only failures raised upstream of where you attached it, so one hook near the source records almost nothing. Second, multiplicity: on a source that runs per subscriber, the hook fires once per subscription, not once per pipeline. Keep the body cheap and non-throwing — a hook that throws replaces or hides the failure you were trying to record.

code

pseudocode · 8 lines
pseudocode
source
  .transform(fetch)        // may fail here
  .on_failure(record)      // records fetch failures
  .transform(parse)        // may fail here too
  .subscribe(
     on_value   = write_to_index,
     on_failure = log       // the only place parse failures appear
  )

go deeper

for a junior

Remember the one-line contract: this kind of hook watches a signal and passes it on unchanged. Adding one is safe; it never makes a failing stream succeed.

for a middle

Explain the mechanics: the hook sees only signals that pass its position, it fires once per subscription rather than once per chain, and a throwing body can cost you the failure you were recording.

for a senior

Show the operating judgment: instrument named boundaries so a record says where the run died, put the work item's identity in the record, and keep hook bodies cheap because they run on the worker that delivered the signal.

for a principal

Frame it as a cost question. Instrumentation density trades observability against per-signal overhead and dashboard noise; decide which boundaries are worth a record and make that the default shape teams start from.

## What observe-only means A reactive pipeline is a chain of stages. Down the chain travel **values**, plus one **terminal signal**: either a normal completion or a **failure**. Back up the chain travels **demand** and, when the consumer stops wanting values, a **cancellation**. An *observe-only hook* is a stage you attach to watch one of those signals go past. It is handed the signal, runs whatever body you gave it, and then the **original signal continues unchanged** to the next stage. It consumes nothing, produces nothing and replaces nothing. That contract is the whole reason it is the right home for logging and metrics: - It cannot turn a failure into a value, so adding instrumentation cannot silently change behaviour. - It cannot stop the sequence terminating, so downstream stages and the final subscriber still see the failure. - It can be added and removed without re-reasoning about the pipeline's outcome. A stage that *replaces* a failure with a value or another source is a different kind of stage entirely, and confusing the two is the most common way an engineer describes a hook wrongly in an interview. ## Position decides what a hook can see A signal only reaches a hook if it passes through that hook's position in the chain. A failure raised in stage three never travels backwards, so a hook attached above stage three records nothing about it. | Hook position | Failures it records | Failures it misses | |---|---|---| | Immediately after the source | Failures the source itself raises | Everything raised by later stages | | Mid-chain | Failures raised by any stage above it | Failures raised below it | | Just before the final subscriber | Every failure still travelling at that point | Any failure an earlier stage already replaced | The practical rule for a long pipeline: one hook near the end tells you *that* the run failed; hooks at a few named boundaries tell you *where*. That is why teams instrument boundaries — after fetching, after parsing, after writing — rather than sprinkling hooks everywhere. ## The job that stops quietly Take a background job that indexes documents: fetch, parse, write to the index, repeat. It stops indexing and nobody notices for a day. The first question on-call asks is whether anything recorded the ending at all. If the only hook sits above the parsing stage, a parsing failure produced no record, and the absence of a log line was read as 'nothing happened' rather than 'the run died past my instrumentation'. The hook was honest; it was in the wrong place. The repair has two halves, and both are about visibility rather than behaviour: 1. Move or add hooks so that every failure still travelling at the end of the chain is recorded once. 2. Record the work item's identity in the hook, not just a count — 'indexing failed' tells you nothing you can act on, 'indexing failed for item 4471 at the parse boundary' does. ## What belongs inside a hook - **Cheap, non-throwing work only.** Increment a counter, append a structured log record, tag a span. - **No risky work.** Implementations differ in how they treat a hook that throws: some let the hook's own failure replace or wrap the one you were recording, some route it away as an unhandled signal. Either way you can lose the original. - **No heavy or blocking call.** The hook runs on whichever worker delivered the signal, and that worker has other work queued behind it. - **No mutation of state the rest of the pipeline reads.** An observe-only stage that quietly changes shared state is no longer observe-only in any useful sense. ## Once per subscription, not once per pipeline On a source that starts its work for each subscriber, the chain is *assembled* once but *run* once per subscription. A hook in that chain therefore fires once per run. If two subscriptions to the same indexing chain both fail, the hook records two failures — which is correct, because two pieces of work failed, even though you wrote the hook once. This is a standing trap for dashboards. A panel labelled 'indexing failures' that is really counting hook invocations will read double when a retrying caller subscribes twice to the same chain, and will read zero when nothing subscribed at all. Counting *subscriptions started* alongside *terminal outcomes* is what makes the panel interpretable. ## What to say in an interview Define the hook by its contract — observes, does not alter — then immediately name the two properties that decide what it is worth: it sees only what passes its position, and it fires per subscription. Finish with the discipline: keep the body cheap, record the identity of the work, and never let instrumentation be the thing that changes the outcome.

  • Two observe-only hooks sit at different points in one chain and report different failure counts. What explains the gap?
    Each hook sees only the signals that reach its position. A stage between them can replace a failure with a value, so the lower hook never sees it, and a stage between them can raise a new failure that only the lower hook sees. Different positions, different populations — not a broken counter.
  • What happens if the body of an observe-only failure hook itself throws?
    You risk losing the failure you were recording. Implementations differ: some let the hook's own failure travel in place of, or wrapped around, the original; others route it away as an unhandled signal. Treat a hook body as code that must not throw, and guard anything inside it that could.

saying these in an interview costs you the question

  • Says an observe-only hook handles the failure, so the sequence continues
  • Believes a hook sees failures raised by stages placed after it
  • Cannot distinguish a hook from a stage that replaces the failure
  • Puts an expensive or blocking call inside the hook body
  • Assumes the hook fires once per pipeline however many subscriptions run
open as a page

A background indexing job stopped a day ago and no failure was logged anywhere — how can an asynchronous failure disappear entirely?

level: seniorimportance: must knowfreq 54%

basics

~20 s

A failure signal must be delivered to something. If the terminal subscription registered only a value handler, it has no destination, so it is raised on an anonymous worker or routed to a process-wide sink nobody collects.

open as a page

An indexing failure's stack trace names only worker-pool frames and none of your code — what happened to the caller?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The trace was captured on the pool worker that raised the failure, and that worker's stack begins at its task loop. The code that built the chain and subscribed ran elsewhere and returned long ago.

open as a page

Your teams keep losing asynchronous failures in production streams — what visibility standard would you set, and how would you make it stick?

level: principalimportance: should knowfreq 38%

basics

~20 s

Set a small outcome contract every stream job must meet, then make the compliant path the easiest one to write rather than a rule to remember. Standardise what must be observable, never which library produces it.

open as a page

Why should a stream run that ends by cancellation be counted as an outcome distinct from one that ends by failure?

level: middleimportance: nice to knowfreq 31%

basics

~10 s

Cancellation means the consumer stopped wanting values, not that anything broke. Folding it into the failure counter inflates the error rate during ordinary disconnects; folding it into success hides work that stopped half done.

open as a page