skip to content

How do failures travel through a chain of asynchronous steps built from futures, and why does wrapping the call in a try/catch at the call site usually fail to catch them?

level: middleimportance: must knowfreq 52%

answer

  1. Two tracks: value and error
  2. Error skips transforms until a handler
  3. try/catch only sees construction-time failures
  4. Unobserved failure = silent
  5. Recover vs enrich vs cleanup-on-both

basics

~20 s

A future completes with either a value or an error, and an error short-circuits the downstream transforms until a handler that accepts errors is reached. A try/catch at the call site only sees synchronous failures, because the function returns a pending handle long before the failure exists.

solid answer

~1 min

A future has two terminal channels. Once it fails, every value-transforming step downstream is skipped and the error is passed along unchanged until something that handles errors — recover, catch, handle — is attached. That is why intermediate steps need no error plumbing. A try/catch around the call site catches only what happens **while building** the chain: argument validation, a null dereference before the async call starts. The interesting failure happens later, usually on another thread or a later event-loop turn, when the calling frame is long gone. There is nothing on the stack to unwind into. The operational consequences are the ones interviewers want: - **Unobserved failures.** A future that fails with no error handler attached vanishes silently. You need a terminal handler on every chain, plus a runtime-level unhandled-failure hook to catch what you missed. - **Fire-and-forget** — starting a future and ignoring the handle — is the same bug. - **Lost causality.** The stack at failure time belongs to the completion machinery, so add context at each hop rather than relying on the trace. - Distinguish **recovering** (substitute a fallback value) from **enriching** (wrap with context and rethrow) from **finally-style** cleanup that must run on both outcomes.

code

text · 10 lines
text
try {
    f = callService()     # returns immediately, still pending
} catch (e) {
    # only reached if building the call threw synchronously
}
# ... calling frame returns here ...
# ... 40 ms later, on another thread, the service call fails ...
# nothing to unwind into: the error lives in f, not on a stack

f.onError(e -> handle(e))   # the only place it can be seen

go deeper

for a junior

Say that a future ends as a value or an error, that errors skip the value steps until a handler, and that try/catch at the call site is too early to see them.

for a middle

Explain why the frame is gone by completion time, and name the unobserved-failure problem plus the terminal-handler rule.

for a senior

Cover operational defences: global unhandled hooks with alerting, context enrichment across hops, cleanup on both tracks, and losers of fan-in combinators failing silently.

for a principal

Treat failure semantics as a contract question — what a partial failure means to callers, whether fallbacks are permissible, and how causality is preserved across an async boundary in your observability design.

## Two channels, one pipeline Every future terminates in exactly one of two ways: a value or an error. Think of a chain of steps as a two-track pipeline. Value-transforming steps run on the value track only. The moment a step fails, the pipeline switches to the error track: subsequent transforms are **skipped**, not executed with a null, and the error is carried forward untouched until it meets a step declared to handle errors. This is why a well-built chain has no error handling in the middle at all. Ten steps and one handler at the end is normal and correct — the same shape as a try block with a single catch, but assembled at runtime out of callbacks. ## Why try/catch at the call site does nothing An async function returns almost immediately, handing back a pending handle. The failure occurs later, and typically on a different thread or a later turn of the event loop. A try/catch is a property of the *stack frame*, and by the time the failure exists that frame has already returned. There is nothing to unwind. Exactly two kinds of failure are still synchronous and therefore catchable at the call site: mistakes made while constructing the chain (validating arguments, a bad configuration lookup) and exceptions thrown by a mapping function attached to an already-completed future, which runs inline on the attaching thread. Well-designed libraries deliberately convert even the first kind into a failed future so that callers have **one** error path to handle rather than two. Sequential await-style syntax restores ordinary try/catch, but only because the compiler rewrites the code into continuations and re-raises the failure at the suspension point. The underlying mechanism is unchanged. ## The unobserved failure The most damaging property of the model: a future that fails and has no error observer is, by default, **silent**. Nobody unwinds, nothing is logged, and the request that was waiting on it just never finishes — or, worse, some other path completed and the failure indicates a half-done side effect that nobody knows about. Defences, in order of importance: 1. **Every chain ends in a terminal step** that handles the error case — completes a response, logs, increments a metric. 2. **Never fire and forget.** Starting a future and discarding the handle discards the error channel too. If the work is genuinely background, attach a handler that logs and reports. 3. **Install a global unhandled-failure hook** if the runtime offers one, and alert on it. It is a bug detector, not a strategy. 4. **In fan-in combinators, log the losers.** A fail-fast *all* propagates the first error only; the other failures are unobserved by construction unless you handle them. ## Loss of causality, and what to do about it The physical stack at failure time shows completion machinery, not the code that requested the operation. Practically: - Attach context at each hop — wrap the error with which call, which entity id, which attempt — so the message carries the story the stack cannot. - Rely on a correlation identifier propagated explicitly through the pipeline, since thread-scoped context is not reliable when continuations hop threads. - Use the library's async causality support if it has one, and know that it costs something to capture. ## Three different responses to failure Candidates blur these; interviewers separate them: - **Recover:** map the error to a fallback value, putting the pipeline back on the value track. Only correct when the fallback is genuinely acceptable — a cached value, an empty list. Recovering to a default because you do not want to think about the failure is how corrupted results reach users. - **Enrich and rethrow:** stay on the error track but add context. Cheap and almost always worth it. - **Cleanup on both tracks:** the finally equivalent. Releasing a connection or a lock must happen for value *and* error outcomes; a cleanup step attached only to the success path leaks under failure — a leak that only appears when something else is already going wrong. A fourth, related concern: a value that arrives after nobody wants it (because a race was lost or the caller cancelled) may itself hold a resource. Dispose it in the terminal handler. ## Timeouts are part of this A timeout is normally implemented as a competing failure: race the work against a timer and fail the composite if the timer wins. Under single assignment, the loser cannot corrupt the outcome — but the losing work keeps running unless you cancel it, so the timeout branch should trigger cancellation too. ## In an interview Say: two terminal channels, errors skip value steps until a handler; try/catch at the call site catches only construction-time failures because the frame is gone by completion time; and the real production risk is the unobserved failure, defended with a terminal handler on every chain plus a global hook.

  • How does an await-style sequential syntax let you use ordinary try/catch again if the underlying mechanism is continuations?
    The compiler splits the function at each suspension point into continuations and records the enclosing handlers. When the awaited future completes with an error, the runtime re-raises it at the resumption point, inside the same lexical try block. So the try/catch works because the language reconstructs the frame, not because the failure unwound a real stack.
  • What is an unobserved or unhandled failure, and how would you detect it in production?
    It is a future that completed with an error while nothing was attached to read the error channel, so no code ever learned about it. Detect it with the runtime's unhandled-failure hook wired to a logger and a metric, alert on that metric being non-zero, and treat any hit as a defect in a chain that lacks a terminal handler. Also review fire-and-forget call sites and fail-fast fan-in combinators, which are the usual sources.

saying these in an interview costs you the question

  • Expecting a try/catch around the call site to catch failures that happen after the function returned
  • Adding error handling to every intermediate step, not knowing that errors skip value transforms automatically
  • Recovering to a default value simply to make the failure go away, hiding real errors from callers
  • Starting futures and discarding the handle, and with it the entire error channel
  • Attaching cleanup only to the success path, so resources leak precisely when things fail

context