skip to content

A background indexing job stopped a day ago and no failure was logged anywhere — how can an asynchronous failure disappear entirely?

level: seniorimportance: must knowfreq 54%

answer

  1. silence is not evidence
  2. who is listening for failure?
  3. no handler, nowhere to deliver
  4. process-wide sink of last resort
  5. the default sink alerts nobody

basics

~20 s

A failure signal must be delivered to something. If the terminal subscription registered only a value handler, it has no destination, so it is raised on an anonymous worker or routed to a process-wide sink nobody collects.

solid answer

~50 s

In a synchronous call a failure propagates up the caller's stack until something catches it, and if nothing does, the thread dies loudly. In an asynchronous pipeline a failure is a *signal* delivered to a subscriber, so it goes wherever that subscriber's failure handler points. Subscribe with only a value handler and there is no such place: implementations differ, but the failure is either raised on the worker that happened to deliver it — where an uncaught-handler nobody configured prints it at best — or routed to a process-wide sink for signals with nowhere to go. Neither path stops the process, so a dead background job leaves a perfectly healthy service behind. The same sink catches failures that arrive after the subscription already ended or was cancelled. Fixes: never subscribe without a failure handler, and configure that global sink at startup to count and log with a stack.

code

pseudocode · 11 lines
pseudocode
// at startup: the last resort, wired to something
set_unhandled_signal_sink(signal -> {
    counter("async.unhandled").increment()
    log_with_stack(signal)
})

// the job: a failure handler is not optional
indexing_chain.subscribe(
    on_value   = write_to_index,
    on_failure = f -> { counter("indexing.failed").increment(); log(f) }
)

go deeper

for a junior

Take away one rule: when you subscribe to a stream, always say what should happen on failure. A subscription with only a value handler can lose the failure completely.

for a middle

Explain the delivery model: a failure is a signal delivered to a registered destination, and with none registered it falls back to a worker's uncaught path or a process-wide sink that is not wired to anything by default.

for a senior

Demonstrate the on-call reasoning: the process being healthy proves nothing, so check the subscription site and the sink's configuration, count sink events as a defect class, and record a terminal outcome per run so an ending is visible.

for a principal

Own the contract across teams: decide whether an unhandled event pages, is counted, or fails the build in test, and make the compliant subscription the easiest one to write rather than a rule people must remember.

## Why silence is possible at all In ordinary synchronous code a failure has an obvious route: it walks up the call stack of the thread that raised it, and if no frame handles it the thread terminates in a way somebody notices. The route is built into the machine. In an asynchronous pipeline there is no such route. A failure is a **terminal signal** handed down the chain to the subscriber, and it goes exactly where that subscriber says to put it — no further. So the question 'where did my failure go?' is really the question **'what did the final subscription ask to be done with failures?'** There are three ways the answer turns out to be 'nowhere': 1. **The subscription registered no failure handler.** Only values were subscribed for. The failure arrives with no destination. 2. **Nothing subscribed at all.** A chain that is built but never subscribed never runs, so there is no failure — and also no work. The symptom is identical from the outside: silence. 3. **The signal arrived after the subscription was already over.** The sequence had terminated, or the consumer had cancelled, and work still in flight failed afterwards. There is no longer a subscriber to deliver to. ## Where an undeliverable failure actually goes Implementations differ here, and a good answer says so rather than asserting one behaviour: - Some **raise it on the worker that was delivering the signal**. That worker belongs to a shared pool; its uncaught-signal handler is whatever the runtime defaults to, which typically prints somewhere nobody is collecting and then returns the worker to the pool. - Some **route it to a process-wide sink** reserved for signals that have nowhere to go — unhandled failures and signals that arrive post-termination. The default behaviour of that sink is usually a log line at a level no alert watches. What almost never happens is the thing engineers expect: the process does **not** die. The worker goes back to serving other work, health looks fine, memory looks fine, and one background job is simply gone. | Symptom on-call sees | What actually happened | Where to look first | |---|---|---| | No log line, no metric movement | No failure handler on the subscription | The subscription site, then the global sink's configuration | | No log line, no work ever started | Nothing subscribed to the chain | A counter of subscriptions started | | A stack in an odd place, no request context | Failure raised on a delivering worker | The pool's uncaught-signal path | | Sink events rising alongside cancellations | Failures landing after the subscriber left | Cancellation and timeout paths | ## The indexing job, diagnosed The job fetches documents, parses them, writes to an index, and was launched once at startup as fire-and-forget: subscribe with a value handler and move on. A malformed document made the parsing stage fail. The failure signal reached the subscription, found no failure handler, and left by one of the two routes above. The sequence terminated — correctly — and with it the subscription that was the entire job. For twenty-four hours the service was healthy and indexed nothing. Notice what did **not** help: the process stayed up, so any check that only asks whether the process is alive stayed green. ## Making it impossible to lose one 1. **Always register a failure handler at the terminal subscription.** This is the single highest-value rule in the whole subject, and it is mechanical enough to enforce by review or by making one wrapper the only sanctioned way to subscribe. 2. **Configure the global sink at startup.** Count every event, log it with whatever stack is available, and tag it as a defect class of its own. An event there means a failure that your code never took responsibility for. 3. **Alert on that counter, not just on the log.** A default sink alerts nobody; it is the last resort, not the plan. 4. **Emit a terminal outcome record per run** — completed, failed or cancelled — so a run that ends is visible as an ending rather than as an absence. 5. **Count subscriptions started as well as endings**, so 'nothing ever subscribed' is distinguishable from 'it ran and failed'. ## What an interviewer is listening for The weak answer is 'the exception was swallowed'. The strong answer names the mechanism: a failure is a delivered signal, delivery needs a registered destination, and when there is none the implementation falls back to a worker's uncaught path or a process-wide sink — neither of which is wired to anything by default, and neither of which stops the service. Then it names the two fixes an on-call engineer can actually check tonight: the subscription site and the sink's configuration.

  • Why does an unhandled asynchronous failure usually leave the process perfectly healthy?
    Because the only thing that ended is one subscription. The failure terminates that sequence and releases its resources; the worker that delivered it returns to its pool and keeps serving everything else. Nothing about the process is wrong — one unit of work is simply gone, which is why liveness of the process proves nothing here.
  • How can a failure arrive when there is no longer anywhere to deliver it?
    The subscription already ended — it completed, failed, or the consumer cancelled — while work was still in flight upstream. When that work fails afterwards, the subscriber is gone. The signal is undeliverable and takes the same last-resort route as an unhandled one, which is why sink events climb during cancellation storms.
  • Why is a default sink that prints at a low level dangerous rather than merely unhelpful?
    It converts a defect into the appearance of a working system. The failure looks handled because something wrote a line, but nothing counts it, nothing alerts on it, and in a busy service the line is rotated away. Counting the event is what makes it real; the log line is only the detail.

saying these in an interview costs you the question

  • Assumes an unhandled asynchronous failure crashes the process
  • Thinks the default global sink alerts somebody
  • Believes a catch block around the assembly code would catch it
  • Reads absence of log lines as evidence that nothing failed
  • Subscribes with a value handler only and calls the job done
  • Cannot tell a job that failed from a job that never started