skip to content

A background task is started fire-and-forget — nobody keeps its handle or waits for it — and it throws. What typically happens to that error, and why is it dangerous?

level: juniorimportance: must knowfreq 50%

answer

  1. no owner = nowhere for the error to go
  2. unhandled rejection vanishes at collection
  3. green dashboards, work not done
  4. dead poller looks alive
  5. every task needs an owner

basics

~20 s

With no owner waiting, the error has nowhere to propagate: it ends that task and lands in a default handler, an unread log line, or a result object nobody inspects. The system keeps reporting healthy while the work silently is not being done.

solid answer

~50 s

An exception propagates to whoever is waiting for the task. In fire-and-forget there is no such party, so depending on the runtime the error goes to a process-wide default handler, gets written once to a log nobody watches, or is stored in a future/promise that is never inspected and disappears when it is collected — the classic "unhandled rejection". The danger is not the lost stack trace, it is the **silence**. The caller already returned success, so metrics, health checks and the user all say fine while the email was never sent, the cache was never invalidated, the audit row was never written. You discover it days later from inconsistent data, with no trace connecting it to the request that caused it. The fix is ownership: every task should have someone who observes its outcome — a scope or supervisor that joins it. Where fire-and-forget is genuinely intended, it must be explicit: an attached failure handler, a counted metric, and an alert on that counter.

code

text · 12 lines
text
# fire-and-forget: nothing observes the outcome
spawn(send_confirmation_email)      # throws -> default handler / silence
return OK                           # caller reports success either way

# owned: the scope joins, so the failure surfaces at the caller
with scope() as s:
    s.spawn(send_confirmation_email)
# leaving the scope joins the child; its failure is raised here

# genuinely detached, but explicit
spawn(refresh_cache)
  .on_failure(e -> { log(e, correlation_id); metric("bg.failure", task="refresh") })

go deeper

for a junior

Say the error has no owner to propagate to, so it is logged at best or lost, and the system looks healthy while the work silently did not happen.

for a middle

Name the concrete outcomes (default handler, unhandled rejection collected later, task ends quietly), the partial-state consequence, and that the fix is a scope that joins the task.

for a senior

Talk about dead background loops looking healthy, alerting on effect rather than liveness (freshness, last-success time, queue depth), correlation ids into background work, and moving must-happen work to a durable queue with retries and a dead-letter path.

for a principal

Make ownership a platform rule: no unowned spawn sites, every long-lived task under a supervisor with a restart policy, and freshness or effect-based SLOs so silent background failure is detectable by construction.

## How errors normally travel In sequential code an error propagates up the call stack until something handles it; if nothing does, the program stops and you notice. Concurrency breaks that chain: a spawned task has its own stack, and the only path back to the starting code is the **handle** — the future, promise, join operation, or scope that represents the task's eventual outcome. Failure propagation is just "the owner observes a failed outcome". Fire-and-forget deletes the path. No handle is kept, nothing joins, so there is no destination for the error. ## What actually happens Runtimes differ, but the outcomes are all variations of *nobody acts*: - **Default handler.** A process-wide hook logs the error with the worker thread's stack. It is disconnected from the request that caused it, often lacks correlation ids, and is buried among unrelated output. - **Silently stored.** The failure is recorded inside a result object that nobody ever inspects; when that object is collected the error is gone. Some runtimes emit a late "unhandled" warning at collection time — minutes later, non-deterministically, sometimes not at all. - **Task ends quietly.** In the worst case the task simply terminates, indistinguishable from success. In none of these does the caller learn, retry, compensate, or fail the request. ## Why silence is the real danger 1. **False health.** The request returned 200 and the dashboards are green, so nobody investigates. Failure is invisible precisely when it is happening. 2. **Partial state.** The visible part committed, the background part did not: an order exists with no confirmation email, a document is indexed with no thumbnail, a permission is revoked in one store and not the other. Divergence accumulates. 3. **Lost causality.** By the time the inconsistency is spotted, the log line — if it exists — is old and carries the wrong stack. Reconstructing which request failed is expensive or impossible. 4. **Nothing retries.** Retry and compensation live at the owner. With no owner, the failure is terminal by default even when it was a transient blip. 5. **Silent absence of a whole service.** A background loop — a poller, a consumer, a scheduled refresh — that dies on its first exception simply stops existing. Everything looks alive; the data just stops updating. This is the single most common production version of the bug. ## The fix: give every task an owner Structured concurrency's core rule is that a task cannot outlive the scope that created it, and the scope **joins** every child. That makes the failure path structural: the child's exception surfaces at the scope's exit, in the caller's own stack, where existing error handling already applies. You cannot accidentally drop it, because you cannot accidentally leave the scope. For work that genuinely must outlive a request — a background refresher, a queue consumer, a periodic compaction — the task still gets an owner, just a longer-lived one: a supervisor that watches its outcome and decides what to do (log and stop, restart with backoff, escalate). The choice of policy is supervision; the requirement is that *somebody observes*. ## When fire-and-forget is acceptable Rarely, and never implicitly. Make it explicit and instrumented: - attach a failure handler at the spawn site — no task is launched without one; - increment a labelled failure metric and alert on it, because a log line alone is not an owner; - carry the correlation id of the originating request into the task so the failure can be tied back; - make the work idempotent and, when it matters for correctness, durable — put it on a queue with retries and a dead-letter destination instead of in memory, since an in-memory fire-and-forget task also dies with the process on every deploy; - confine it to work whose loss is genuinely acceptable (a best-effort metric ping), and be honest that anything else is not fire-and-forget but an unreliable queue. ## Detecting it in an existing codebase Search for spawn sites whose result is discarded, tasks submitted to shared/global executors, timers and event callbacks registered with no failure handler, and long-lived loops with no restart policy. A quick operational test: force an exception in each background task in a test environment and check whether *anything* — an alert, a metric, a failed request — reacts within a minute. If nothing does, the task is unowned.

  • A long-running background poller dies on its first exception. What symptoms does the team actually see?
    Usually none from the service itself: health checks pass, CPU and memory look normal, and no request fails, because the poller was not on any request path. The symptom is downstream and delayed — data that stops changing, a cache that goes stale, a queue whose depth grows without bound. The lesson is to alert on the work's effect (last-success timestamp, freshness, queue depth) as well as on the process being up.
  • When is fire-and-forget actually acceptable?
    When the work is genuinely best-effort and its loss changes nothing important — a metric ping, an optional cache warm — and even then it should be explicit: a failure handler attached at the spawn site plus a counted, alertable metric. Anything whose loss creates inconsistency belongs on a durable queue with retries and a dead-letter path, because an in-memory background task also dies on every deploy and restart.

It is mailing a letter with no return address to an address that may not exist: if it fails to arrive, nothing comes back to tell you — you find out much later from the consequences.

saying these in an interview costs you the question

  • Assumes the runtime will surface an exception from a task nobody is waiting on
  • Thinks logging the error inside the task counts as handling it
  • Uses fire-and-forget for work that must happen, such as writes or notifications
  • Never checks whether a background loop is still alive after its first failure
  • Believes an in-memory background task is a substitute for a durable queue

context