skip to content

Some concurrent work legitimately outlives any single request — background schedulers, cache refreshers, connection keep-alive loops, metric flushers. How do you reconcile that with a discipline that forbids orphan tasks?

level: principalimportance: should knowfreq 28%

answer

  1. longer-lived owner, not no owner
  2. lifetime ladder: app > component > connection > request
  3. owned copies + own resources + own identity
  4. shutdown: reverse order, grace budget, escalate and log
  5. dead daemon = orphan; supervise and cap

basics

~20 s

Give them a longer-lived owner instead of no owner. Each background task belongs to an explicit component or application-lifetime scope that starts it, watches its failures, and cancels and joins it in a defined shutdown order. Ownership moves up the tree; it never disappears.

solid answer

~50 s

The rule is *every task has an owning scope*, not *every task ends with the request*. Long-lived work gets a long-lived scope. In practice: - **Name the lifetimes.** Application scope, component/module scope, connection scope, request scope. Each owns tasks whose lifetime matches it, and each nests inside the one above. - **Own the data too.** A background task must not borrow request-scoped context or pooled resources; it gets owned copies and its own resources, so the lifetime nesting is real. - **Define shutdown.** Cancel innermost-first or in reverse dependency order, with a grace budget per level and a hard deadline, and log what did not stop. - **Supervise failures.** A daemon that dies silently is an orphan by another name: the owning scope decides restart-with-backoff versus fail the component. - **Bound them.** Long-lived scopes are where unbounded spawning hides, so cap concurrency inside them. The test: at any instant you can name the scope responsible for every live task.

go deeper

for a junior

Know the core idea: background work still belongs to something — an application-level scope that starts it and stops it at shutdown.

for a middle

Add that the task must not borrow request-scoped context or pooled resources, and that shutdown cancels and waits for it.

for a senior

Design the lifetime ladder and the shutdown sequence: reverse dependency order, per-level grace budgets, escalation, supervision with backoff, bounded workers.

for a principal

Make it an enforceable platform invariant: named scopes with owners, a shutdown budget derived from the platform kill deadline, health signals driven by supervision policy, and introspection that lists live scopes so leaks are found before production.

## Restating the rule correctly Structured concurrency is often summarized as "tasks must not outlive the function that started them", which makes daemons look like a contradiction. The precise invariant is weaker and more useful: **every task has an owning scope that will eventually join it, can cancel it, and receives its failures**. Nothing in that requires the owner to be a request. Long-lived work needs a long-lived owner. ## Build a lifetime hierarchy Make the lifetimes explicit and nested, largest first: ```text application scope (process lifetime: schedulers, metric flushers) └ component scope (module/subsystem: cache refresher, pool keep-alive) └ connection scope (one client session: reader loop, heartbeat) └ request scope (one call: fan-out children) ``` Each task is started in the scope matching its intended lifetime. A cache refresher does not belong to the request that happened to trigger a miss; it belongs to the cache component. This single placement decision fixes most "why is this still running?" incidents, because it is now answerable by construction. ## Ownership of data must move with the task Promoting a task to a longer-lived scope is only sound if it stops depending on shorter-lived state. A background task must not capture a request-scoped context, a pooled connection, or a mutable object the request thread will keep changing. Hand it **owned copies** of the values it needs (ids, immutable snapshots) and give the long-lived scope **its own resources** — its own connection or client, its own identity or service credential rather than the triggering user's. If a background task needs the caller's authority, that authority must be materialized into something with a matching lifetime, not borrowed. ## Shutdown is the real design work A long-lived scope's contract is exercised exactly once: at shutdown. Decide and write down: - **Order.** Generally stop accepting new work at the edge first, drain request scopes, then component scopes, then application scope — reverse dependency order, so a task is never cancelled while something still depends on it. - **Grace budget.** Each level gets a bounded time to finish in-flight work; the total must fit the platform's kill timeout (an orchestrator's termination grace period), otherwise the process is killed mid-drain and you get exactly the truncated-work failure the discipline was meant to prevent. - **Escalation.** Request politely (cancel), then hard-cancel, then abandon and log with the task's name. An unstoppable task must be *visible*, not silently waited on forever. - **Idempotent, bounded cleanup.** Cleanup runs when the deadline has already expired, so it needs its own small budget and must be safe to run twice. ## Supervision: a dead daemon is an orphan too If a long-lived task fails and nobody notices, you have re-created silent failure inside a scope. The owning scope must have a policy: restart with backoff and a failure budget (with a cap, so a permanently broken dependency does not become a hot loop), or mark the component unhealthy and fail readiness so traffic drains away. Either way the failure is *reported* — counted, logged with the scope name, and visible in health checks. "Restart forever, log nothing" is the anti-pattern. ## Bounded concurrency and observability Long-lived scopes are where unbounded growth reappears, because request scopes no longer bound it: a component scope that starts one task per incoming item will happily accumulate millions. Cap it — a fixed worker set consuming a bounded queue, with an explicit policy when the queue is full. And expose the tree: a debug endpoint or dump listing live scopes, their names, child counts, and ages turns "something is still running" into a one-line answer, and makes leaked lifetimes obvious in staging. ## Handing tasks upward Some runtimes let a task be *moved* to an ancestor scope rather than being started there — useful when a request discovers work that must continue after it returns. That is the sanctioned escape hatch: ownership is transferred, never dropped. If your platform lacks it, model the same thing explicitly by submitting a self-contained job (owned data only) to a component-scoped worker. ## The acceptance test At any instant, for every live task, you can name the scope responsible for it, the deadline or cancellation signal that will end it, and where its failure will be reported. If any of the three is "nowhere", it is an orphan wearing a costume.

  • A request discovers work that must continue after the response is sent. What is the sanctioned way to handle it?
    Transfer ownership upward rather than dropping it: either hand the running task to an ancestor scope if the runtime supports it, or package the work as a self-contained job — owned data copies only, no request context or pooled resources — and submit it to a bounded worker pool owned by a component scope. Either way the work has a scope that will cancel and join it at shutdown, and a place where its failure is reported.
  • How do you choose the grace period for draining long-lived scopes at shutdown?
    Work backwards from the hard kill deadline imposed by the platform — for example an orchestrator's termination grace period — and reserve a margin so the process finishes voluntarily rather than being killed mid-drain. Allocate the remainder in reverse dependency order, giving each level roughly its observed p99 in-flight duration, and treat anything that regularly exceeds its slice as a design bug: it needs its work chunked, checkpointed, or made resumable rather than a longer grace period.

saying these in an interview costs you the question

  • "Daemons prove structured concurrency doesn't work in real systems" — they need a longer-lived scope, not no scope.
  • Letting a background task keep the triggering request's context, credentials, or pooled connection.
  • Restarting a failing daemon forever without backoff, a failure budget, or any health signal.
  • Treating shutdown as "wait for everything" with no ordering, no budget, and no escalation path.
  • Assuming long-lived scopes bound concurrency by themselves — they are exactly where unbounded spawning reappears.

context