skip to content

How would you design a long-lived service scope so that independent background tasks fail in isolation, are observable, and are all cancelled on shutdown?

level: principalimportance: nice to knowfreq 25%

answer

  1. Scope = SupervisorJob + dispatcher + CoroutineExceptionHandler + CoroutineName
  2. Supervisor isolates; handler observes; cancel shuts down
  3. shutdown(): cancelAndJoin on the scope's Job
  4. Restartable task = try/catch loop, rethrow CancellationException
  5. viewModelScope/lifecycleScope are this pattern

basics

~10 s

Build the scope on a SupervisorJob plus a dispatcher and a CoroutineExceptionHandler so one task's failure doesn't kill the others. On shutdown, cancel the scope to stop everything, and log failures via the handler.

solid answer

~40 s

Create the scope as `CoroutineScope(SupervisorJob() + dispatcher + CoroutineExceptionHandler { _, e -> log/metric(e) } + CoroutineName(...))`. The SupervisorJob isolates child failures so independent tasks don't take each other down; the handler makes every uncaught failure observable (logged/metered); the dispatcher bounds where work runs. For shutdown, expose a `close()`/`stop()` that calls `scope.cancel()` (optionally `scope.coroutineContext[Job]?.cancelAndJoin()` to await orderly stop) — because cancellation still flows downward through a SupervisorJob, this stops all children. Keep the scope's Job as a field so its lifecycle is tied to the component (Android ViewModel does this with viewModelScope, a SupervisorJob-backed scope). For restartable tasks, wrap each in its own try/catch or a relaunch loop with backoff. Avoid passing Jobs into individual launches (that detaches them). This is the canonical resilient-scope pattern.

go deeper

for a junior

Can name SupervisorJob as the way to keep tasks independent.

for a middle

Assembles scope = SupervisorJob + dispatcher + handler and cancels on shutdown.

for a senior

Adds observability, cancelAndJoin for orderly stop, and per-task retry with correct cancellation handling.

for a principal

Reasons about lifecycle ownership, testability via injected dispatchers, anti-patterns, and ties it to real frameworks like viewModelScope.

## Goals 1. **Isolation** — one task's failure must not cancel the others. 2. **Observability** — every failure is logged/metered, never silently lost. 3. **Clean shutdown** — a single call cancels everything and (optionally) waits. 4. **No leaks** — every task is in the scope's tree. ## The scope ```kotlin import kotlinx.coroutines.* class BackgroundService( dispatcher: CoroutineDispatcher = Dispatchers.Default, ) { private val handler = CoroutineExceptionHandler { ctx, e -> log.error("task failed in ${ctx[CoroutineName]}", e) metrics.increment("task.failure") } private val scope = CoroutineScope( SupervisorJob() + dispatcher + handler + CoroutineName("bg-service"), ) fun startPoller() = scope.launch(CoroutineName("poller")) { /* ... */ } fun startSync() = scope.launch(CoroutineName("sync")) { /* ... */ } suspend fun shutdown() { scope.coroutineContext[Job]?.cancelAndJoin() // cancels all children, awaits } } ``` ## Why each piece - **SupervisorJob()** — children fail independently. A crash in `poller` does not cancel `sync`. - **CoroutineExceptionHandler** — under a SupervisorJob, direct children are roots, so this handler receives their uncaught failures. Without it, failures hit the thread's default uncaught handler. - **CoroutineName** — makes log lines and thread dumps identify the failing task. - **dispatcher** — bounds threading; inject it for testability (`StandardTestDispatcher`). - **cancelAndJoin()** on shutdown — cancellation flows downward through the SupervisorJob, stopping all children; `join` awaits orderly completion. ## Restartable tasks The handler is terminal — it cannot resume a failed coroutine. For tasks that must keep running, build resilience **inside** the task: ```kotlin scope.launch(CoroutineName("poller")) { while (isActive) { try { pollOnce() } catch (e: CancellationException) { throw e } // never swallow cancellation catch (e: Exception) { log.warn("retrying", e); delay(backoff) } } } ``` Note the explicit rethrow of `CancellationException` — swallowing it would break cooperative cancellation. ## Anti-patterns to avoid - `scope.launch(SupervisorJob()) { }` — reparents/detaches the child; it leaks past shutdown. - Catching `Throwable` (or a bare `catch (e: Exception)`) **without** rethrowing `CancellationException` — defeats cancellation. - Relying on the handler to recover state. - Sharing one scope across unrelated lifecycles. ## Real-world reference Android's `viewModelScope` and `lifecycleScope` are exactly this: SupervisorJob-backed scopes tied to a component lifecycle, cancelled when the component is destroyed. Server-side, a per-request or per-component scope follows the same shape.

  • Why must the per-task retry loop rethrow CancellationException?
    Cancellation is delivered as a CancellationException. Swallowing it in a broad catch makes the task ignore cancellation, leaking work and breaking shutdown. Always rethrow it (or catch it separately).
  • Why inject the dispatcher rather than hardcode Dispatchers.Default?
    Testability and control: in tests you can supply a TestDispatcher to control virtual time, and in production you can route work to a bounded pool or a confined dispatcher.

saying these in an interview costs you the question

  • No CoroutineExceptionHandler -> failures vanish into the default handler
  • Using a plain Job so one failure cancels the whole service
  • Passing SupervisorJob() into launch and detaching tasks
  • Catching CancellationException without rethrowing
  • Expecting the handler to restart a failed task

context