You run multiple independent background subtasks and one failing must not kill the others. How do you structure the scope, and what are the failure-observability trade-offs?
answer
- Independent + resilient → supervisorScope / SupervisorJob
- Isolation costs the single rethrow → must report per task
- CoroutineExceptionHandler: root launch only, not async
- Supervision applies to direct children only
- Still rethrow CancellationException + cancel scope on shutdown
basics
~10 sUse supervisorScope (or a scope built on SupervisorJob) so one task's failure doesn't cancel the others. The trade-off: each task must handle or report its own errors, or failures get lost.
solid answer
~40 sFor independent subtasks where partial failure is acceptable, use `supervisorScope { }` or a long-lived `CoroutineScope(SupervisorJob() + dispatcher)`. A `SupervisorJob` stops failure from propagating upward, so a failing child does not cancel its siblings or the scope. The cost is **observability**: with isolation, no single rethrow tells you something failed. You must (a) wrap each `launch` body in try/catch, (b) install a `CoroutineExceptionHandler` (fires for uncaught failures in root `launch` coroutines), or (c) for `async`, `try/catch` each `await()`. You also still own cancellation/teardown: cancel the scope on shutdown to avoid leaks, and rethrow `CancellationException` inside any broad catch. Choose `coroutineScope` instead when subtasks are dependent and any failure should abort the batch (fail-fast). The decision is really 'isolate-and-report' vs 'fail-fast-and-rethrow'.
code
kotlin · 20 linesimport kotlinx.coroutines.*
suspend fun main() = supervisorScope {
val results = mutableListOf<String>()
val tasks = listOf("ok1", "bad", "ok2")
tasks.forEach { name ->
launch {
try {
if (name == "bad") throw RuntimeException("task $name failed")
results += name
} catch (e: CancellationException) {
throw e
} catch (e: Exception) {
println("isolated failure: ${e.message}")
}
}
}
// join all, then:
// siblings ok1/ok2 still complete despite 'bad' failing
}go deeper
Knows to reach for supervisorScope so one failure doesn't kill the others.
Builds the supervisorScope with per-child try/catch and understands SupervisorJob isolation.
Adds CoroutineExceptionHandler/async nuances, owned-scope teardown, and cancellation hygiene.
Frames the isolate-vs-fail-fast decision, designs observability and teardown for long-lived background work, and reasons about supervision-not-inherited edge cases.
## The design choice - **Dependent subtasks / fail-fast** → `coroutineScope` (regular Job): first failure cancels the rest and rethrows. One place to catch. - **Independent subtasks / resilient** → `supervisorScope` or a `SupervisorJob`-based scope: failures are isolated; siblings survive. ## Structuring a resilient scope ```kotlin // Short-lived, scoped to a suspend function: suspend fun runAll(tasks: List<suspend () -> Unit>) = supervisorScope { tasks.forEach { task -> launch { try { task() } catch (e: CancellationException) { throw e } // never swallow catch (e: Exception) { reportFailure(e) } // isolate + report } } } // Long-lived, owned by a component: class Worker { private val scope = CoroutineScope( SupervisorJob() + Dispatchers.Default + CoroutineExceptionHandler { _, e -> log.error("uncaught", e) } ) fun start(task: suspend () -> Unit) = scope.launch { task() } fun shutdown() = scope.cancel() // structured teardown, no leaks } ``` ## The observability trade-off Isolation removes the single rethrow that would otherwise tell you 'the batch failed'. To keep visibility: - **Per-child try/catch** in each `launch` body — most explicit; lets you decide retry/skip/metric per task. - **CoroutineExceptionHandler** — catches *uncaught* exceptions from **root** `launch` coroutines under the SupervisorJob. It does **not** fire for `async` (the exception lives in the `Deferred`) nor for children whose parent already handles failures. - **async**: you must `try { d.await() }` per task; an un-awaited failing `async` under a SupervisorJob is effectively silent. ## SupervisorJob nuance: only the direct hierarchy is supervised A `SupervisorJob` only changes propagation for its **direct** children. If a supervised child itself opens a plain `coroutineScope` internally, failures within that inner scope still cancel that inner scope normally — supervision is not inherited transitively into nested regular scopes. ## Cancellation correctness still applies Resilience does not exempt you from cancellation hygiene: - Rethrow `CancellationException` inside broad catches (or use `ensureActive()`), otherwise shutdown cancellation is swallowed. - For suspending cleanup during teardown, use `withContext(NonCancellable)`. - Cancel the owned scope on component shutdown to prevent leaks. ## Picking the model — checklist - Must all-or-nothing? → coroutineScope. - Independent, tolerate partial failure? → supervisorScope / SupervisorJob. - Need a failure signal? → per-child try/catch or CoroutineExceptionHandler. - Long-lived background work? → owned CoroutineScope(SupervisorJob()+dispatcher), cancel on teardown. ## Summary Supervision buys resilience at the cost of explicit per-task error handling and reporting. Always pair it with cancellation hygiene and a deliberate teardown story.
- If a supervised child internally calls coroutineScope and a grandchild fails, are the child's siblings protected?The grandchild failure cancels the inner coroutineScope (regular Job) and propagates out of that child as a normal failure. Whether siblings of the *child* survive depends on the SupervisorJob: yes, the SupervisorJob isolates the child's failure from its siblings, but the inner scope itself was torn down.
- Why might an installed CoroutineExceptionHandler never fire for your failing tasks?Because the tasks use async (exception is stored in Deferred, not 'uncaught'), or because each child catches its own exception, or the handler was placed on a non-root coroutine. It only handles uncaught exceptions of root launch coroutines.
supervisorScope is a manager who lets each report fail without firing the team, but then it's the manager's job to notice and log who failed.
saying these in an interview costs you the question
- Recommending supervisorScope but forgetting that failures now need explicit reporting
- Expecting CoroutineExceptionHandler to catch async exceptions
- Assuming supervision is inherited into nested coroutineScope blocks
- Building a long-lived scope and never cancelling it on shutdown (leak)
- Swallowing CancellationException inside the per-task catch