skip to content

You own a CoroutineScope(SupervisorJob() + Dispatchers.IO) in a server component and need a graceful shutdown that stops accepting new work and waits for in-flight coroutines to finish (with a timeout) before the process exits. How do you implement it?

level: principalimportance: should knowfreq 28%

answer

  1. cancel() returns immediately - not a wait
  2. 3 phases: quiesce, drain/deadline-cancel, release
  3. join children with withTimeoutOrNull(grace)
  4. cancelChildren() on timeout, then scope.cancel()
  5. NonCancellable for suspending cleanup

basics

~10 s

Stop launching new work, cancel the children, then wait for the scope's Job to finish using join with a timeout. If it overruns, force-cancel and exit. Cancel the scope itself last.

solid answer

~40 s

scope.cancel() returns immediately — it doesn't wait for children to wind down — so a naive teardown can drop in-flight work. For graceful shutdown: (1) flip a flag so no new coroutines are launched; (2) decide between draining (let current work finish) or cancelling (stop it). To drain, capture the scope's Job and runBlocking/await with a timeout: withTimeoutOrNull(grace) { job.children.toList().joinAll() } or join the Job after stopping new launches. To hard-stop, call scope.coroutineContext.cancelChildren() then join with a timeout. Because cancel() is non-suspending and fire-and-forget, you typically do the waiting in a runBlocking block during JVM shutdown (@PreDestroy / shutdown hook). Use withContext(NonCancellable) for any cleanup that must run. Finally scope.cancel() to release ownership. The key insight is separating 'stop new work', 'finish/stop existing work within a deadline', and 'release the scope' as distinct steps.

code

kotlin · 9 lines
kotlin
fun shutdownGracefully(grace: Duration) = runBlocking {
    accepting = false                          // 1. stop new work
    val job = scope.coroutineContext[Job]!!
    val drained = withTimeoutOrNull(grace) {   // 2. drain within deadline
        job.children.toList().joinAll()
    }
    if (drained == null) job.cancelChildren()  //    overran -> stop them
    scope.cancel()                             // 3. release ownership
}

go deeper

for a junior

Knows cancel() stops work but likely can't implement a bounded graceful drain.

for a middle

Can stop new work and cancel children with a timeout, but may miss the non-suspending nature of cancel() or NonCancellable cleanup.

for a senior

Implements the quiesce/drain/release phases correctly with withTimeoutOrNull and joinAll, handling cleanup.

for a principal

Designs the shutdown contract end-to-end (platform hook, deadlines, SupervisorJob drain semantics, observability) and standardizes it across services.

## Why naive scope.cancel() isn't graceful `scope.cancel()` cancels the Job and returns **immediately**; it does not block until children stop. During JVM shutdown that means in-flight requests can be abandoned mid-write. Graceful shutdown needs three explicit phases: 1. **Quiesce** — stop accepting/launching new work. 2. **Drain or deadline-cancel** — let in-flight coroutines finish, bounded by a timeout. 3. **Release** — cancel the scope to free ownership. ## Phase 1 — stop new work Guard your launch path so nothing new starts: ```kotlin @Volatile private var accepting = true fun submit(task: suspend () -> Unit) { if (!accepting) return scope.launch { task() } } ``` Set `accepting = false` at shutdown. ## Phase 2a — drain (preferred for graceful) Wait for current children to complete, bounded by a grace period. `cancel()` is non-suspending, so do the waiting from a blocking context (shutdown hook / `@PreDestroy`): ```kotlin fun shutdownGracefully(grace: Duration) = runBlocking { accepting = false val job = scope.coroutineContext[Job]!! val finished = withTimeoutOrNull(grace) { job.children.toList().joinAll() // wait for in-flight work } != null if (!finished) job.cancelChildren() // deadline exceeded -> stop them scope.cancel() // phase 3: release } ``` `joinAll()` suspends until each child completes; `withTimeoutOrNull` caps the wait without throwing. ## Phase 2b — deadline-cancel (stop in-flight, then wait briefly) If you'd rather cancel immediately but still ensure cleanup runs: ```kotlin scope.coroutineContext.cancelChildren() withTimeoutOrNull(grace) { job.children.toList().joinAll() } ``` Children get `CancellationException`; their `finally`/`use` blocks run (use `withContext(NonCancellable)` for cleanup that itself suspends). ## Phase 3 — release `scope.cancel()` last so the scope's own Job reaches the terminal state and ownership is gone. After this the scope is single-use, which is correct at process end. ## SupervisorJob interaction With a `SupervisorJob`, draining via `job.children.joinAll()` is robust because one child failing during drain won't cancel the others — you still drain the rest within the deadline. ## Wiring into the platform - Spring: `@PreDestroy fun close() = shutdownGracefully(10.seconds)`. - Plain JVM: `Runtime.getRuntime().addShutdownHook(Thread { shutdownGracefully(...) })`. ## Key APIs/keywords - `Job.children` — sequence of direct child Jobs. - `joinAll()` / `Job.join()` — suspend until completion. - `withTimeoutOrNull(duration)` — bound the wait, return null on timeout (no throw). - `cancelChildren()` vs `cancel()` — stop children vs. terminate the scope. - `NonCancellable` — context so cleanup completes during cancellation. - `runBlocking` — bridge from non-suspending shutdown hook to suspending drain.

  • Why can't you just await scope.cancel() to know children are done?
    cancel() is non-suspending and returns immediately. To wait you must join the scope's Job or its children separately, ideally with a timeout.
  • How do you ensure a coroutine's cleanup (e.g. flushing a buffer) runs even though it's being cancelled?
    Put the cleanup in a finally/use block and wrap suspending cleanup in withContext(NonCancellable) so it isn't itself cancelled.
  • How does SupervisorJob help during a drain?
    If one child fails while draining, SupervisorJob keeps the siblings running, so the rest still finish within the grace period instead of being torn down.

Like closing a shop: lock the door to new customers, let those inside finish (up to closing time), then usher out anyone remaining and turn off the lights.

saying these in an interview costs you the question

  • Assumes scope.cancel() blocks until children finish
  • No timeout/deadline on the drain (can hang shutdown forever)
  • Forgets to stop accepting new work first
  • Doesn't use NonCancellable for cleanup that suspends
  • Calls scope.cancel() before draining, abandoning in-flight work

context