skip to content

Structured Concurrency

The rule that every coroutine belongs to a scope and a scope cannot finish before its children — which is what makes leaks, orphans, and forgotten cancellation structurally impossible. It is the idea interviewers most want you to articulate.

part ofKotlinoverview, primer and where to startread it →
on this pageshow

explore

questions

25

What is a Job in Kotlin coroutines, and what are the states in its lifecycle?

level: juniorimportance: must knowfreq 70%

answer

  1. launch -> Job, async -> Deferred (a Job)
  2. States: New/Active/Completing/Completed/Cancelling/Cancelled
  3. No enum: isActive / isCompleted / isCancelled
  4. Cancelled is also Completed (terminal)
  5. Job lives in CoroutineContext: coroutineContext[Job]

basics

~10 s

A Job is a handle to a running coroutine. It can be active, can complete, or can be cancelled. You can wait for it, cancel it, or check if it is still running.

solid answer

~40 s

A Job is the lifecycle handle for a coroutine, returned by launch (async returns a Deferred, a Job subtype). It moves through states: New (created but not started, e.g. with start = CoroutineStart.LAZY), Active (running), Completing (its body finished but it waits for children), Completed, Cancelling, and Cancelled. You inspect it via isActive, isCompleted, and isCancelled, control it with start(), cancel(), and join() (suspends until done), and react with invokeOnCompletion {}. A Job lives in the CoroutineContext, so a coroutine can read its own Job via coroutineContext[Job]. Most states are not directly observable as an enum — the three booleans together encode them.

code

kotlin · 10 lines
kotlin
val job = scope.launch(start = CoroutineStart.LAZY) {
    delay(100)
}
println(job.isActive)    // false (New)
job.start()
println(job.isActive)    // true (Active)
job.cancel()
job.join()
println(job.isCancelled) // true
println(job.isCompleted) // true (Cancelled is terminal)

go deeper

for a junior

Knows launch returns a Job and can name active/cancelled/completed and use cancel()/join().

for a middle

Explains the Completing state and that a cancelled Job is also completed; reads state via the three booleans.

for a senior

Distinguishes Job vs Deferred, knows the Job lives in the context, and uses invokeOnCompletion with its cause semantics.

for a principal

Reasons about why the API exposes orthogonal booleans rather than an enum, and how Completing underpins structured concurrency guarantees.

## What a Job is A `Job` is a cancellable handle to the lifecycle of a coroutine — a background unit of work. When you call `launch { ... }` you get back a `Job`; `async { ... }` returns a `Deferred<T>`, which **is a Job** (a subtype that also carries a result via `await()`). A `Job` is also stored in the coroutine's `CoroutineContext`, so the running coroutine can find it with `coroutineContext[Job]`. ## The lifecycle states A Job conceptually passes through these states: - **New** — created but not started yet. Only happens with `CoroutineStart.LAZY`. Call `start()` or `join()` to begin it. - **Active** — started and running (the default after `launch`). - **Completing** — the coroutine's own code finished, but it is **waiting for its children** to finish before it can complete. This is internal/transient. - **Completed** — fully done, including all children. - **Cancelling** — `cancel()` was called (or an exception was thrown); it is winding down. - **Cancelled** — finished because it was cancelled. Cancelled is a **terminal, completed** state. ## Observing the state There is no public enum. You read three boolean properties: ```kotlin val job = scope.launch { /* work */ } println(job.isActive) // true while running println(job.isCompleted) // true once finished (normally OR cancelled) println(job.isCancelled) // true if cancelled or failed ``` Note the subtlety: a **cancelled** Job has `isCompleted == true` AND `isCancelled == true`. "Completed" here means "reached a terminal state", not "succeeded". ## Controlling a Job - `start()` — begin a lazy job; returns `true` if it actually started it. - `join()` — a `suspend` function that suspends the caller until the job is completely done (does not throw on cancellation). - `cancel(cause)` — requests cancellation. - `invokeOnCompletion { cause -> ... }` — register a callback fired when the job completes; `cause` is `null` on normal completion, a `CancellationException` if cancelled, or the failure. ```kotlin val job = scope.launch(start = CoroutineStart.LAZY) { doWork() } // state: New job.start() // -> Active job.invokeOnCompletion { cause -> println("done, cause=$cause") } job.join() // suspend until terminal ``` Understanding these states is the foundation for parent/child links: a parent sits in **Completing** while it waits for children, which is what makes structured concurrency work.

  • Why is there no single State enum exposed?
    The library exposes three orthogonal booleans (isActive/isCompleted/isCancelled) because some states are internal/transient (Completing, Cancelling) and the combinations cover what callers actually need.
  • What does join() do if the job was cancelled?
    join() simply returns when the job reaches a terminal state; it does not re-throw the cancellation cause. await() on a Deferred, by contrast, re-throws the failure.

A Job is like a tracking number for a delivery: you can check its status, cancel it, or wait at the door until it arrives.

saying these in an interview costs you the question

  • Saying isCompleted is false for a cancelled job
  • Thinking launch returns a Deferred / async returns a plain Job
  • Claiming there is a public Job.State enum to switch on
  • Believing a coroutine starts Active when created with LAZY
  • Confusing join() (waits, no throw) with await() (waits, re-throws)

context

open as a page

What does the coroutineScope { } builder do, and how is it different from launching with GlobalScope?

level: juniorimportance: must knowfreq 65%

basics

~20 s

coroutineScope runs a block, starts child coroutines inside it, and waits until all of them finish before returning. GlobalScope launches coroutines that are not tied to anything and can outlive your code, so they can leak.

open as a page

Why does an object with a lifecycle (e.g. a ViewModel, a service, a request handler) usually own its own CoroutineScope, and what does cancelling that scope accomplish?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The object creates a scope to launch background work tied to its life. When the object is done, it cancels the scope, which stops every coroutine started in it so nothing keeps running and leaking.

open as a page

What is structured concurrency in Kotlin coroutines, and what core guarantee does it give you?

level: juniorimportance: must knowfreq 80%

basics

~10 s

Every coroutine runs inside a scope, and that scope won't finish until all the coroutines it started have finished. This stops background work from being forgotten or leaking.

open as a page

What is a SupervisorJob in Kotlin coroutines, and how does it differ from a regular Job when one child fails?

level: juniorimportance: must knowfreq 70%

basics

~10 s

A SupervisorJob is a special parent for coroutines. If one child crashes, the others keep running and the parent stays alive. With a regular Job, one crash cancels everything.

open as a page

Why does a parent coroutine not complete until its children finish, and what is the 'Completing' state?

level: middleimportance: must knowfreq 55%

basics

~20 s

A parent waits for all its children to finish before it counts as done. Even after the parent's own code runs out, it stays in a waiting phase until every child it started has completed.

open as a page

What is the difference between coroutineScope and supervisorScope when one of their child coroutines throws an exception?

level: middleimportance: must knowfreq 78%

basics

~10 s

In coroutineScope, if one child fails, all the other children are stopped and the failure is rethrown. In supervisorScope, a failing child does not stop its siblings; each child fails independently.

open as a page

When constructing CoroutineScope(SupervisorJob() + dispatcher) for a long-lived owner, why is SupervisorJob preferred over a plain Job, and what is the consequence at the scope level?

level: middleimportance: must knowfreq 62%

basics

~10 s

With a plain Job, one child crashing cancels the whole scope and its siblings. With SupervisorJob, a child's failure stays local — siblings keep running and the scope stays alive to launch more work.

open as a page

Contrast launching work with coroutineScope { } versus GlobalScope.launch { }. Which respects structured concurrency, and what goes wrong with the other?

level: middleimportance: must knowfreq 70%

basics

~10 s

coroutineScope ties work to the current call and waits for it to finish. GlobalScope ties work to the whole app, so nobody waits for it or cancels it, and it can leak.

open as a page

How does supervisorScope { } differ from coroutineScope { } with respect to child failures?

level: middleimportance: must knowfreq 65%

basics

~10 s

Both wait for all child coroutines to finish. In coroutineScope, one child failing cancels all the others and throws. In supervisorScope, a failing child does not cancel its siblings.

open as a page

How does cancellation propagate through a Job hierarchy when you cancel a parent?

level: seniorimportance: must knowfreq 60%

basics

~10 s

Cancelling a parent cancels all of its children, and their children, all the way down the tree. Each coroutine stops at its next suspension point by getting a cancellation exception.

open as a page

Why is CancellationException treated specially during structured exception propagation, and what bug does swallowing it cause?

level: seniorimportance: must knowfreq 70%

basics

~20 s

Cancellation works by throwing a special CancellationException. It only stops the coroutine being cancelled, not its siblings, and the parent treats it as normal. If your catch block swallows it, the coroutine ignores cancellation and keeps running.

open as a page

scope.cancel() is called but a child coroutine keeps running. Explain why coroutine cancellation is cooperative, and how to make CPU-bound or blocking code actually stop.

level: seniorimportance: must knowfreq 55%

basics

~10 s

Cancellation only takes effect at suspension or check points. A tight loop or a blocking call never checks, so it keeps running. You make it stop by suspending, checking isActive, or calling ensureActive()/yield().

open as a page

What is the difference between scope.cancel() and scope.coroutineContext.cancelChildren() for a scope you own, and when would you choose each?

level: middleimportance: should knowfreq 40%

basics

~10 s

cancel() cancels the scope's Job itself, so the scope is dead and can't be reused. cancelChildren() stops the current coroutines but leaves the Job alive, so you can launch new work afterward.

open as a page

Explain precisely when a coroutineScope { } block returns. If the block's last statement runs but a launched child is still working, what happens?

level: middleimportance: should knowfreq 55%

basics

~10 s

The block returns only after every coroutine started inside it has finished. Even if the last line ran, if a child is still working the block keeps waiting for it.

open as a page

What is the practical difference between a coroutine launched as a structured child versus one given its own root Job, and why does it matter for lifecycle?

level: seniorimportance: should knowfreq 40%

basics

~20 s

A structured child is tracked by its parent: the parent waits for it and cancels it. A coroutine given its own fresh Job becomes independent, so the parent neither waits for it nor cancels it — risking leaks.

open as a page

Inside coroutineScope you start two async tasks; one throws before you call await(). Does the scope fail, and when do you observe the exception?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Yes — because async is still a child of the scope, the failure cancels the scope and its siblings right away. But the actual exception value is re-thrown when you call await() on that task.

open as a page

How does the structured concurrency principle shape the design of suspend functions that do internal concurrency? Why shouldn't such a function expose its own background coroutines?

level: seniorimportance: should knowfreq 45%

basics

~10 s

A suspend function should finish all the concurrency it starts before it returns, using coroutineScope. That way callers never inherit hidden background work they didn't ask for.

open as a page

Where do uncaught exceptions from launch children go under a SupervisorJob, and how do you handle them correctly?

level: seniorimportance: should knowfreq 50%

basics

~10 s

A launch child's uncaught exception under a SupervisorJob goes to a CoroutineExceptionHandler. The handler must be on the child (or its scope), not added when you call launch on a coroutineScope's child.

open as a page

A developer writes scope.launch(SupervisorJob()) { ... } expecting that child's failures to be isolated. Why is this often a mistake?

level: seniorimportance: should knowfreq 35%

basics

~10 s

Passing a SupervisorJob into launch makes it the coroutine's parent job, which breaks structured concurrency: the new coroutine is no longer tied to the scope, so cancelling the scope won't cancel it.

open as a page

You run multiple independent background subtasks and one failing must not kill the others. How do you structure the scope, and what are the failure-observability trade-offs?

level: principalimportance: should knowfreq 42%

basics

~10 s

Use supervisorScope (or a scope built on SupervisorJob) so one task's failure doesn't cancel the others. The trade-off: each task must handle or report its own errors, or failures get lost.

open as a page

You own a CoroutineScope(SupervisorJob() + Dispatchers.IO) in a server component and need a graceful shutdown that stops accepting new work and waits for in-flight coroutines to finish (with a timeout) before the process exits. How do you implement it?

level: principalimportance: should knowfreq 28%

basics

~10 s

Stop launching new work, cancel the children, then wait for the scope's Job to finish using join with a timeout. If it overruns, force-cancel and exit. Cancel the scope itself last.

open as a page

Before structured concurrency, ad-hoc background tasks (raw threads, GlobalScope, detached callbacks) were the norm. What concrete classes of bugs does the structured concurrency principle eliminate, and what is the cost or constraint it imposes?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

It removes leaked and orphaned tasks, lost errors, and forgotten cancellation, by making every task have an owner that waits for it. The cost is that you must always have a scope and your code suspends until children finish.

open as a page

How would you design a long-lived service scope so that independent background tasks fail in isolation, are observable, and are all cancelled on shutdown?

level: principalimportance: nice to knowfreq 25%

basics

~10 s

Build the scope on a SupervisorJob plus a dispatcher and a CoroutineExceptionHandler so one task's failure doesn't kill the others. On shutdown, cancel the scope to stop everything, and log failures via the handler.

open as a page