Explain how a cancellation signal travels through a tree of parent and child tasks. When is a cancellation considered complete, and what should a parent do while its subtree is winding down?
answer
- down = cancel, up = failure
- late spawns start cancelled
- signal fast, stop slow
- complete = whole subtree terminated
- detached spawn escapes the tree
basics
~20 sCancelling a node marks it and recursively signals every descendant, and tasks spawned afterwards start already-cancelled. Each leaf stops at its next checkpoint and cleans up. Cancellation is complete only when the whole subtree has terminated, so the parent must wait — with a grace deadline and escalation for stragglers.
solid answer
~60 sCancellation propagates **downward** through the ownership tree. Marking a node sets its signal and recursively signals its children; anything spawned into that scope after the fact starts in the cancelled state, which closes the race where a task is created just as cancellation begins. Signalling is instant; stopping is not. Each descendant reaches its next checkpoint, unwinds, runs cleanup, and reports its outcome. The **subtree's** cancellation latency is therefore the worst path through it: one un-checkpointed leaf or one slow cleanup pins everything above it. Cancellation is complete only when **every descendant has terminated** and released its resources — that is the structured-concurrency invariant, so the parent must join, not just signal. A parent that returns after calling cancel has orphaned its subtree and lied about being done. While waiting, the parent: enforces a grace deadline; escalates when it expires (close the resources the stragglers are blocked on, log which tasks are stuck); and preserves the original cause rather than letting cancellation exceptions from children overwrite it. Tasks spawned into an unrelated global scope escape the tree entirely — that is how the invariant is usually broken in practice.
code
text · 7 linesscope.cancel() # marks scope + all descendants; later spawns start cancelled
if not scope.join(within=GRACE):
for t in scope.stuck_tasks():
log("missed grace deadline", t.name, t.current_op)
t.resource.close_hard() # force the blocked call to fail
scope.join(within=HARD_LIMIT) # last wait before alarming
# only here is cancellation actually completego deeper
Say the signal goes down the tree to all descendants and that the parent has to wait for them to actually finish before it is done.
Add the poisoned-spawn rule for late children, the distinction between signalling and stopping, and that subtree latency is the worst path through the tree.
Cover the grace deadline and the escalation ladder, diagnosing which task missed the deadline, preserving the original cause under a flood of cancellation outcomes, and how detached spawns break the invariant.
Own the budget: allocate the shutdown grace across tree depth, mandate that all spawning goes through owned scopes, and instrument missed-deadline tasks so uncancellable code is found before it pins a deployment.
## The tree In structured concurrency, tasks form a tree by ownership: a scope owns the tasks started inside it, those tasks may open their own scopes, and no task outlives the scope that created it. Cancellation rides that structure. ## Downward propagation Cancelling a node does three things: 1. **Marks the node** — its cancellation state becomes set, so its own checkpoints will observe it. 2. **Recursively signals descendants** — every child scope and task under it is marked too. This is why cancelling the root of a request handler stops the eight fan-out calls, their retries, and their sub-tasks, without any of them knowing about the others. 3. **Poisons future spawns** — a task started into an already-cancelled scope begins cancelled (or is rejected outright). Without this rule there is a race: cancellation sweeps the children, and a task created a microsecond later never gets the signal and escapes. Note the direction. *Downward* is cancellation. *Upward* — a child failing and causing its siblings and parent to stop — is failure propagation, a different policy that typically triggers downward cancellation as its mechanism. ## Signalling is fast; stopping is not Setting the signal is O(nodes) and immediate. Actually stopping is bounded by each leaf's checkpoint interval plus its cleanup time (see cancellation latency). Aggregated over a tree: **subtree latency ≈ max over all root-to-leaf paths of (checkpoint gap + cleanup + any shielded spans)** One badly written leaf — a 40-second compute loop with no poll, a blocking call that ignores the signal, an unbounded shielded cleanup — sets the floor for its entire ancestry. This is why cancellability is a whole-system property: your handler cannot shut down faster than its slowest descendant. ## When is cancellation complete? The useful definition: **when every descendant has reached a terminal state and released its resources.** Not when cancel() returned; not when the parent's own code finished. This is exactly the structured-concurrency invariant — a scope does not exit until all tasks inside it have ended — and cancellation is the case that tests it. A parent that signals and returns has: - orphaned tasks still holding connections and still producing side effects; - resources it is about to close that a child may still be using (use-after-close); - reported "stopped" to its own parent while its subtree is alive, so shutdown proceeds over live work. So the parent's obligation while winding down is to **join**: wait for each child's outcome, then complete. ## What the parent does while waiting - **Wait, with a deadline.** The grace period should be an explicit budget, and the parent should not wait forever on principle. - **Escalate at the deadline.** The escalation ladder is: signal → wait → take away what the straggler is blocked on (close its socket/handle, which makes its blocked call fail) → give up on it and alarm → at the extreme, restart the process. Escalation matters because cooperation cannot be enforced. - **Diagnose, don't hide.** Log *which* task is stuck, with enough identity (name, id, current operation, ideally a stack) to fix it later. A metric for "tasks that missed the grace deadline" turns an invisible class of bug into a tracked one. - **Preserve causality.** Children being cancelled will report cancellation outcomes. If the cancellation was triggered by a real failure somewhere, that failure is the cause; the flood of cancellation outcomes from siblings is a consequence and must not overwrite or outrank it in what the parent reports. - **Do not start new work.** A parent in the process of cancelling should not spawn additional children — the poisoned-spawn rule usually enforces this, and code that tries is a design smell. ## How the tree gets broken in practice - **Detached spawns.** Submitting to a global executor, a module-level scope, or a fire-and-forget helper hands the task to a different owner. It no longer receives the signal, and the parent no longer waits for it. This is the single most common way a nicely structured system develops unstoppable work. - **Handing a raw handle upward** so that a longer-lived object keeps the child alive past its scope. - **A shield around a whole subtree**, making an entire branch immune. - **Callbacks scheduled on timers or event loops** that outlive the scope that registered them, unless the scope explicitly unregisters them during cleanup. ## A concrete picture Root handler with a 5-second client timeout fans out to three services; one of those does its own parallel fetch. The client disconnects at t=1s. The root is cancelled; the signal reaches all six nodes within microseconds. Two leaves are awaiting I/O and abort immediately. One leaf is mid-compute and stops 30 ms later at its next checkpoint. One leaf's cleanup releases a distributed lease under a shielded 500 ms budget. The root completes at roughly t=1.5s — not at t=1s — and only then is the request truly finished.
- Why must a task spawned into an already-cancelled scope start out cancelled?Otherwise there is a race window: cancellation walks the existing children, and a task created just after that walk never receives the signal and escapes the tree. Poisoning future spawns closes the window without any locking between the canceller and the spawner, and it also expresses the intent — the scope is winding down, so no new work belongs in it.
- Where does the boundary sit between a parent waiting for a subtree and simply giving up on it?At an explicit grace deadline, followed by escalation rather than an indefinite wait. When the deadline passes, force the issue by closing the resources the stragglers are blocked on, record which tasks missed the deadline with enough identity to debug them, and only then abandon them with an alarm — with process restart as the final option. Waiting forever converts one uncancellable task into a stuck deployment.
saying these in an interview costs you the question
- Returns from the parent as soon as cancel is called, without joining the children
- Assumes propagation is instantaneous because the signal is
- Spawns into a global executor and still expects the parent's cancellation to reach it
- Lets a flood of child cancellation outcomes overwrite the original failure cause
- Waits indefinitely for stragglers with no grace deadline or escalation