skip to content

questions

19

In a concurrent system, what does it actually mean to cancel a running task, and why is cancelling a task different from simply ignoring the result it eventually produces?

level: juniorimportance: must knowfreq 55%

answer

  1. request, not a kill
  2. abandon result ≠ stop work
  3. third outcome: cancelled
  4. resources keep burning
  5. side effects still land

basics

~20 s

Cancelling means asking a running task to stop early and release what it holds; the task must notice the request and wind down. Ignoring its result stops nothing: it keeps burning CPU, connections and memory, and its side effects still happen.

solid answer

~50 s

Cancellation is a **request delivered to a running task** telling it the result is no longer wanted and it should unwind. It has two halves: someone sets a cancellation signal, and the task observes that signal at safe points, abandons the work, and runs its cleanup on the way out. Abandoning the result is a completely different thing. If the caller stops waiting but the task keeps running, you have work nobody owns: it still holds a connection, still writes the file, still sends the outbound request, still grows the heap. Under load that becomes pool exhaustion, duplicated side effects, and a shutdown that never finishes because those tasks never end. So a task has three outcomes, not two: succeeded, failed, or **cancelled**. Cancelled is its own outcome. Normally you do not retry it and do not log it as a system error, because the caller is the one who asked for it.

code

text · 10 lines
text
# ignore the result — task keeps running
handle = spawn(long_job)
wait_for(handle, 2s) or return FALLBACK   # long_job still running, still holding a connection

# cancel — task is told to stop and we wait for it to finish stopping
handle = spawn(long_job)
if not wait_for(handle, 2s):
    handle.cancel()        # sets the signal
    handle.join()          # wait until the task has actually unwound + cleaned up
    return FALLBACK

go deeper

for a junior

Say cancellation is a request the task must observe, name the three outcomes, and give one concrete cost of abandoning work instead (a leaked connection or a duplicate side effect).

for a middle

Add the signal/observe split, cancellation latency, and why 'cancelled' has to be kept distinct from 'failed' in logs and metrics.

for a senior

Talk about what it means for shutdown and pool exhaustion, escalation after a grace period by closing the underlying resource, and how you audit which tasks in a service are actually cancellable.

for a principal

Frame it as an ownership invariant: every task has an owner that can cancel it and observes its outcome; discuss where the system draws hard boundaries (process, connection, resource close) because cooperation is not a guarantee.

## What a task is, and what cancelling one means A concurrent task is a unit of work running independently of the code that started it: a thread, a lightweight/green thread, a coroutine, an actor processing a message, a job submitted to a worker pool. **Cancellation** is the mechanism by which something outside that task — its parent, a timeout, a user who closed the connection, a shutdown routine — communicates *stop; the result is no longer wanted; release what you hold*. In essentially every mainstream runtime this is a **request, not a kill**. The signal is recorded somewhere the task can see (a flag, a token, a per-task or per-thread bit). The task is responsible for looking at it and reacting. That split — one side signals, the other side observes — is the whole model, and it is why cancellation is described as *cooperative*. ## Why abandoning the result is not cancellation A very common beginner move is to stop waiting: return early from the caller, drop the handle, resolve the caller's own result with a fallback. Nothing about that reaches the running task. It continues to: - **Consume resources.** CPU on a busy pool, a checked-out database connection, an open socket, a file handle, buffers on the heap. Ten abandoned tasks per second against a ten-connection pool is an outage. - **Perform side effects.** It still sends the payment call, still writes the row, still publishes the event. The user saw a timeout; the system did the work anyway. Now you have state the user does not expect, sometimes twice if they retried. - **Prevent shutdown.** Graceful shutdown means "wait for in-flight work to end". Work nobody can stop never ends, so you either hang or hard-kill the process mid-write. - **Hide failures.** When it eventually throws, nobody is listening; the error lands in a default handler or vanishes. So: *ignoring a result is a decision about the caller; cancellation is a decision about the work.* ## The three outcomes Modelling cancellation properly means a task terminates in one of three ways: 1. **Success** — produced a value. 2. **Failure** — threw or returned an error caused by the work itself. 3. **Cancelled** — stopped because it was asked to. Keeping (3) distinct from (2) matters operationally. Cancelled tasks should not page anyone, should not usually be retried (the caller already gave up or a shutdown is in progress), and their "cancelled" exception or status must not overwrite an *earlier real failure* that triggered the cancellation in the first place. Conflating the two produces dashboards full of fake errors during every deploy, and root-cause analysis that blames the shutdown rather than the bug. ## What the task has to do A well-behaved task, on observing the signal: 1. Stops as soon as it reaches a point where stopping is safe. 2. Releases everything it acquired, in reverse order of acquisition — close sockets, delete temp files, undo or complete a half-written record. 3. Reports the cancelled outcome to whoever owns it, rather than swallowing it and pretending success. And critically: it does not have to stop *instantly*. Cancellation latency is real and bounded by how often the task looks at the signal. ## What cancellation cannot promise Because it is cooperative, a task that never checks — a tight loop with no checkpoint, a blocking call that ignores the signal, a third-party library that swallows it — is *uncancellable*. Systems therefore pair cancellation with escalation: a grace period, then closing the underlying resource out from under the task (closing a socket makes a blocked read fail), and at the very last resort killing the whole process, where the operating system reclaims everything. Politeness gets you a clean stop; resource and process boundaries get you a guaranteed one. ## Practical shape A useful mental checklist for any long-running task: *Who can cancel me? Where do I look at the signal? What do I release when I stop? Who learns that I stopped?* If any of the four has no answer, the task will one day be the thing keeping a deployment from finishing.

  • If a task is free to never notice the signal, how can a system ever guarantee that shutdown completes?
    It cannot guarantee it through cooperation alone, so cooperation is paired with escalation. You give a bounded grace period, then take away what the task is blocked on — close the socket or file so the blocked operation fails — and as a last resort terminate the process, where the OS reclaims memory, handles and locks. The guarantee comes from resource and process boundaries; cooperation only buys you a *clean* stop rather than an abrupt one.
  • Should a cancelled task be reported as an error?
    Usually no — it should be a distinct outcome. The caller normally requested the cancellation, so it is expected, not a defect: do not page on it, do not blindly retry it, and do not let a cancellation status overwrite the original failure that caused the cancellation. Counting cancellations as a separate metric is still valuable, because a spike in them means callers are timing out.

Cancelling is pulling the fire alarm in a kitchen: cooks stop at a safe point, turn off the burners and walk out. Ignoring the result is just leaving the restaurant — the stoves are still on.

saying these in an interview costs you the question

  • Says you can just kill the thread or task from outside and be done
  • Believes that when the caller stops waiting, the work stops
  • Treats every cancellation as an error to log and retry
  • Assumes cancellation takes effect instantly and synchronously
  • Cancels without any cleanup, leaving sockets, locks or temp files behind

context

open as a page

A background task is started fire-and-forget — nobody keeps its handle or waits for it — and it throws. What typically happens to that error, and why is it dangerous?

level: juniorimportance: must knowfreq 50%

basics

~20 s

With no owner waiting, the error has nowhere to propagate: it ends that task and lands in a default handler, an unread log line, or a result object nobody inspects. The system keeps reporting healthy while the work silently is not being done.

open as a page

A developer starts a background task by spawning it and immediately returning, keeping no handle to it. What can go wrong with such fire-and-forget tasks?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Nobody owns it. Failures vanish silently, nobody waits for it, it can outlive the data and resources it borrowed, it cannot be cancelled or given a deadline, and shutdown can kill it mid-work. The leak stays invisible.

open as a page

What is a cancellation checkpoint, and how would you make a 30-second CPU-bound computation loop cancellable? What determines how quickly cancellation actually takes effect?

level: middleimportance: must knowfreq 52%

basics

~20 s

A checkpoint is a point where the task tests the cancellation signal and abandons the work if it is set. Blocking and awaiting operations usually provide them implicitly; a pure compute loop has none, so you poll the signal every N iterations. Cancellation latency is roughly the time between checkpoints plus cleanup time.

open as a page

Why do modern concurrency runtimes make cancellation cooperative — the task must observe a signal — instead of forcibly terminating it from outside, and where does a thread-level interruption flag fit into that model?

level: middleimportance: must knowfreq 60%

basics

~20 s

Forced termination stops a task at an arbitrary instruction, so invariants are half-updated, locks and buffers are left in unknown states, and no cleanup runs. Cooperative cancellation stops only at safe points. A thread interruption flag is the same idea at thread level: a bit that blocking operations honour by aborting early.

open as a page

You start three concurrent child tasks and wait for all of them. One fails after 50 ms while the others would run for another 10 seconds. What should happen to the siblings, and why is that the default in structured concurrency?

level: middleimportance: must knowfreq 58%

basics

~20 s

By default the scope fails fast: the first failure determines the overall result, so the siblings are cancelled immediately, the scope waits for them to finish unwinding, and the failure is reported to the caller after about 50 ms instead of 10 seconds.

open as a page

In structured concurrency, a scope (also called a nursery or task group) cannot exit until every task started inside it has finished. What guarantees does that rule buy, and what does it cost?

level: middleimportance: must knowfreq 52%

basics

~20 s

It makes concurrency nest like a block of code: on exit no task from that block is still running. So lifetimes are bounded, errors have somewhere to propagate, and cancellation has a subtree. Cost: the slowest child gates the block, so you must add timeouts and cancellation.

open as a page

When a timeout fires and the caller receives a timeout error, what has actually happened to the work that was in flight, and what does the system still have to do?

level: middleimportance: must knowfreq 48%

basics

~20 s

Usually nothing has stopped. A timeout only means the caller stopped waiting; it must then cancel the work region, wait for it to unwind, and release its resources. And a timed-out write is ambiguous — it may still have been applied on the other side.

open as a page

Compare giving each call in a request chain its own fixed timeout with carrying a single absolute deadline through the whole chain. What breaks with per-call timeouts, and what does a propagated deadline give you?

level: middleimportance: must knowfreq 52%

basics

~20 s

Per-call timeouts add up, so the total is bounded only by the sum, not by what the user will wait. A deadline is one point in time passed down the chain: every hop computes the time remaining, so the whole request is bounded and any hop can fail fast when the time is already gone.

open as a page

A task that is cancelled still has to release what it acquired — close connections, delete temp files, undo a half-written record. How do you run that cleanup reliably when the task is already being cancelled and the cleanup itself may need to block or await?

level: seniorimportance: must knowfreq 48%

basics

~20 s

Attach cleanup to scope exit so it runs on success, failure and cancellation alike, release in reverse acquisition order, and run any blocking or awaiting cleanup inside a small shielded (non-cancellable) region with its own timeout — otherwise a sticky cancellation aborts the cleanup itself. Never shield the whole task.

open as a page

Two of five concurrent children fail within milliseconds of each other, and cancelling the remaining three produces cancellation errors of its own. How do you report this to the caller without losing information or blaming the wrong cause?

level: seniorimportance: must knowfreq 44%

basics

~20 s

Pick the first genuine failure as the primary cause, attach the second as a suppressed or secondary cause, and drop the cancellation errors entirely — they are consequences of the first failure, not causes. Keep each error's identity (which child) and never let a wrapper hide the type callers catch on.

open as a page

A service retries a failing dependency up to three times with a 2-second timeout per attempt, while its own caller expects an answer within 3 seconds. What is wrong with that configuration, and how would you fix it?

level: seniorimportance: must knowfreq 46%

basics

~20 s

Attempts multiply: worst case is over 6 seconds against a 3-second promise, so attempts 2 and 3 run after the caller has already given up — pure wasted load. Fix: retry inside the remaining deadline, set each attempt's timeout to min(cap, remaining), and skip retries when too little time is left.

open as a page

Explain how a cancellation signal travels through a tree of parent and child tasks. When is a cancellation considered complete, and what should a parent do while its subtree is winding down?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Cancelling a node marks it and recursively signals every descendant, and tasks spawned afterwards start already-cancelled. Each leaf stops at its next checkpoint and cleans up. Cancellation is complete only when the whole subtree has terminated, so the parent must wait — with a grace deadline and escalation for stragglers.

open as a page

Sometimes you want the opposite of fail-fast: query five replicas and return the first good answer, or gather whatever four of six services returned before a deadline. How do you design a concurrent scope that deliberately collects partial results, and what guarantees must still hold?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Separate the completion policy from the structure: make each child's failure a value rather than an escaping error, so the scope decides. Common policies are all-succeed, any-succeed, k-of-n quorum and best-effort-by-deadline. The scope still owns and joins every child, and partial results must be labelled as partial.

open as a page

Why is a concurrent task that outlives the code region that started it dangerous when that region owns resources such as a pooled connection, a reusable buffer, a request-scoped context, or an open transaction?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Because release is tied to the region, not the task. When the region ends it closes or recycles the resource, and the surviving task keeps using it — writing into a buffer or connection now owned by someone else. That is corruption and cross-request data leakage, not just a crash.

open as a page

A long-lived component supervises a set of worker tasks that can crash. Walk through the strategies available when a child dies — resume, restart, stop, escalate — and explain how you decide among them.

level: principalimportance: should knowfreq 38%

basics

~20 s

A supervisor watches its children's outcomes and applies a policy: resume (keep state, log), restart (discard possibly corrupt state, start from a known-good initial state), stop (the child is no longer viable), or escalate (fail itself so a higher supervisor recovers a bigger subtree). Bound restarts with a budget and backoff.

open as a page

Some concurrent work legitimately outlives any single request — background schedulers, cache refreshers, connection keep-alive loops, metric flushers. How do you reconcile that with a discipline that forbids orphan tasks?

level: principalimportance: should knowfreq 28%

basics

~20 s

Give them a longer-lived owner instead of no owner. Each background task belongs to an explicit component or application-lifetime scope that starts it, watches its failures, and cancels and joins it in a defined shutdown order. Ownership moves up the tree; it never disappears.

open as a page

You own a user-facing API with a 500 ms latency target that fans out to several downstream services, some of which call further services. How would you assign timeout and deadline budgets across that call graph?

level: principalimportance: should knowfreq 30%

basics

~20 s

Start from the user-visible budget and divide it downward: subtract fixed overhead and a cleanup reserve, split the rest across sequential hops, share it among parallel ones, and propagate the remainder as a deadline every hop enforces and sheds on. Size each hop from its measured p99, not from a guess.

open as a page

Describe request hedging — sending a second copy of an in-flight request before the first one has failed. When does it improve tail latency, and what does it cost?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

If a request is still pending at roughly the p95 latency, send a duplicate to another replica, take whichever answers first, and cancel the loser. It cuts tail latency when slowness is per-request bad luck, and costs about 5% extra load plus a strict idempotency requirement.

open as a page