skip to content

A client abandons its request while three asynchronous calls it triggered are still pending. What does cancelling a future actually mean, and what has to be true for the pending work to genuinely stop?

level: seniorimportance: should knowfreq 42%

answer

  1. Cancel the reader vs stop the producer
  2. Cooperative: poll a signal, interrupt, close the resource
  3. Race losers keep running
  4. Timeout without cancel amplifies load
  5. Cancel/complete race; cancel must be idempotent

basics

~20 s

Cancelling usually only completes the read side with a cancelled outcome so the consumer stops waiting. The producer keeps working unless it cooperates — polling a cancellation signal, being interrupted, or having its socket closed. Real cancellation needs a channel back to the producer plus resource cleanup.

solid answer

~60 s

Separate two things that the word cancel conflates. 1. **Detaching the reader.** The future completes as cancelled and downstream steps stop. Cheap and always possible. 2. **Stopping the producer.** Only possible if the producer cooperates: it periodically checks a cancellation signal, it is waiting on an interruptible operation, or something closes the underlying resource such as the socket. A bare future has no channel back to whoever holds the write side, so cancel often means only (1). The work keeps burning CPU, connections and rate limit, and its eventual completion is discarded — this is how abandoned requests turn into an overloaded downstream. The other hard parts: - **Propagation.** Cancelling a composed result should cancel its inputs; many libraries do not do this upstream, so a raced or timed-out call leaves siblings running. - **Races.** Cancel and completion can arrive together. Single assignment picks a winner; cancellation must be idempotent, and a value that arrives after cancellation may hold a resource that now nobody will close. - **Timeouts are cancellation with a clock** and should trigger the same propagation.

code

text · 13 lines
text
function longJob(signal):
    p = new Promise()
    spawn {
        for chunk in chunks:
            if signal.cancelled: p.fail(Cancelled); return
            process(chunk)
        p.complete(result)
    }
    return p.future()

f = longJob(signal)
onTimeout: signal.cancel()          # ask the producer to stop
f.onComplete(v -> if cancelled: dispose(v))   # late value still owns a resource

go deeper

for a junior

Know that cancelling mainly means the consumer stops waiting, and that stopping real work requires the work itself to check for cancellation.

for a middle

Name the three cooperation mechanisms — polled signal, interruptible wait, closing the underlying resource — and note that cancel is asynchronous.

for a senior

Discuss propagation to siblings and upstream, the completion race and disposal of late values, and why a timeout must cancel to avoid amplifying downstream load.

for a principal

Argue for scope-owned lifetimes so cancellation is structural rather than remembered case by case, and describe the shutdown and overload behaviour you want as a system property.

## Cancel means two different things Ask a candidate to cancel a future and you will learn whether they distinguish the **read side** from the **write side**. - Completing the *read* side with a cancelled outcome makes the consumer stop waiting: downstream steps are skipped, the caller can respond, the memory of the chain can be released. This is always available and costs nothing. - Stopping the *work* requires reaching the producer, and a future by itself contains no such channel. Whoever holds the write side is somewhere else entirely — a thread in a pool, an outstanding I/O operation, a remote server that is already computing your answer. Because of that asymmetry, in many systems cancel means only *I stopped listening*. The work continues to completion and its result is thrown away. ## What cooperation looks like There are exactly three practical mechanisms: 1. **A cancellation signal the work polls.** A token, flag or context object shared with the producer, checked between units of work. This is the general solution and works for CPU-bound loops and multi-stage jobs alike. It only helps at the granularity of the checks: a task that checks once every ten seconds cancels once every ten seconds. 2. **An interruptible wait.** If the producer is parked in an operation the runtime can interrupt, the runtime delivers the interruption and the operation aborts. This depends on the operation being interruptible; many are not. 3. **Destroying the resource underneath.** Closing the socket or file handle makes the pending operation fail. This is blunt and effective, and it is how request cancellation reaches a remote server — the transport signals abort, or the connection drops and the server (if it honours that) stops. Without one of these, cancellation is advisory. Interviewers listen for the phrase *cooperative cancellation*: the runtime asks, the work agrees. ## Why it matters operationally Abandoned work is not free. Under load, clients time out and retry; if abandoned calls keep running, the downstream service sees the original load plus the retries, and the system degrades exactly when cancellation would have helped most. This is a classic overload amplifier: a timeout without cancellation converts client impatience into extra server work. Abandoned work also holds resources: a connection from a bounded pool, a permit, a lock, memory for buffered results. That is a leak on the failure path, which is where leaks hurt. ## Propagation, in both directions - **Downstream:** naturally handled. If a future is cancelled, its dependent steps never run. - **Upstream:** the harder half. When you cancel a composite (a fan-in, a race, a timeout wrapper), do the *inputs* get cancelled? Frequently not, unless the library or your code arranges it. Whoever wins a race leaves the losers alive; a fail-fast fan-in leaves the siblings alive. If you use racing for hedged requests or timeouts, you must cancel the losers explicitly or you have built a work amplifier. The general solution to propagation is a **scope that owns the child operations**, so that cancelling or failing the scope cancels everything started inside it — the structured approach to concurrency, which exists precisely because ad hoc future graphs make this so easy to get wrong. ## Races and idempotency Cancellation and completion can arrive simultaneously. Under single assignment one wins: - If completion won, the cancel is a **no-op** — so `cancel()` must return or report whether it actually took effect, and callers must not assume it did. - If cancellation won, a value may still arrive from the producer afterwards. If that value owns a resource — an open stream, a leased object — nobody is going to close it unless you attach a disposal step for exactly this case. - Cancel must be **idempotent**: cancelling twice, or cancelling an already-completed future, must be harmless. Also beware of the intermediate state: work that has been asked to stop but has not stopped yet. Shutdown logic that assumes cancel is synchronous will race live work; if you need *stopped*, you must wait for the work to actually report termination, not merely for the cancel call to return. ## Design checklist - Give every long-running producer a cancellation signal and check it at meaningful boundaries. - Bind cancellation to the resource: closing the connection is the enforcement of last resort. - Make timeouts cancel, not merely fail. - Cancel the losers of any race or hedge. - Dispose values that arrive after cancellation. - Distinguish *requested* from *finished* in shutdown paths. ## In an interview Lead with the two meanings, name cooperative cancellation as the only real mechanism, and then give the operational sting: a timeout that does not cancel amplifies load on the downstream service that is already struggling.

  • Your timeout fires and you return an error to the client. Is that enough?
    No. Failing the composite only detaches the reader; the underlying call is still in flight, still holding a connection and still loading the downstream service. Under load that turns client timeouts and retries into amplified traffic against a service that is already struggling. A timeout should also trigger cancellation of the work and release of the resources it holds.
  • What guarantees does a cancel call give you about the state of the work when it returns?
    Almost none by itself. It records a request to stop; the work may have already completed, may stop at its next check, or may be stuck in a non-interruptible operation. Treat cancel as asynchronous, make it idempotent, and if you need certainty that the work has ceased, wait for a separate termination signal from the task rather than for the cancel call to return.

Hanging up on a chef does not stop the meal. Unless the kitchen has a way to hear that the order was withdrawn, they cook it anyway and it goes in the bin — using the same oven someone else is queuing for.

saying these in an interview costs you the question

  • Assuming cancelling a future stops the underlying computation with no cooperation from the producer
  • Treating cancel as synchronous — believing the work has stopped once the call returns
  • Forgetting that the losers of a race or hedged request keep running and holding resources
  • Ignoring the cancel-versus-complete race, so a late value holding a resource is never disposed
  • Implementing timeouts as failure only, which amplifies load on an already-overloaded downstream

context