skip to content

A task that is cancelled still has to release what it acquired — close connections, delete temp files, undo a half-written record. How do you run that cleanup reliably when the task is already being cancelled and the cleanup itself may need to block or await?

level: seniorimportance: must knowfreq 48%

answer

  1. cleanup on all exits, reverse order
  2. sticky cancel kills blocking cleanup
  3. shield = small + own deadline
  4. close beats graceful goodbye
  5. cleanup error must not eat the cause

basics

~20 s

Attach cleanup to scope exit so it runs on success, failure and cancellation alike, release in reverse acquisition order, and run any blocking or awaiting cleanup inside a small shielded (non-cancellable) region with its own timeout — otherwise a sticky cancellation aborts the cleanup itself. Never shield the whole task.

solid answer

~60 s

Two rules. First, cleanup belongs on **scope exit**, not on the success path: a handler that runs for all three outcomes (success, failure, cancelled), releasing in reverse order of acquisition. Anything conditional on "the work finished" is a leak waiting to happen. Second, in runtimes where cancellation is **sticky** — once cancelled, any further blocking or awaiting call in that task fails immediately — naive cleanup that itself blocks (send an abort, delete a remote object, roll back) is aborted on its first line. The fix is a **shielded region**: a small, explicitly non-cancellable block around only the cleanup, with **its own bounded deadline** drawn from the shutdown grace budget. Discipline around it: keep the shield tiny (shielding the whole task body just makes the task uncancellable); prefer closing resources over waiting for graceful protocol exchanges; make cleanup idempotent, since it may run after a partial one; and never let a failure inside cleanup replace the original cause — attach it as a secondary/suppressed error. The same shield is what protects a genuinely atomic critical section, with checkpoints placed on either side of it rather than inside.

code

text · 12 lines
text
conn = pool.acquire()
try:
    tmp = create_temp()
    try:
        do_work(conn, tmp)          # cancellable: has checkpoints
    finally:                        # runs on success, failure AND cancel
        shielded(timeout=2s):       # cancellation masked, but bounded
            best_effort(delete(tmp))
finally:
    shielded(timeout=1s):
        best_effort(conn.abort())   # graceful attempt
    conn.close_hard()               # always: local, cannot hang

go deeper

for a junior

Say cleanup must run on every exit including cancellation, release in reverse order, and note that a cancelled task still owns whatever it acquired.

for a middle

Explain sticky cancellation aborting blocking cleanup and the need for a small non-cancellable region, plus idempotency and not losing the original error.

for a senior

Budget it: cleanup gets an explicit slice of the shutdown grace period, graceful attempts are timeboxed before hard-closing, and cancellation is tested at several points with resource counters asserted back to baseline.

for a principal

Set the policy — where shields are allowed, the maximum shielded span, the shutdown budget split across the tree — and push toward designs where compensation is idempotent and recovery is possible after an abrupt stop, so cleanup never becomes load-bearing for correctness.

## The problem Cancellation unwinds a task mid-flight. At that moment the task may hold: a pooled connection, an open file and a half-written temp file, a lease or distributed lock, a partially applied remote change, an entry in an in-memory registry. If unwinding does not release these, cancellation converts one abandoned request into a permanent leak — and cancellations arrive in bursts (client disconnects, deploys, timeouts), so leaks arrive in bursts too. ## Rule 1: cleanup on scope exit, in reverse order Every acquisition should be paired at the point of acquisition with a release that runs on **all** exits. Whatever the language calls it — finally block, defer, scope guard, RAII destructor, using/with block — the property required is: *runs on success, on failure, and on cancellation*. Order matters: release in reverse order of acquisition, because later resources often depend on earlier ones (you close the stream before returning the connection that owns it). Nesting the guards naturally produces this order. A useful discipline is to make the *owner* obvious: whoever acquires releases, and a resource handed to a child task becomes the child's responsibility only if the child's lifetime is guaranteed to end inside the parent's scope. ## Rule 2: cancellation is often sticky, so cleanup needs a shield Many runtimes make cancellation **sticky**: once the task is cancelled, every subsequent operation that would block or suspend immediately fails with the cancellation outcome. This is deliberate — it stops a cancelled task from starting new work and drags it toward termination quickly. But cleanup frequently *is* blocking work: sending an abort/rollback to a server, deleting a remote temp object, flushing a buffer, releasing a distributed lease. Without protection, the first such call in the cleanup path fails instantly and the release never happens. The mechanism to protect it is variously called a **shield**, an uncancellable block, a non-cancellable context, or a cleanup budget: inside it, the cancellation signal is masked so blocking calls behave normally. Using a shield safely: - **Keep it narrow.** Wrap only the release calls. The moment you shield the main work "to keep it simple", the task becomes uncancellable and you have traded a leak for a hang. - **Bound it.** A shielded region has, by construction, opted out of the cancellation that would have stopped it — so it needs its own deadline. Give cleanup an explicit timeout carved from the shutdown grace period ("cleanup gets 2 seconds of a 10-second budget"). When that deadline is exceeded, escalate: abandon the graceful path, hard-close the handle, log loudly. - **Prefer closing to negotiating.** Hard-closing a socket or handle is fast, local, and cannot block indefinitely; a polite protocol-level goodbye can hang exactly when the peer is the thing that is broken. Try the graceful step under a short timeout, then close. - **Make it idempotent.** Cleanup can run after a previous partial cleanup, or twice across a retry. Deleting an already-deleted temp file or releasing an already-released lease must be a no-op, not an error. ## Errors inside cleanup Cleanup itself can fail. The invariant is: **the original cause must survive**. If the task was unwinding because of a failure, and the cleanup then throws, replacing the exception makes the root cause disappear and every investigation starts from the wrong stack. Attach the cleanup error as a suppressed/secondary cause and keep the first one primary. If the task was unwinding because of cancellation, a cleanup failure is usually logged rather than propagated, since the caller already stopped caring about the result — but it must still be *visible* in metrics, because "cleanup failing" is exactly how leaks start. ## Shielding critical sections (not just cleanup) The same tool applies to sections that must not be abandoned half-done: a two-phase commit acknowledgement, an idempotency record write that must accompany a side effect, a message acknowledgement after processing. Wrap the atomic span in a shield and place checkpoints *before* and *after* it, never inside. Then the task is highly cancellable overall while still having a few short spans that will always complete. Keep those spans short precisely because they set the floor on cancellation latency. ## Anti-patterns - **Cleanup on the success path only** — "we close it at the end of the happy path". - **A shield around everything**, producing tasks that ignore cancellation entirely. - **Unbounded cleanup** — a shielded region with no timeout can hang shutdown forever, which is worse than the leak. - **Starting new work in cleanup**: spawning children, retrying business calls, or logging over the network to a system that is currently down. - **Swallowing the original error** when the cleanup throws. - **Non-idempotent compensation** that errors when it finds nothing to undo, masking the real outcome. ## How you verify it Test cancellation at multiple points of the task, not just at the start: cancel before acquisition, after acquisition but before the work, and mid-work, and assert in each case that the resource counters return to their baseline (open connections, file handles, temp files) and that the task terminates within the cleanup budget.

  • What goes wrong if you shield the entire task body instead of just the cleanup?
    The task becomes uncancellable: the signal is masked for its whole duration, so timeouts and shutdown have no effect and the parent scope waits indefinitely for it. You have converted a resource leak into a hang, which is usually worse because it blocks deploys and pins the whole subtree. Shields must cover only spans that would do harm if abandoned, and those spans must be short.
  • Cleanup itself throws while the task is unwinding from a real failure. What do you report?
    The original failure stays primary and the cleanup error is attached as a secondary or suppressed cause. Replacing the cause makes the true root disappear from logs and sends every investigation to the wrong stack. If the unwinding was due to cancellation rather than failure, the cleanup error is typically logged and counted rather than propagated, since the caller has already abandoned the result — but it must remain visible, because failing cleanup is how leaks begin.

A shield is the surgeon finishing the stitch after the fire alarm sounds: brief, bounded, and only for the step that would do harm if abandoned — not an excuse to complete the whole operation.

saying these in an interview costs you the question

  • Releases resources only on the success path
  • Wraps the whole task in a non-cancellable block to 'make cleanup reliable'
  • Shielded cleanup with no timeout, so shutdown can hang forever
  • Lets an exception thrown in cleanup replace the original failure
  • Cleanup that is not idempotent and errors when there is nothing left to undo
  • Starts new blocking work — retries, network logging, spawning children — during cleanup

context