skip to content

When a timeout fires and the caller receives a timeout error, what has actually happened to the work that was in flight, and what does the system still have to do?

level: middleimportance: must knowfreq 48%

answer

  1. timeout = caller stops waiting, worker doesn't stop
  2. cancel the whole region, not one call
  3. request stop, then await stopped (grace budget)
  4. blocking I/O stops only when the handle is closed
  5. timed-out write = unknown, not failed

basics

~20 s

Usually nothing has stopped. A timeout only means the caller stopped waiting; it must then cancel the work region, wait for it to unwind, and release its resources. And a timed-out write is ambiguous — it may still have been applied on the other side.

solid answer

~60 s

A timeout is **the caller giving up**, not the work stopping. Three obligations follow. 1. **Turn it into cancellation of a region.** The expiry must cancel the whole task subtree for that operation, not just the one call that noticed. Otherwise siblings keep running for a result nobody will read. 2. **Actually wait for the unwind.** Cancellation is cooperative: tasks stop at cancellation points, and I/O usually stops only when the socket is closed or aborted. Returning the timeout error before the children have unwound leaks the very tasks and resources you were trying to reclaim — which is why a scope joins them, with its own small grace budget, since the original budget is already spent. 3. **Treat the outcome as ambiguous.** For a remote mutation, timeout means "unknown", not "did not happen": the request may have been received, applied, and only the response lost. Blind retry of a non-idempotent operation double-applies it. Use idempotency keys or a status check. Cleanup must be bounded and idempotent, and its own failure must not mask the timeout.

go deeper

for a junior

Say the key thing plainly: the timeout stops the waiting, not the work; something still has to cancel it and clean up.

for a middle

Add that cancellation is cooperative and must be awaited, and that a timed-out write may still have been applied.

for a senior

Cover the grace budget and escalation, blocking I/O needing the handle closed, idempotency keys for ambiguous mutations, and metrics for tasks that never stopped.

for a principal

Frame it as a system property: timeouts must shed load rather than add it, ambiguity must be designed for with idempotency and reconciliation, and unbounded calls in cleanup paths should be prevented by the platform.

## What a timeout actually is A timeout is a decision by the *waiter*: "I will not wait any longer." It says nothing about the *worker*. In the simplest broken implementation, the caller abandons a future and returns an error while the underlying task keeps running, keeps its connection, keeps its memory, and eventually completes into the void. The system now has more concurrent work than it thinks it has — the classic timeout-induced overload, where timeouts under stress *increase* load instead of shedding it. ## Obligation 1: expiry must cancel a region, not one call An operation is usually a tree: fetch three things in parallel, one of which fans out further. If the deadline cancels only the call that noticed, siblings continue. So the deadline should be attached to a **scope** covering the whole operation, and expiry cancels that scope, transitively. This is why timeouts are described as *scoped cancellation*: the scope names what to stop, the deadline says when. ```text with scope(deadline = now + 200ms): scope.start(fetchA) scope.start(fetchB) # at expiry: cancel A and B, then wait for both to unwind, then raise Timeout ``` ## Obligation 2: cancellation is cooperative and must be awaited A running task does not stop at an arbitrary instruction. It observes cancellation at defined points: when it checks a flag, when it awaits something cancellable, or when a blocking operation is interrupted. Pure CPU loops with no check will run to completion regardless. Blocking I/O typically stops only if the runtime closes or aborts the underlying socket or file handle — which is why a socket read that ignores the deadline is the single most common source of "the timeout fired but the thread is still stuck". So timeout handling has two phases: *request stop*, then *wait for stopped*. Skipping the second phase means returning while children still hold the connection, the buffer, and the request context — the leak the timeout was supposed to prevent. Because the deadline has already expired at this point, cleanup needs its **own small grace budget** (a cleanup deadline), plus an escalation path: request cooperative stop, then hard-abort the resource (close the socket), then abandon and log with enough identity to be diagnosed. A task that never stops must be *visible*, not silently awaited forever. ## Obligation 3: the result is ambiguous, not negative For anything with a side effect on another system, a timeout has three possible truths: the request never arrived; it arrived and failed; it arrived and *succeeded* but the response was lost or slow. The caller cannot distinguish them. Treating timeout as failure and retrying a non-idempotent operation is how duplicate charges, duplicate orders and double-applied increments happen. Remedies, in order of preference: make the operation idempotent so retries are safe; attach a client-generated idempotency key so the server can deduplicate; or, if neither is possible, query the operation's status before retrying and design the read path to expose it. Note that cancelling locally does *not* undo remote work — at best cancellation propagates as a signal the peer may honour, and even then it races with completion. ## Cleanup rules Cleanup runs in an already-degraded state, so it needs care: - **Bounded.** Never issue an unbounded call during cleanup — a compensating request with no timeout of its own reproduces the original hang inside the error path. - **Idempotent.** Cleanup can run twice: once on timeout and once when the abandoned task finally completes. - **Non-masking.** An error thrown by cleanup must not replace the timeout error the caller needs to see; attach it as suppressed or log it separately. - **Ordered.** Release in reverse acquisition order, and make sure nothing releases a resource a still-unwinding child is using. ## What good looks like operationally Count timeouts separately from other errors, and separately count *abandoned tasks that never stopped* — that number should be zero and is a strong leak indicator. Record how long unwinding took; a growing cleanup time is an early signal that some path ignores cancellation. And whenever a timed-out operation had a side effect, log the idempotency key so the ambiguous case can be reconciled later.

  • After a timeout fires, why not just return the error immediately and let the abandoned task finish on its own?
    Because it is no longer accounted for anywhere. It still holds a thread or connection, still writes to buffers and contexts that the caller is about to release, and its failure has nowhere to go. Under load, abandoning tasks means concurrency keeps climbing while the caller believes it shed the work, which turns a slow dependency into a capacity collapse. Waiting for the unwind, with a bounded grace budget and a log for anything that refuses to stop, keeps the accounting honest.
  • How should a client treat a timeout on a request that creates a resource, such as a payment or an order?
    As an unknown outcome. The request may have been applied with only the response lost, so retrying blindly risks a duplicate. The right design is an idempotency key generated by the client and reused across retries, so the server recognizes the repeat and returns the original result. Where that is impossible, query the operation's status by a client-supplied identifier before retrying, and make the reconciliation path a first-class part of the design rather than a manual cleanup.

Hanging up on a phone call does not stop the person on the other end from doing what you asked. You only know that you stopped listening — and if you call back to repeat the request, they may do it twice.

saying these in an interview costs you the question

  • "The timeout killed the request" — it only ended the wait; the work usually continues.
  • Retrying a non-idempotent write after a timeout as if the timeout proved it did not happen.
  • Returning the timeout error without waiting for the cancelled children to unwind and release resources.
  • Assuming a blocking socket read will stop when a timer fires, without closing or aborting the handle.
  • Performing cleanup with an unbounded call, so the error path hangs exactly like the original operation.

context