How should a retry loop wait between attempts so a cancelled context.Context ends it immediately?
answer
- sleeping ignores the shutdown signal
- wait on two things at once
- a closed channel is always ready
- select over Done and the timer channel
- return ctx.Err so the loop stops
basics
~10 sNever time.Sleep. Start a time.Timer for the delay and select over ctx.Done() and the timer's channel, returning ctx.Err() when cancellation wins. Stop the timer on the way out so it is released early.
solid answer
~40 s`time.Sleep` is uninterruptible: a worker parked in a thirty-second wait ignores a shutdown signal for thirty seconds, and that is how a retry loop outlives the context that was supposed to stop it. Instead create a `time.Timer` and `select` over `ctx.Done()` and `timer.C`, returning `ctx.Err()` when the context fires so the caller stops attempting; `defer timer.Stop()` releases it as soon as the wait ends either way. Check the context before each attempt too, and treat `errors.Is(err, context.Canceled)` from an attempt as terminal — a cancelled context will not un-cancel, so another attempt can only fail identically. The same shape covers deadlines: the wait ends early, the loop returns, and the shutdown path is not held hostage by whatever backoff delay happened to be in flight.
code
go · 10 linesfunc wait(ctx context.Context, d time.Duration) error {
t := time.NewTimer(d)
defer t.Stop()
select {
case <-ctx.Done():
return ctx.Err()
case <-t.C:
return nil
}
}go deeper
Remember that time.Sleep cannot be interrupted, and that waiting on ctx.Done together with a timer channel in a select is the standard way to make a delay cancellable.
Explain the mechanics: a closed Done channel is immediately ready, ctx.Err distinguishes cancellation from a deadline, and defer Stop releases the timer when the other case wins.
Show where the three context checks go around a real delivery loop, and why a cancelled context is terminal rather than retryable — including using errors.Is on a wrapped error.
Own the shutdown contract for the whole worker: how long a drain is allowed to take, and the fact that any uninterruptible wait inside it becomes the floor on your deployment's stop time.
## The failure this prevents A webhook delivery worker drains a queue, posts each payload, and on failure waits before trying again. Shutdown arrives: the process cancels the worker's `context.Context` and waits for it to return. Nothing happens for half a minute, and then the orchestrator kills the process mid-delivery. The cause is one line: ```go time.Sleep(delay) ``` `time.Sleep` parks the goroutine for a fixed duration and takes no arguments that could ever end it early. It does not know about contexts, signals, or anything else. Every second of backoff is a second the worker cannot be stopped, and the worst case scales with the longest delay in the schedule. ## The cancellable wait ```go func wait(ctx context.Context, d time.Duration) error { t := time.NewTimer(d) defer t.Stop() select { case <-ctx.Done(): return ctx.Err() case <-t.C: return nil } } ``` A `select` blocks until one of its cases is ready, so this waits for whichever comes first: the delay elapsing, or the context being cancelled or hitting its deadline. `ctx.Done()` returns a channel that is closed on cancellation, and a closed channel is always ready to receive, which is what makes cancellation win instantly. Returning `ctx.Err()` propagates *why* it ended — `context.Canceled` for an explicit cancel, `context.DeadlineExceeded` for a deadline — so the caller can distinguish a shutdown from a genuine timeout without inspecting anything else. The caller then simply refuses to continue: ```go if err := wait(ctx, delay); err != nil { return err // shutdown or deadline: stop retrying } ``` ## Why Stop still matters `defer t.Stop()` releases the timer as soon as the wait finishes. Since Go 1.23 an unstopped timer is collectable once it becomes unreachable, and timer channels became unbuffered, so a forgotten timer no longer pins memory until it fires the way `time.After` used to inside a hot `select`. Stopping is still the clearest way to express "I am done with this", it costs nothing, and it keeps the code correct on older toolchains a service might still be built with. ## Check the context in three places A loop that is genuinely cancellable checks in three places, and each catches a different window: 1. **Before an attempt.** If the context is already done, do not open a connection you are about to abandon. `if err := ctx.Err(); err != nil { return err }`. 2. **During an attempt.** Build the request with `http.NewRequestWithContext(ctx, ...)` so cancellation aborts the in-flight call rather than only being noticed afterwards. 3. **During the wait.** The `select` above. ## Cancellation is terminal, not transient When an attempt returns an error, the loop has to decide whether another attempt could plausibly do better. A cancelled context never becomes uncancelled, so the answer there is always no: ```go if errors.Is(err, context.Canceled) || errors.Is(err, context.DeadlineExceeded) { return err } ``` `errors.Is` is the right test rather than `err == context.Canceled` because the error you get back from an HTTP call is wrapped by the layers it passed through, and comparing with `==` silently misses the wrapped case — which then shows up as a loop that keeps retrying a context that is already dead, burning the shutdown window on attempts that cannot succeed. ## Per-attempt versus overall The wait above is driven by whatever context the worker holds. If each attempt also gets its own shorter deadline derived from that context, remember that the derived context is cancelled by the parent too, so the shutdown path still wins; the parent's cancellation cascades to every child. What you must not do is derive an attempt context from `context.Background()` to "avoid the parent's deadline" — that severs the link and reintroduces exactly the worker that will not stop. ## Testing it The honest test does not sleep for real. Run one delivery against an `httptest.NewServer` handler that always fails, cancel the context from the test goroutine after the first attempt, and assert that the worker's function returns within a few milliseconds with an error satisfying `errors.Is(err, context.Canceled)`. With `time.Sleep` in the loop that test hangs for the full backoff and is the first thing to go flaky in CI; with the `select` it is deterministic.
- Why use errors.Is rather than comparing the error to context.Canceled with ==?Because the error that comes back from an HTTP call has been wrapped on the way up, so the value you hold is rarely the sentinel itself. errors.Is unwraps the chain and matches the sentinel wherever it sits. A == comparison silently returns false on a wrapped error, and the loop keeps retrying a context that is already dead.
- Is time.After in the select an acceptable alternative to time.NewTimer?Functionally yes, and since Go 1.23 the timer it creates is collectable once unreachable rather than pinned until it fires, so the old leak-in-a-loop objection is gone on current toolchains. An explicit time.NewTimer with defer Stop still reads better in a retry loop because the timer has a name and an obvious lifetime.
- Should each attempt get its own deadline derived from the worker's context?Usually yes, so a single hung attempt cannot consume the whole window. Derive it with context.WithTimeout from the worker's context, never from context.Background, and always call the returned CancelFunc. Deriving from the parent keeps the shutdown cascade intact: cancelling the worker cancels the in-flight attempt too.
time.Sleep is an oven timer you cannot cancel: the only way to stop waiting is to wait. The select is a doorbell next to it — whichever rings first ends the wait.
saying these in an interview costs you the question
- Uses time.Sleep between attempts and calls it good enough
- Checks the context only before the first attempt
- Compares an error to context.Canceled with == instead of errors.Is
- Derives the attempt context from context.Background to dodge the parent
- Retries after a cancellation as if it were a transient failure