In a transcoding job runner, context.WithCancel is called per job but cancel runs only on the error path. What leaks, and how do you find every such site?
answer
- the happy path is the buggy path
- the slope tracks completed work, not concurrent work
- reachability, not scope, decides
- the parent's children set holds the reference
- go vet has an analyzer named for this exact bug
basics
~20 sEach successful job leaves its derived context registered in the long-lived parent's children set, so the parent retains every child. Memory creeps in step with throughput. Sweep the module with go vet, whose lostcancel check flags exactly this.
solid answer
~50 sOn the success path the child context is never detached, so the runner's long-lived parent context keeps it in its children set forever — one retained node per completed job, plus anything reachable from each child, including any values stored on it. The symptom is a slow, throughput-proportional memory creep with no unbounded queue and no goroutine explosion, which is why it survives casual review. Any goroutine parked on that child's `Done()` also never wakes. The fix is the two-line idiom: `defer cancel()` on the line after the derivation, so it fires on the success path too; cancelling after the job finished undoes nothing. To find every site, run `go vet ./...` — the `lostcancel` analyzer reports a discarded cancel value and a return path reachable without the cancel variable being used. Then hand-audit what vet cannot see: cancel funcs stored on a struct or passed to another function.
code
go · 11 linesfunc runJob(root context.Context, chunks [][]byte) error {
ctx, cancel := context.WithCancel(root)
for _, c := range chunks {
if err := encodeChunk(ctx, c); err != nil {
cancel() // only here
return err
}
}
return nil // child stays registered with root, forever
}go deeper
You are not expected to diagnose this, but recall the rule it enforces: derive with context.WithCancel and defer cancel on the very next line, on every path including success.
Explain why the child survives: the parent's children set holds a reference, so reachability rather than lexical scope decides. Note that the growth is proportional to completed jobs.
Show the audit as an engineer inheriting the codebase would run it: go vet across the module for lostcancel, then a hand pass over the sites vet cannot reason about, and a controlled comparison confirming the slope really was this.
Own the durable fix rather than the incident: vet in CI, the two-line idiom as a stated convention, and a clear rule about who holds a cancel func when a context outlives the function that made it.
## The mechanism behind the creep `context.WithCancel(parent)` does not just hand back a child; it **links** the child into the tree so that cancelling an ancestor can reach it. Concretely, the package walks up to the nearest cancellable ancestor and adds the new child to that ancestor's set of children. That is a strong reference held by the ancestor. A job runner typically has a single long-lived root — one context created at startup and used for the process's lifetime. Every job derives from it. Every job that finishes without calling its cancel func leaves its node in the root's children set, and the node in turn keeps alive whatever the child chain references: further derived contexts, any values attached along the way, and the closure state of anything registered against it. So the retained set grows monotonically at exactly the rate the runner completes work. A machine running a hundred jobs an hour looks fine in a smoke test and fine in a one-hour soak; a week in staging is where you see it. That is the profile of the failure: **not a spike, a slope**. ## Why it hides Three properties make this defect unusually easy to miss. 1. **The happy path is the buggy path.** Everything the tests exercise as a failure — a bad input, a timeout, a downstream error — takes the branch that *does* call cancel. Only success leaks, and success is what you stop looking at. 2. **Each leaked node is small.** A `cancelCtx` is a handful of words. The visible growth comes from what hangs off it, which varies by job, so the slope is noisy and easy to attribute to something else. 3. **Nothing fails.** There is no error, no log line, no panic. The service degrades over days and gets restarted by a deployment before anyone charts it, which is how these survive for months. A second, related consequence: anything that registered against the child's `Done()` and is waiting for it — a helper goroutine, an `AfterFunc` registration — never fires, because that context is now unreachable in the sense that matters (nothing will ever cancel it) even though it is very much reachable by the garbage collector. ## The fix The fix is not a smarter cleanup path; it is the idiom: ``` ctx, cancel := context.WithCancel(parent) defer cancel() ``` adjacent lines, always. Two objections come up in review and both are wrong: - *'We already returned the result, cancelling now would be harmful.'* It would not. Cancel closes a channel and sets an error on a context nobody will use again. The results are already produced and returned. Cancelling a finished unit of work is a no-op with a bookkeeping side effect — the detach — which is the whole point. - *'The child is short-lived, the GC will handle it.'* The GC handles unreachable objects. The parent's children set makes the child reachable. Reachability, not scope, is what decides. ## Finding every site `go vet` ships an analyzer called `lostcancel`, on by default, aimed at precisely this bug. Running `go vet ./...` over the module reports two shapes: - a cancel value that is thrown away outright, and - a function where a `return` is reachable without the cancel variable having been used on that path. The second is the one that catches this job runner: the error path calls cancel, the success path returns without it, and vet walks the control-flow graph and points at the offending `return`. **Know its blind spots**, because an audit that trusts vet alone will declare victory early. The analyzer reasons about a local variable within a function. It cannot follow a cancel func that is: - stored in a struct field and expected to be invoked by some other method, - passed as an argument to another function that may or may not call it, - captured by a closure handed to something else, - or called only inside a goroutine that might exit before reaching the call. Each of those is a hand-audit item. The grep that finds candidates is the derivation call itself: every `context.WithCancel(` in the tree, checked for a `defer cancel()` on the following line, with the remainder read individually. On a codebase you inherited, that grep is a bounded, finishable piece of work, and it converts a vague memory-creep investigation into a list. Once the sweep is clean, the way it stays clean is CI: `go vet ./...` in the pipeline, so the next instance is caught in review rather than in staging six weeks later. ## Confirming it was the cause Before rewriting anything, tie the slope to the mechanism. The cheap confirmation is a controlled comparison: run the same workload against a build with `defer cancel()` added and one without, over enough jobs for the slope to be visible, and check whether the growth is proportional to completed jobs rather than to concurrent ones. Proportionality to *completed* work is the fingerprint here — a normal live-set growth tracks concurrency, and a retention bug like this tracks throughput.
- Reviewers object that cancelling after the job already succeeded could abort something. How do you answer?Cancel closes a channel and sets an error on a context nothing will use again; the results are already produced and returned. The only lasting effect is the detach from the parent, which is exactly what you want. That is why `defer cancel()` is safe on the success path.
- What can go vet's lostcancel check not see?It reasons about a local variable inside one function. A cancel func stored on a struct field, passed to another function, captured by a closure handed elsewhere, or invoked only inside a goroutine that may exit first is all invisible to it. Those need a hand audit driven by grepping the derivation sites.
- How would you distinguish this retention from ordinary growth in the live set?Check what the growth is proportional to. Normal live-set size tracks concurrent work and flattens when concurrency does. This retention tracks completed jobs, so it keeps climbing at steady concurrency and never comes back down between bursts.
- How do you stop the same defect recurring after you have fixed every site?Put `go vet ./...` in CI so the analyzer runs on every change, and keep the two-line idiom — derivation then `defer cancel()` on the next line — as a review convention. A rule enforced by a tool survives turnover; one that lives in reviewers' heads does not.
saying these in an interview costs you the question
- Says the garbage collector will clean up short-lived child contexts
- Claims cancelling after success would abort completed work
- Adds cleanup only to error paths and calls the audit done
- Treats a passing go vet run as proof no sites remain
- Blames the memory slope on the collector's pacing without evidence