Why does a runtime.AddCleanup cleanup never run when its closure captures the pointer it was attached to?
answer
- the registration is itself a root
- you captured the thing you are waiting on
- alive because of its own registration
- pass a key, never the object
- silent leak, no panic
basics
~20 sThe runtime holds the cleanup function and its argument until the cleanup runs, so capturing the pointer keeps the object permanently reachable. It never becomes garbage, so its memory is never freed and the cleanup can never fire.
solid answer
~50 s`runtime.AddCleanup(ptr, cleanup, arg)` stores the function and its argument in the runtime's own bookkeeping, and that bookkeeping is a root: whatever is reachable from it is live. If the closure captures `ptr`, or if `arg` leads back to the object through a field, then the object is reachable from the very registration that only fires once the object stops being reachable. It can never fire. The result is a silent leak of both the object and whatever resource the cleanup was meant to release — no panic, no error, just a heap that grows. The rule is to pass the smallest independent value the cleanup needs, such as a map key, a file descriptor, or a small state struct that the object points at but which does not point back. The same shape explains why a reference cycle containing an object with a `runtime.SetFinalizer` finalizer is not guaranteed to be collected.
code
go · 7 linese := &entry{key: "doc-42"}
// WRONG: the closure captures e, so e is reachable from its own registration
runtime.AddCleanup(e, func(string) { dropKey(e.key) }, e.key)
// RIGHT: the cleanup is handed the key and can reach nothing else
runtime.AddCleanup(e, func(k string) { dropKey(k) }, e.key)go deeper
Remember one rule: the cleanup function must not mention the object. Give it the small value it needs, such as a key or a file descriptor, and nothing that leads back to the object itself.
Explain why the capture is fatal. The runtime keeps the function and its argument alive as a root, so an object reachable from its own registration can never become unreachable and the cleanup can never become eligible.
Treat this as a review reflex and a test you write. Show the test that drops the reference, forces collection and fails on a timeout, and be able to say what a timeout in CI actually means.
Decide whether these hooks are allowed in shared code at all, and what evidence you demand before they ship: a leak test in the package's own suite, or a counter proving the backstop is not silently carrying the codebase.
## Roots, and why a registration is one Go's collector decides what is alive by tracing from a set of roots — goroutine stacks, globals, and various runtime-internal structures. Anything reachable from a root is live; everything else is garbage. When you call `runtime.AddCleanup(ptr, cleanup, arg)`, the runtime must keep `cleanup` and `arg` somewhere until it fires, and that somewhere is traced like any other root. Now put the two facts together. The cleanup becomes eligible only when `ptr` is unreachable. But if `cleanup` or `arg` can reach `ptr`, then as long as the registration exists the object is reachable, and the registration exists until the cleanup fires. The condition for firing is exactly the condition the registration itself prevents. Nothing breaks the loop. ## The two ways to make the mistake **The closure capture.** The natural, tempting way to write a cleanup is to reach for the object's fields inside the function body: ``` runtime.AddCleanup(e, func(string) { dropKey(e.key) }, e.key) ``` That closure captures `e`. The closure is stored by the runtime. `e` is now permanently live. **The argument that leads back.** More subtle: `arg` need not *be* the object to pin it. If `arg` is a pointer to a struct that has a field pointing at the object — even several hops away — the object is still reachable from the registration. Reachability is transitive. Whatever `arg` can reach stays alive from registration until the cleanup runs. The API refuses the most obvious form of the mistake, passing the watched pointer itself as the argument, but nothing can see through a closure's captured variables or a chain of struct fields. It is a code-review property, not a compile-time one. ## Why it is so hard to notice The failure is entirely silent. No panic, no error return, no log line. The object simply never dies, so the heap grows in proportion to how many of these you create, and the resource the cleanup would have released is never released either. In a service, this looks like a slow, steady memory climb with no obvious owner, and the object graph in a heap profile shows the objects retained by the runtime rather than by any application structure — which reads as "retained by nothing" to somebody skimming. ## The symmetrical trap in the older API `runtime.SetFinalizer` has a version of the same problem built into it: the finalizer is *always* handed the pointer, so any two objects that reference each other where at least one carries a finalizer form a cycle with no safe finalization order. The documented behaviour is that such a cycle is not guaranteed to be collected and the finalizers are not guaranteed to run. `AddCleanup` was designed to remove that class of bug — which it does for cycles among the watched objects, while leaving the capture mistake possible if you write it by hand. ## How to write it so it cannot happen Decide, at registration time, the minimum data the cleanup needs, and copy it out of the object then. A map key. An integer descriptor. A pointer to a small state block that the object holds a pointer to, but which holds no pointer back to the object — that direction matters and is the standard way to give a cleanup access to mutable state, such as a `sync.Once` guarding a double release. When a method is the natural body, use a method expression on that separate state value rather than on the object itself, so the function value carries no reference to the wrapper. ## Proving it in a test Because the failure is silent, the useful move is to make it loud in the test suite. Register a cleanup whose argument is a channel, drop the only reference to the object, force collection, then wait on the channel with a timeout: - if the channel closes, the object really did become unreachable and the cleanup fired; - if the timeout fires, something still references the object — which is precisely the bug. Calling `runtime.GC()` more than once is the usual belt-and-braces, and the timeout has to be generous, because cleanup execution is asynchronous even after the object is found unreachable. This test is worth having in any package that ships a runtime hook: it is the only mechanical way to catch a capture that a future edit reintroduces.
- How would you prove in a test that a cleanup actually runs?Register a cleanup whose argument is a channel, drop the only reference to the object, call `runtime.GC()` (twice is the usual belt-and-braces), then wait on the channel with a generous timeout. The timeout branch is the assertion that matters: it fires exactly when something still references the object, which is the bug the test exists to catch.
- What is the equivalent trap with runtime.SetFinalizer?Two objects that point at each other where at least one carries a finalizer. There is no order that respects the dependencies, so the documented behaviour is that the cycle is not guaranteed to be collected and the finalizers are not guaranteed to run. `runtime.AddCleanup` has no such problem, because the callback holds no pointer into the cycle.
- If a cleanup needs several fields of the object, how do you register it safely?Copy them out at registration time into a separate value and pass that as `arg` — a small struct, or a pointer to a state block that the object references but which holds no pointer back. Keep it minimal: everything `arg` can transitively reach stays alive from registration until the cleanup runs.
saying these in an interview costs you the question
- Wraps the object in the closure to reach its fields
- Expects a panic or an error when the cleanup can never run
- Thinks the runtime holds the cleanup argument only weakly
- Blames the collector rather than the registration
- Passes a struct pointer that points back at the object