skip to content

Where must the deferred recover live when a controller runs each reconcile in its own goroutine?

level: middleimportance: must knowfreq 60%

answer

  1. the starter has already moved on
  2. defers belong to one stack only
  3. one boundary per go statement
  4. a goroutine has no return value
  5. route the failure back on a channel

basics

~20 s

Inside the goroutine's own function, at the top, not in the code that started it. A recovery only takes effect in a function deferred by the goroutine that panicked, and a panic that escapes any goroutine terminates the whole process.

solid answer

~50 s

Every `go` statement needs its own boundary. The function that launched the goroutine has already moved on, and its deferred recovery belongs to a different stack, so it can never see the child's panic; if nothing in the child recovers, the runtime kills the entire process, not just that goroutine. So the goroutine's function literal opens with `defer func() { if r := recover(); r != nil { ... } }()`. The second half is getting the failure back out: a goroutine has no return value, so the handler must hand the error somewhere explicit — send it on the results channel the reconcile loop reads, or record it against the item's key under a mutex — otherwise the boundary contains the crash but loses the report. In practice I ban bare `go doWork(...)` calls and route every spawn through one helper that installs the boundary.

code

go · 12 lines
go
type Key struct{ Namespace, Name string }

func (c *Controller) startReconcile(k Key) {
	go func() {
		defer func() {
			if r := recover(); r != nil {
				c.results <- fmt.Errorf("reconcile %v: recovered panic: %v", k, r)
			}
		}()
		c.results <- c.reconcile(k)
	}()
}

go deeper

for a junior

Remember the rule and its consequence: the deferred recovery goes inside the goroutine's own function, and a panic nobody recovers ends the entire program, not just that goroutine.

for a middle

Explain the mechanics — defer chains are per-goroutine, the starter's frame is already gone — and show how the recovered failure is routed back, since a goroutine has no return value to carry it.

for a senior

Talk about making the boundary structural: a single spawn helper, identity of the failed item in the error, requeue behaviour, and honesty about what the boundary cannot catch.

for a principal

Set the standard for concurrency in the codebase: who may write a bare go statement, what every spawned unit must report, and how that reporting feeds alerting and backpressure rather than just a log line.

## Why placement is the whole question Recovery in Go is scoped to a single goroutine's stack. `recover` returns a value only when it is called by a function that the **panicking goroutine itself** deferred. A controller that spawns one goroutine per reconcile therefore has as many potential crash sites as it has in-flight goroutines, and the loop that started them protects none of them. The consequence is severe and specific to Go: an unrecovered panic in *any* goroutine terminates the **whole process**. It does not quietly kill one worker the way an uncaught exception on a worker thread does in several other runtimes. One malformed object in one reconcile can take down every other reconcile in flight. ## The wrong shape ```go func (c *Controller) start(k Key) { defer func() { if r := recover(); r != nil { /* never fires */ } }() go c.reconcile(k) } ``` Two separate things are wrong. The deferred closure belongs to `start`'s stack, and `start` returns immediately — long before the child goroutine does any work. And even if it were still running, the recovery would be on the wrong goroutine's defer chain. ## The right shape ```go func (c *Controller) startReconcile(k Key) { go func() { defer func() { if r := recover(); r != nil { c.results <- fmt.Errorf("reconcile %v: recovered panic: %v", k, r) } }() c.results <- c.reconcile(k) }() } ``` The boundary is the first statement of the function the `go` statement runs. Everything the goroutine does afterwards is inside it. ## Getting the failure back out A goroutine's return values are discarded — `go f()` has nowhere to put them — so the recovery handler must route the failure deliberately. The usual choices: - **A channel.** Send the wrapped error on the same channel the goroutine would have sent a normal result on. The reconcile loop receives one value per item either way and can requeue the failed key. - **Shared state under a lock.** Record the failure against the item's key in the status map, with the mutex held. Never write a shared variable from the handler without synchronisation; that swaps a panic for a data race. - **A callback supplied by the starter.** Useful for a library where the caller decides what a failed unit of work means. Whichever you choose, include the key of the item that failed. A recovered panic with no identity is nearly useless when a thousand objects are reconciling and one of them is poisonous. ## Making it structural rather than a habit The failure mode here is a missed `go` statement, not a misunderstood one. Every developer on the team knows the rule; someone still writes `go c.reconcile(k)` in a hurry. Two practices help: 1. **One spawn helper.** Give the controller a single `func (c *Controller) spawn(k Key, f func() error)` that installs the boundary and publishes the result, and require all concurrency to go through it. Now the boundary exists in exactly one place and can be reviewed once. 2. **Review rule.** A bare `go` keyword in a long-lived service is a review comment by default. It is far easier to spot in a diff than to debug after a 3am process restart. ## Panics the boundary cannot catch Be honest about the limits when you describe this. A per-goroutine recovery stops panics raised on that goroutine's stack. It does nothing about a goroutine that a library starts internally, and it does nothing about runtime failures that are not panics at all — concurrent map writes, for instance, abort the program without giving deferred functions a chance to run. The boundary is a containment mechanism for bugs in *your* reconcile path, not a general guarantee that the process survives. ## Summary The deferred recovery goes at the top of the function that `go` runs, once per spawn, with the failure routed explicitly back to the loop that cares. The starter's own defers are on the wrong stack and will never fire, and the price of getting it wrong is the entire process rather than one unit of work.

  • The recovery handler has the panic value but the goroutine has no return path. How do you get it back to the reconcile loop?
    Send it where a normal outcome would have gone: the results channel the loop already receives from, so a failed item and a successful one arrive through the same path and the loop can requeue the key. Recording it in a shared status map works too, but only with the mutex held — an unsynchronised write turns the panic into a data race.
  • How do you stop someone reintroducing a bare `go reconcile(k)` six months later?
    Make the boundary structural rather than remembered: a single `spawn` helper on the controller that takes the work function, installs the deferred recovery and publishes the outcome, with a review rule that a bare `go` keyword in long-lived service code needs justification. One reviewed implementation beats a convention everyone has to recall.
  • Does a per-goroutine recovery mean the process can no longer die unexpectedly?
    No. It contains panics raised on the goroutines you wrapped. Goroutines started inside a dependency have their own stacks and their own missing boundaries, and some runtime failures are not panics at all — a concurrent map write aborts the program without running deferred functions, so no recovery can intercept it.

saying these in an interview costs you the question

  • Putting the deferred recovery in the function that starts the goroutine
  • Believing an unrecovered panic only kills the goroutine it happened in
  • Recovering in the goroutine but never reporting the failure anywhere
  • Writing the error to a shared variable with no synchronisation
  • Assuming a wrapper around the outermost call protects spawned work