skip to content

How do you stop one subscriber's panicking callback from killing a webhook dispatcher process?

level: seniorimportance: should knowfreq 45%

answer

  1. the guard has to ride with the work
  2. wrap the callback, not the go statement
  3. one recover per unit, not per worker
  4. a panic becomes a value on a channel
  5. exactly one send on every path

basics

~20 s

Run each callback in a goroutine whose body begins with a deferred recover, so the panic is contained there and becomes a value. Send that value back on a results channel, attributed to the subscriber.

solid answer

~50 s

The dispatcher cannot guard callbacks from outside, so each dispatch goroutine runs a wrapper whose first statement is `defer func() { if r := recover(); r != nil { … } }()`. The wrapper calls the registered callback and, on the recovery path, turns the recovered value into an error identifying which subscriber failed, sending it on the results channel exactly as the success path sends a nil error — one send on every path. Placement is what people get wrong: the recover must sit inside the per-callback function, so one bad subscriber ends only that dispatch. A single recover at the top of a worker that loops over callbacks still lets the panic unwind the loop, and the worker exits. Containment also proves nothing about state the callback corrupted, so the recovered panic must be logged and counted, never swallowed.

code

go · 14 lines
go
type result struct {
	sub string
	err error
}

func runOne(sub string, h func(), out chan<- result) {
	defer func() {
		if r := recover(); r != nil {
			out <- result{sub: sub, err: fmt.Errorf("handler panicked: %v", r)}
		}
	}()
	h()
	out <- result{sub: sub} // skipped when h panics
}

go deeper

for a junior

Know the shape by heart: the function a goroutine runs starts with defer func() { if r := recover(); r != nil { ... } }(), and that guard only helps if it is inside that goroutine.

for a middle

Explain why placement decides the outcome: a defer runs at function return, so a recover above a loop lets the panic unwind the loop and end the worker, while a per-callback wrapper keeps the rest running.

for a senior

Demonstrate the full containment: attribute the panic to a subscriber, send exactly one result on every path, log at error severity with a metric, and say plainly what the recover cannot repair — a held lock, half-mutated shared state, a goroutine the callback started.

for a principal

Own the standard: which boundaries are allowed to recover, what a recovered panic must emit before the service continues, and how a rising recovered-panic rate turns into an owner and a fix instead of permanent background noise.

## The situation A dispatcher holds a registry of callbacks supplied by other packages — one per subscriber — and runs them when an event arrives. It starts a goroutine per dispatch so a slow subscriber does not hold up the rest. Then one subscriber's callback dereferences a nil pointer and the entire dispatcher process dies, dropping every delivery in flight for every subscriber. The crash log has one panic in it; the platform records a restart. Because a panic anywhere terminates the process, and because a recover only works on the goroutine that panicked, the fix must live inside the goroutine body. ## The wrapper ```go type result struct { sub string err error } func runOne(sub string, h func(), out chan<- result) { defer func() { if r := recover(); r != nil { out <- result{sub: sub, err: fmt.Errorf("handler panicked: %v", r)} } }() h() out <- result{sub: sub} } ``` Three properties make this work: 1. **The deferred function is registered by the goroutine that will run `h`.** `go runOne(sub, h, out)` puts the recover on the right stack. Nothing the dispatcher writes around the `go` statement could do this. 2. **The recovered value becomes an ordinary value.** A panic cannot be returned; a value can. The dispatcher learns about the failure the same way it learns about a successful delivery. 3. **Exactly one send happens on every path.** If `h` panics, the trailing send is skipped and the deferred send happens instead. If a collector is counting one result per subscriber, that invariant is what stops it from blocking forever or double-counting. ## Placement: per unit of work, not per worker The most common half-fix is a single deferred recover at the top of a long-lived worker goroutine that loops over callbacks: ```go go func() { defer func() { _ = recover() }() // wrong granularity for job := range jobs { job.handler() } }() ``` This prevents the process from dying, but deferred functions run when the function **returns**, and a panic unwinds the whole frame including the loop. The recover therefore fires as the worker exits, and the worker is gone: every remaining callback goes unrun and the queue backs up silently. The recover must be inside a function that covers **one** callback — either a wrapper called per iteration, or the goroutine body when the goroutine handles a single dispatch. ## What the wrapper does not buy you - **It does not repair state.** If the callback took a mutex and panicked before releasing it (no `defer` on the unlock), that mutex stays locked forever and the next dispatch deadlocks. If it half-updated a shared structure, the wrapper cannot tell. - **It does not cover goroutines the callback starts.** If the subscriber's code writes its own `go`, a panic in that goroutine has its own stack, no recover on it, and it kills the process regardless of your wrapper. - **It does not make the failure benign.** A recovered panic is still a bug in someone's code, and if the dispatcher hides it at debug level, the subscriber never learns and the defect lives forever. So the contract for a recovered panic is: attribute it (which subscriber, which event), log it at error severity, count it as a metric, and treat a rising rate as an incident. That is what keeps containment honest rather than turning it into a silencer. ## Reporting the failure back Sending a `result` on a channel is the plain version. Two variants come up: - The dispatcher waits on a `sync.WaitGroup` and each wrapper writes its outcome into a preallocated slot; no channel capacity to size, and the recover path just assigns to the slot. - The dispatcher uses `golang.org/x/sync/errgroup` to collect the first error. Note that `errgroup.Group.Go` does **not** recover panics for you, and neither does `sync.WaitGroup.Go`; both run your function on a fresh goroutine, so a panic in it still kills the process. The deferred recover is still yours to write inside the function you hand them. Whichever you use, make sure the wrapper cannot leave the collector waiting: an unbuffered results channel with a collector that has already given up will block the recover path's send and leak the goroutine, so pair the send with buffering or with a select on the dispatcher's cancellation channel. ## What to verify in review For each `go` statement that runs foreign code: is the launched function a wrapper? Is `defer` + `recover` its first statement? Does every exit path — normal, panic, early return — send or record exactly one outcome? Is the recovered value logged with the subscriber's identity? If all four hold, one bad callback costs one delivery instead of the whole process.

  • A worker goroutine defers one recover at the top of its body and then loops over callbacks. What happens after the third callback panics?
    The panic unwinds the entire function frame, loop included, so the deferred recover runs as the goroutine returns. The process survives, but the worker is gone and every remaining callback is never run — usually a silent stall rather than a crash. The recover has to be inside a per-callback function for the loop to continue.
  • What must the wrapper guarantee about the results channel so the dispatcher does not hang?
    Exactly one send on every path: the panic path sends instead of, not in addition to, the success path. The collector counting one result per subscriber then always completes. And if the dispatcher can give up early, the sends need a buffer or a select on cancellation, or the recover path blocks forever and leaks the goroutine.
  • Does the wrapper protect you if the subscriber's callback starts its own goroutine?
    No. That goroutine has a separate stack, your deferred recover is not on it, and an unrecovered panic there terminates the process just as before. Containment only reaches code the wrapper's own goroutine executes, which is a real limit when running callbacks registered by other packages.
  • What is left broken even after a successful recover?
    Anything the callback mutated before it panicked. A mutex taken without a deferred unlock stays locked and the next dispatch deadlocks; a shared map or struct may be half-updated; an external write may be half-applied. Recovery contains the crash, not the corruption — which is why the recovered panic must be reported and fixed rather than absorbed.

saying these in an interview costs you the question

  • Puts the recover around the go statement in the dispatcher
  • One recover at the top of a worker that loops over many callbacks
  • Swallows the recovered value with no log, metric or attribution
  • Assumes errgroup or WaitGroup.Go recovers panics for you
  • Claims a recover restores state the callback already corrupted
  • Sends on the results channel twice, or not at all, on the panic path