A Go HTTP service has recovery middleware, yet it still exits with a panic stack trace naming a handler. Why?
answer
- recover has no reach across goroutines
- grep the handler for a go statement
- the crash text bypassed your logger
- read the created by line
- exit status 2, not a graceful stop
basics
~20 sThe middleware's deferred recover only protects the goroutine running ServeHTTP. If a handler starts work with the go statement and that goroutine panics, no recover is on its stack, so the runtime prints a traceback to stderr and kills the whole process.
solid answer
~50 sRecovery middleware defers a `recover()` in the goroutine that runs `ServeHTTP`, and `recover` only ever sees panics from its own goroutine. A handler that fires off background work — `go audit(...)`, a cache refresh, a fire-and-forget notification — creates a goroutine whose stack has no deferred recover at all, so a panic there goes straight to the runtime: it prints the traceback and exits the process with status 2. The tell is where the output lands. Recovered panics appear in your structured logs or on `Server.ErrorLog`; this one appears raw on standard error, and the traceback's `created by` line names the handler that spawned it. The fix is that every goroutine started by a handler needs its own deferred recover, usually through a single spawn helper everyone uses, and that goroutine must not touch the `http.ResponseWriter` once the handler has returned. Better still, hand the work to a queue with a guarded worker rather than spawning per request.
code
go · 4 linesfunc recordHandler(w http.ResponseWriter, r *http.Request) {
go audit(r.URL.Path) // new goroutine, no recover on its stack
w.WriteHeader(http.StatusAccepted)
}go deeper
Remember the rule that recover only works inside the goroutine that panicked, so a wrapper around a handler cannot protect work the handler starts separately.
Explain the mechanics: the middleware's defer belongs to the ServeHTTP goroutine, a go statement creates a fresh unguarded stack, and an unrecovered panic anywhere exits the process.
Show the diagnosis from the artefacts — output on stderr rather than your logger, the created by line, exit status 2 — and then give a structural fix rather than patching one call site.
Decide the standard for background work in a codebase other teams contribute to: a single guarded spawn helper, a queue instead of per-request goroutines, and a review rule that makes bare go statements visible.
## The symptom A service has a recovery middleware, panics show up as 500s in the logs — and then one night the container restarts with a stack trace in its output and no matching 500. The middleware looks fine. It is fine. It is simply not in the picture. ## Why recover is per-goroutine `recover()` stops the unwinding of **the goroutine that is panicking**, and it only works when called from a function that goroutine deferred. There is no supervisor, no parent-child relationship, no way to install a handler that covers goroutines you started. When a goroutine panics and its own stack has no recover, the runtime treats it as fatal for the entire program: it prints the panic value and a traceback, and calls exit with status 2. Every other goroutine — including the ones with requests half-served — is killed where it stands. That is the whole explanation. The middleware wraps `next.ServeHTTP`, which runs on the request's goroutine. `go doWork()` starts a *different* goroutine. The wrapper's defer never runs on it. ## How handlers end up spawning goroutines It is rarely a deliberate concurrency design. It is usually one of: - fire-and-forget side effects: audit records, analytics events, webhook deliveries, cache warming; - "don't make the user wait" work that was moved off the request path in a hurry; - a fan-out that a later refactor left un-awaited; - a helper deep in a shared package that spawns internally, so the handler author never sees a `go` statement at all. The last one is the nastiest on a platform running handler code contributed by other teams: nothing in the handler's own file suggests a goroutine exists. ## Diagnosing it at 3am 1. **Look at where the output landed.** A recovered panic goes wherever your middleware logs, and a panic the server itself recovered goes to `Server.ErrorLog`. A fatal traceback is written by the runtime directly to standard error and is not routed through any logger. If the crash text is in the raw container stream and not in your structured logs, no recover was involved. 2. **Read the first stack, then the created by line.** The traceback starts with the goroutine that panicked. Below its frames is a `created by ...` line naming the function that started it — that is your handler or the helper it called, which is what ties a process crash back to an HTTP route. 3. **Check the exit code.** An unrecovered panic exits with status 2, which in an orchestrator shows as a crash rather than a graceful shutdown. 4. **Widen the traceback if you need the neighbours.** `GOTRACEBACK` defaults to printing only the panicking goroutine; `GOTRACEBACK=all` prints every goroutine's stack, which shows what else was in flight when the process died. 5. **Distinguish a real fatal error.** Text beginning `fatal error:` rather than `panic:` — concurrent map writes, the deadlock detector, out of memory — is a runtime throw that no recover could have caught anyway, and the fix is different. ## Fixing it - **Give every spawned goroutine its own deferred recover**, and route it into the same log and metric as the handler recoverer so a background panic is as visible as a request one. - **Funnel spawning through one helper** so the guard cannot be forgotten, and forbid bare `go` statements in handler packages by review convention. - **Prefer not spawning.** Do the work synchronously if it is fast, or push it onto a queue consumed by a small number of long-lived workers that already have their guard. A goroutine per request is also unbounded concurrency, which is a second problem. - **Never touch `w` after the handler returns.** The `http.ResponseWriter` is only valid for the duration of the call; a background goroutine that writes to it after the handler has returned is a bug independent of the panic. - **Detach the lifetime deliberately.** Work meant to outlive the request must not depend on the request's cancellation, and it must have a bound of its own. ## The sentence that answers the question "Recovery middleware protects one goroutine — the one running the handler. Anything the handler starts with `go` is unprotected, and an unrecovered panic in any goroutine takes the whole process with it."
- Why does that crash not appear in Server.ErrorLog along with the other panics?Server.ErrorLog only receives what net/http itself chooses to log, and the server only logs panics it recovered on a connection goroutine. A fatal, unrecovered panic is printed by the runtime straight to standard error before the process exits, so it never passes through any logger you configured. In a container it lands in the raw stream instead of your structured logs.
- How would you stop this class of bug rather than fix one instance?Route all background work through one helper that starts the goroutine with a deferred recover and reports through the same log and metric as the request recoverer, then ban bare go statements in handler packages by review convention. Where the work is not genuinely per-request, replace the spawn with a queue and a small pool of long-lived guarded workers, which also bounds concurrency.
- What does GOTRACEBACK change about the output you get from that crash?By default the runtime prints only the panicking goroutine's stack. Setting GOTRACEBACK=all prints every goroutine's stack, which shows what else was in flight when the process died — often the fastest way to see how many requests you dropped. Higher settings add runtime frames and can raise SIGABRT for a core dump, at the cost of much larger output.
The middleware is a safety net strung under one trapeze. The handler quietly sent a second acrobat to a different rig in the same tent — and when that one falls, the tent comes down.
saying these in an interview costs you the question
- Claims one recover in the middleware covers spawned goroutines
- Says a panic in a background goroutine only kills that goroutine
- Suggests recovering in the parent goroutine after the go statement
- Has the background goroutine write to the ResponseWriter after return
- Confuses a fatal runtime error with an unrecovered panic