For a Go service hosting other teams' handlers, do you recover panics into 500s or let the process crash?
answer
- surviving is already the default
- one request against every tenant
- what did the unwinding skip
- make recovery loud, not quiet
- shed the route before killing the process
basics
~20 sDefault to recovering into a 500 so one bad request does not drop everyone else's in-flight work, but make each panic expensive: a high-severity event with the route and an owner. Reserve crash-only for handlers whose failure could leave shared state unsound.
solid answer
~60 sStart from what Go already does: `net/http` recovers per connection, so surviving is the default and crash-only is something you have to implement on purpose by re-panicking or exiting from your recoverer. On a platform running other teams' handlers I keep the recover, because a panic is usually one unlucky request and crashing punishes every other tenant with in-flight work, plus it invites a crashloop when a poison request repeats. The price of recovering is that the process continues past an invariant nobody proved is intact — a mutex left locked because the unlock was not deferred, a half-applied mutation — so recovery must never be quiet: every panic emits a high-severity event carrying the route and the owning package, panics are rate-alerted rather than ignored, and a route panicking above a threshold gets shed with a 503 instead of taken out on the whole process. Crash-only is the right posture for a narrow class: handlers that mutate process-wide state where continuing risks serving wrong answers. That choice belongs to whoever carries the pager, so I agree it with SRE and write it down.
go deeper
Know that Go's HTTP server keeps running after a handler panic by default, and that turning that into a 500 is what a recovery middleware is for.
Be able to argue both sides briefly: recovering limits the blast radius to one request, crashing avoids serving from a process whose invariants may already be broken.
Bring the operational detail — poison requests and crashloops, locks not released because the unlock was not deferred, per-route shedding, and alerting on panic rate rather than on each occurrence.
Own the policy end to end: one platform-wide recoverer, panic events with a named owner and an error budget, a written list of crash-only cases, and an agreed contract with whoever carries the pager.
## The question behind the question This is not "is recover good". It is: when your process discovers that one of its own invariants was violated, do you keep serving or do you stop? On a service that hosts handler code you did not write and cannot fix, the answer is a policy, and the person carrying the pager has a legitimate veto. ## Start from Go's default, because it is already a choice `net/http` recovers panics at the connection layer. So a Go HTTP service **already** survives handler panics with no code from you. "Let it crash" is therefore not the lazy option; it is an active decision you implement by re-panicking out of your recoverer, or by calling `os.Exit` after logging. State this early — candidates who think they are choosing between crashing and recovering have the baseline wrong. ## The case for recovering into a 500 - **Blast radius.** A panic is normally one request meeting one unhandled shape — a nil map write on an unusual payload, an index computed from untrusted input. Crashing converts a one-request bug into an outage for every in-flight request of every other team on the box. - **Poison requests.** If a client retries the request that panicked, crash-only turns into a crashloop, and an orchestrator's backoff makes the outage longer than the bug deserves. - **Observability.** Recovering lets you attach the route, the request id and the owning package to the failure. A crash gives you a traceback on stderr and nothing else correlated. - **The client gets an answer.** A 500 is a status your callers, your load balancer and your SLO maths can all reason about. A dropped connection looks like a network fault and pollutes someone else's error budget. ## The case for crashing - **A panic means you were wrong about something.** The unwinding skipped every cleanup that was not deferred. A `sync.Mutex` unlocked on the happy path only is now locked forever, so the next request touching it hangs — a far worse failure than a restart, and much harder to diagnose. - **Shared mutable state.** If the panicking handler was midway through mutating a process-wide cache, an in-memory index or a file, continuing can serve wrong answers silently. Wrong answers beat down trust faster than 500s do. - **Restarting is cheap and honest.** With a fast start-up, a readiness probe and connection draining, an orchestrator restart is measured in seconds and gives you a known-good state. - **Recovering makes panics free.** If a panic costs nobody anything, the contributing team never fixes it, and the recoverer quietly becomes the error-handling strategy. ## The posture I would actually own 1. **Recover by default, at the edge, exactly once.** One recoverer for the whole platform, not one per team, so the behaviour is uniform and auditable. 2. **Make every panic expensive.** A structured high-severity event with route, method, panic value, stack, and the package that appears first in the trace — which names the owning team. A panic without an owner is a platform bug by default. 3. **Alert on rate, not on instances.** One panic per week is a ticket; a route panicking on 5% of requests is a page. Give panics an error budget like any other failure class. 4. **Shed before you crash.** If a single route is panicking above a threshold, disable that route and return 503 for it. That contains a broken tenant without punishing the rest, and it is a much better lever than process restart. 5. **Crash-only for a named list.** Handlers holding process-wide mutable state whose consistency you cannot verify after an abort. Write down which ones, and why, so the decision is reviewable rather than folklore. 6. **Accept what you cannot recover.** Fatal runtime errors — concurrent map writes, the deadlock detector, out of memory — end the process no matter what you decide, and panics in goroutines a handler spawns do too. Your policy must cover being restarted anyway: fast startup, drain on shutdown, no work lost that was not already durable. 7. **Agree the operational contract.** What happens after N panics per minute? Does the instance fail readiness and get rotated — which gets you a clean restart *and* graceful draining, the best of both — or does it page a human? That is SRE's call as much as yours, and it should be in the runbook, not in a middleware nobody has read. ## What a weak answer sounds like "Always recover, a 500 is fine" — ignores that the process may be unsound. "Crash-only is correct, that is how Erlang does it" — borrows a supervision model Go does not have; Go has no per-goroutine supervisor and no isolated heaps, so a crash is process-wide, not actor-scoped. "It depends" with no criteria — the criteria are blast radius, state soundness, restart cost and who is paged. ## The closing line Recover so that one request's bug stays one request's bug; make the recovery loud enough that it gets fixed; and reserve crashing for the cases where continuing would mean serving answers you cannot defend.
- What would make you switch a specific route to crash-only?Evidence that a panic there can leave process-wide state unsound — it mutates a shared in-memory index, holds a lock released without defer, or writes a file in stages — combined with a cheap restart and effective draining. The test is whether continuing risks serving wrong answers rather than errors; wrong answers justify a restart, 500s usually do not.
- How do you stop recovered panics from becoming background noise nobody fixes?Give them an owner and a budget. Every panic event carries the route and the first non-platform package in the stack, which names the contributing team; panics are counted per route with a threshold that pages; and a route over budget is shed with a 503 until it is fixed. Reviewing the panic list in the same forum as other incidents keeps the count from drifting upward quietly.
- SRE wants the process to restart on any panic. How do you resolve that?Separate the two things they actually want: a known-good process and a loud signal. Offer failing readiness after a panic rate threshold so the instance is drained and rotated rather than killed mid-request, plus per-route shedding for a single bad tenant. If they still want crash-only for a specific class of state, agree that list explicitly and put it in the runbook rather than arguing per incident.
saying these in an interview costs you the question
- Recovers everything and treats a 500 as the end of the story
- Insists crash-only is universally correct in Go
- Assumes a recovered process is in a consistent state
- Ignores that a poison request turns crash-only into a crashloop
- Has no owner or threshold attached to panic events
- Forgets that fatal runtime errors end the process regardless