skip to content

For a Go service hosting other teams' handlers, do you recover panics into 500s or let the process crash?

level: principalimportance: should knowfreq 26%

answer

  1. surviving is already the default
  2. one request against every tenant
  3. what did the unwinding skip
  4. make recovery loud, not quiet
  5. shed the route before killing the process

basics

~20 s

Default to recovering into a 500 so one bad request does not drop everyone else's in-flight work, but make each panic expensive: a high-severity event with the route and an owner. Reserve crash-only for handlers whose failure could leave shared state unsound.

solid answer

~60 s

Start from what Go already does: `net/http` recovers per connection, so surviving is the default and crash-only is something you have to implement on purpose by re-panicking or exiting from your recoverer. On a platform running other teams' handlers I keep the recover, because a panic is usually one unlucky request and crashing punishes every other tenant with in-flight work, plus it invites a crashloop when a poison request repeats. The price of recovering is that the process continues past an invariant nobody proved is intact — a mutex left locked because the unlock was not deferred, a half-applied mutation — so recovery must never be quiet: every panic emits a high-severity event carrying the route and the owning package, panics are rate-alerted rather than ignored, and a route panicking above a threshold gets shed with a 503 instead of taken out on the whole process. Crash-only is the right posture for a narrow class: handlers that mutate process-wide state where continuing risks serving wrong answers. That choice belongs to whoever carries the pager, so I agree it with SRE and write it down.

go deeper

for a junior

Know that Go's HTTP server keeps running after a handler panic by default, and that turning that into a 500 is what a recovery middleware is for.

for a middle

Be able to argue both sides briefly: recovering limits the blast radius to one request, crashing avoids serving from a process whose invariants may already be broken.

for a senior

Bring the operational detail — poison requests and crashloops, locks not released because the unlock was not deferred, per-route shedding, and alerting on panic rate rather than on each occurrence.

for a principal

Own the policy end to end: one platform-wide recoverer, panic events with a named owner and an error budget, a written list of crash-only cases, and an agreed contract with whoever carries the pager.

## The question behind the question This is not "is recover good". It is: when your process discovers that one of its own invariants was violated, do you keep serving or do you stop? On a service that hosts handler code you did not write and cannot fix, the answer is a policy, and the person carrying the pager has a legitimate veto. ## Start from Go's default, because it is already a choice `net/http` recovers panics at the connection layer. So a Go HTTP service **already** survives handler panics with no code from you. "Let it crash" is therefore not the lazy option; it is an active decision you implement by re-panicking out of your recoverer, or by calling `os.Exit` after logging. State this early — candidates who think they are choosing between crashing and recovering have the baseline wrong. ## The case for recovering into a 500 - **Blast radius.** A panic is normally one request meeting one unhandled shape — a nil map write on an unusual payload, an index computed from untrusted input. Crashing converts a one-request bug into an outage for every in-flight request of every other team on the box. - **Poison requests.** If a client retries the request that panicked, crash-only turns into a crashloop, and an orchestrator's backoff makes the outage longer than the bug deserves. - **Observability.** Recovering lets you attach the route, the request id and the owning package to the failure. A crash gives you a traceback on stderr and nothing else correlated. - **The client gets an answer.** A 500 is a status your callers, your load balancer and your SLO maths can all reason about. A dropped connection looks like a network fault and pollutes someone else's error budget. ## The case for crashing - **A panic means you were wrong about something.** The unwinding skipped every cleanup that was not deferred. A `sync.Mutex` unlocked on the happy path only is now locked forever, so the next request touching it hangs — a far worse failure than a restart, and much harder to diagnose. - **Shared mutable state.** If the panicking handler was midway through mutating a process-wide cache, an in-memory index or a file, continuing can serve wrong answers silently. Wrong answers beat down trust faster than 500s do. - **Restarting is cheap and honest.** With a fast start-up, a readiness probe and connection draining, an orchestrator restart is measured in seconds and gives you a known-good state. - **Recovering makes panics free.** If a panic costs nobody anything, the contributing team never fixes it, and the recoverer quietly becomes the error-handling strategy. ## The posture I would actually own 1. **Recover by default, at the edge, exactly once.** One recoverer for the whole platform, not one per team, so the behaviour is uniform and auditable. 2. **Make every panic expensive.** A structured high-severity event with route, method, panic value, stack, and the package that appears first in the trace — which names the owning team. A panic without an owner is a platform bug by default. 3. **Alert on rate, not on instances.** One panic per week is a ticket; a route panicking on 5% of requests is a page. Give panics an error budget like any other failure class. 4. **Shed before you crash.** If a single route is panicking above a threshold, disable that route and return 503 for it. That contains a broken tenant without punishing the rest, and it is a much better lever than process restart. 5. **Crash-only for a named list.** Handlers holding process-wide mutable state whose consistency you cannot verify after an abort. Write down which ones, and why, so the decision is reviewable rather than folklore. 6. **Accept what you cannot recover.** Fatal runtime errors — concurrent map writes, the deadlock detector, out of memory — end the process no matter what you decide, and panics in goroutines a handler spawns do too. Your policy must cover being restarted anyway: fast startup, drain on shutdown, no work lost that was not already durable. 7. **Agree the operational contract.** What happens after N panics per minute? Does the instance fail readiness and get rotated — which gets you a clean restart *and* graceful draining, the best of both — or does it page a human? That is SRE's call as much as yours, and it should be in the runbook, not in a middleware nobody has read. ## What a weak answer sounds like "Always recover, a 500 is fine" — ignores that the process may be unsound. "Crash-only is correct, that is how Erlang does it" — borrows a supervision model Go does not have; Go has no per-goroutine supervisor and no isolated heaps, so a crash is process-wide, not actor-scoped. "It depends" with no criteria — the criteria are blast radius, state soundness, restart cost and who is paged. ## The closing line Recover so that one request's bug stays one request's bug; make the recovery loud enough that it gets fixed; and reserve crashing for the cases where continuing would mean serving answers you cannot defend.

  • What would make you switch a specific route to crash-only?
    Evidence that a panic there can leave process-wide state unsound — it mutates a shared in-memory index, holds a lock released without defer, or writes a file in stages — combined with a cheap restart and effective draining. The test is whether continuing risks serving wrong answers rather than errors; wrong answers justify a restart, 500s usually do not.
  • How do you stop recovered panics from becoming background noise nobody fixes?
    Give them an owner and a budget. Every panic event carries the route and the first non-platform package in the stack, which names the contributing team; panics are counted per route with a threshold that pages; and a route over budget is shed with a 503 until it is fixed. Reviewing the panic list in the same forum as other incidents keeps the count from drifting upward quietly.
  • SRE wants the process to restart on any panic. How do you resolve that?
    Separate the two things they actually want: a known-good process and a loud signal. Offer failing readiness after a panic rate threshold so the instance is drained and rotated rather than killed mid-request, plus per-route shedding for a single bad tenant. If they still want crash-only for a specific class of state, agree that list explicitly and put it in the runbook rather than arguing per incident.

saying these in an interview costs you the question

  • Recovers everything and treats a 500 as the end of the story
  • Insists crash-only is universally correct in Go
  • Assumes a recovered process is in a consistent state
  • Ignores that a poison request turns crash-only into a crashloop
  • Has no owner or threshold attached to panic events
  • Forgets that fatal runtime errors end the process regardless