Should every goroutine in your Go service defer a recover, or should panics be allowed to crash it?
answer
- it is an availability posture, not a style rule
- trust boundary versus core path
- restart cost sits on one side of the scale
- containment must never be silent
- shared-state mutation is the no-recover zone
basics
~20 sNeither extreme. Recover at trust boundaries where a goroutine runs code you do not own on one independent unit of work; let panics crash the process in your own core paths, because continuing on broken invariants is worse than a restart.
solid answer
~50 sThis is an availability posture, not a coding preference, and it belongs to whoever owns the service's uptime. Crashing is the honest default: a panic means an invariant already broke, and a restart returns the process to a known state with a loud, correlated signal — one panic in the crash log next to a restart record. Recovering buys blast-radius containment, which matters when one goroutine runs a callback registered by another team and a single bad handler would otherwise drop every subscriber's in-flight work. I draw the line at trust and independence: recover where the goroutine runs foreign code on one re-drivable unit of work, crash where the goroutine was mutating shared state. Anything recovered must be attributed, logged at error severity and counted, with an alert on the rate, or containment quietly becomes a way of shipping permanent bugs.
go deeper
You are not expected to set this policy, but know both directions exist: a recover keeps the process alive, and a crash gives a clean restart plus a loud signal.
Be ready to describe the tradeoff concretely — what a restart costs in in-flight work and cache warmth, and what a recover cannot repair in state the panicking code already touched.
Show that you would apply it per go statement using trust and independence as the tests, and that you attach attribution, severity, a metric and an alert to anything you recover.
Own the posture across the service: name the boundaries where recovery is permitted, state what a recovered panic must emit, weigh restart cost against blast radius, and be able to defend the call to an on-call owner who wants everything recovered.
## Why this is a decision and not a rule Go forces the question by design. A panic anywhere ends the process, and recovery can only be written inside each goroutine, so "what should happen when a goroutine panics" cannot be answered once in a framework — it is answered `go` statement by `go` statement. That makes it a policy somebody sets and can be overruled on, typically by whoever carries the pager or runs the incident review. ## The case for letting it crash - **A panic means the invariant is already gone.** Nil dereference, index out of range, a library's own consistency check. The code that continues after a recover is running on state nobody has vouched for. - **Restart is a known-good state.** A fresh process has no half-updated maps, no mutex left locked by a goroutine that died between `Lock` and `Unlock`, no partially written cache entry. - **The signal is loud and correlated.** A restart record next to a panic block is an unambiguous artefact for the crash log. A recovered panic logged at the wrong level is nearly invisible. - **It keeps the bug's cost on the team that owns it.** Recovery is a subsidy: it makes the defect survivable, and defects that are survivable rarely get fixed. ## The case for containing it - **Blast radius.** In a dispatcher that runs a callback per subscriber, a crash punishes every subscriber for one team's bug. Containment converts a total outage into a single failed delivery. - **Restart is not free.** Everything in flight is lost, caches start cold, and if the panic is input-triggered and that input keeps arriving, the process crash-loops — the platform's restart backoff then turns a bug into an outage of increasing duration. - **Some code is not yours to fix.** When goroutines execute handlers registered by other packages, you cannot make their code panic-free, and you should not let their quality decide your availability. ## Where I actually draw the line Two tests, both applied per `go` statement: 1. **Trust.** Is the goroutine executing code this service owns, or code registered by someone else? Foreign code gets a wrapper. 2. **Independence.** Is the goroutine's work one self-contained, re-drivable unit — one delivery, one message, one item — or is it mid-way through mutating state other goroutines will read? Independent units may be abandoned; shared-state mutation may not. That yields a small, defensible policy: wrappers at the trust boundary, no wrappers in the core, and an explicit ban on recovering inside code that holds a lock or is part-way through a persistent write. It is also easy to review, which matters more than elegance for a rule applied at hundreds of call sites. ## The conditions that keep containment honest A recover is only acceptable with all of these attached, and this is the part a lead should refuse to compromise on: - **Attribution** — which subscriber, which event, which handler. A recovered panic with no identity cannot be assigned to an owner. - **Severity** — logged at error, never debug, and never discarded with `_ = recover()`. - **A metric and an alert on the rate.** One recovered panic a week is a ticket; fifty a minute is an incident, and the difference must be visible without reading logs. - **An owner and a deadline.** Containment is a shock absorber, not a resolution. If a subscriber's handler keeps panicking, the escalation is to disable that subscriber, not to keep absorbing it forever. ## Constraints that legitimately change the answer - **Restart cost.** A stateless process that restarts in two seconds makes crashing cheap. One that rebuilds a large in-memory index for a minute, or loses thousands of queued deliveries, tips the balance towards containment. - **Who is harmed.** If a crash takes down work belonging to tenants who did not cause it, that is a multi-tenant fairness argument for containment independent of engineering taste. - **What the platform does.** If the supervisor restarts instantly and the deployment can absorb the churn, fail-fast is a smaller decision than it looks. If a crash loop trips a rollback or takes the fleet down instance by instance, it is a much bigger one. - **Correctness stakes.** In a path that moves money or writes durable records, continuing after an unexplained panic is a correctness risk that outranks availability, and the crash is the safer answer even when it is the more expensive one. ## How the disagreement usually goes The on-call owner argues that a partner's bad handler must never page them at 3am. The correctness owner argues that a service which recovers everything eventually corrupts something and nobody notices for a week. Both are right about their own risk. The settlement is not a percentage — it is naming the boundaries where recovery is permitted, writing down what a recovered panic must emit, and agreeing that a rising recovered-panic rate is treated as an incident rather than as background noise. Then the decision is auditable, and whoever owns uptime can change it deliberately rather than discovering it in a review.
- What must a recovered panic emit for the policy to stay honest?Attribution — which handler, which input, which tenant — a log line at error severity, and a counter with an alert on its rate. Without those, containment is indistinguishable from hiding the bug, and the defect survives indefinitely because nobody is paged, ticketed or embarrassed by it. Silent `_ = recover()` is the anti-pattern.
- Where would you refuse to recover even under a recover-everywhere policy?Anywhere the goroutine was mutating shared or durable state: holding a lock, mid-way through a multi-field update, or part-way through a write that must be all-or-nothing. Continuing there means running on state nobody can vouch for, and the resulting corruption is discovered far from its cause. Crash and restart is cheaper than that debugging.
- How does restart cost change the call?It sets the price of the fail-fast default. A stateless service that comes back in seconds makes crashing nearly free. One that loses thousands of in-flight deliveries, warms a cache for a minute, or trips a rollback on repeated restarts makes containment worth its risks — especially when the panic is input-triggered and the same input will crash the replacement process too.
- Who gets to overrule this decision?Whoever owns the service's availability and whoever owns its correctness, and they usually pull in opposite directions. The lead's job is to settle it explicitly — naming the boundaries where recovery is allowed and what it must report — so the posture is a written, auditable choice rather than an accident of whichever engineer wrote each go statement.
saying these in an interview costs you the question
- Wraps every goroutine in a recover as a blanket rule
- Treats recover as making the program correct again
- Discards the recovered value with no log or metric
- Ignores crash-loop and in-flight-work costs of fail-fast
- Recovers inside code holding a lock or mid durable write
- Frames it as personal style rather than an availability posture somebody owns