skip to content

Panic Policy and Boundaries

panic is Go's answer to a broken program, not to a failed operation, and recover works only inside a deferred function. The policy question is where crashing is honest and where it is rude.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

explore

questions

17

How do you stop a panic from escaping a Go function and return it to the caller as an error?

level: juniorimportance: must knowfreq 68%

answer

  1. the caller should never see a panic
  2. only a deferred call can intervene
  3. an anonymous result cannot be reassigned
  4. set err inside the deferred closure
  5. register it before the risky work

basics

~20 s

Give the function a named error result and defer a closure that calls recover. When recover returns non-nil, the closure assigns a wrapped error to that named result, so the caller sees an ordinary error instead of a panic.

solid answer

~40 s

The conversion has three parts. First, name the error result — `func Reconcile(k Key) (err error)` — because only a named result can be reassigned once the function has stopped running normally. Second, `defer` the recovery closure at the very top, before anything that can fail; a call to `recover` only has an effect inside a function the panicking goroutine deferred. Third, inside that closure write `if r := recover(); r != nil { err = fmt.Errorf("reconcile %v: recovered panic: %v", k, r) }`. The function then returns normally carrying that error. I keep the boundary narrow — one per exported entry point or per goroutine top, never sprinkled through internal helpers — and the message always says the error came from a panic, so nobody mistakes a bug for an expected failure.

code

go · 10 lines
go
type Key struct{ Namespace, Name string }

func Reconcile(k Key) (err error) {
	defer func() {
		if r := recover(); r != nil {
			err = fmt.Errorf("reconcile %v: recovered panic: %v", k, r)
		}
	}()
	return apply(k)
}

go deeper

for a junior

Be ready to write the pattern from memory: named error result, deferred closure at the top, recover checked for non-nil, error assigned. Know that calling recover straight in the function body does nothing.

for a middle

Explain why the result must be named and why the closure must be deferred: the assignment happens after the body has stopped running, and only an already-registered deferred call executes while the stack unwinds.

for a senior

Show where you put this boundary in a real service and how the error stays actionable — which unit of work failed, that it came from a panic, and enough detail for whoever reads it at 3am.

for a principal

Own the policy: which boundaries in the codebase may convert panics, what those errors must contain, and why recovery scattered through helpers hides bugs rather than containing them.

## What the boundary is for A panic in Go is not an exception you catch at a convenient place. It unwinds the goroutine's stack, running deferred functions as it goes, and if nothing stops it the whole **process** exits with a traceback. A recovery boundary is the one place where you deliberately stop that unwinding and turn the failure into a value the surrounding code already knows how to handle: an `error`. ## The three ingredients ### 1. A named result parameter ```go func Reconcile(k Key) (err error) { ... } ``` `err` here is a real variable that lives for the whole call, not just a type in the signature. A deferred closure can read and write it. If the signature were `func Reconcile(k Key) error`, the result would be anonymous — the closure has nothing to name, so it cannot change what the caller receives. This is the single most common mistake when writing the pattern for the first time: the recovery runs, the panic is stopped, and the function returns `nil` because nobody could assign to the result. ### 2. A deferred closure, registered before the risky work ```go defer func() { if r := recover(); r != nil { err = fmt.Errorf("reconcile %v: recovered panic: %v", k, r) } }() ``` Only functions that were already deferred when the panic happened get to run during unwinding. A `defer` placed after the call that panics is never registered, so it never runs. Put the recovery closure at the top of the function body. A related subtlety: deferred calls run last-in-first-out. If the function also defers cleanup that sets `err` (closing a file, committing or rolling back), registering the recovery closure **first** means it runs **last**, so it sees and can overwrite whatever the cleanup put in `err`. That ordering is usually what you want at a boundary. ### 3. Checking the recovered value `recover()` returns the value passed to `panic`, typed as `any`. It is `nil` when the goroutine is not panicking, so the `if r := recover(); r != nil` guard is what distinguishes a normal return from a recovered one. The value can be anything — a string, an `error`, a custom type — so format it with `%v` rather than assuming it is an `error`, or type-switch on it if you care. ## What the resulting error should say A converted panic is a bug, not an expected outcome, and the error should read that way. `fmt.Errorf("reconcile %v: recovered panic: %v", k, r)` tells a reader three things: which unit of work failed, that the failure was a panic rather than a returned error, and what the panic value was. If callers need to branch on it, define a small error type that carries the recovered value (and, for a real service, the stack captured at the same moment) instead of flattening everything into a string. ## Where the boundary belongs One boundary per *unit of work that must not take the process down*: the top of each goroutine you start, an exported entry point of a library that must not panic into a caller's code, the per-item step of a loop that processes many independent items. Not inside every helper. Blanket recovery deep in the call tree converts programming bugs into vague errors far from where they happened and lets the program continue on state that a half-finished function left behind. Equally, never recover into silence. `defer func() { recover() }()` compiles, stops the panic, and returns the zero values — a function that reports success after a crash is worse than one that crashes. ## Worked shape ```go func Reconcile(k Key) (err error) { defer func() { if r := recover(); r != nil { err = fmt.Errorf("reconcile %v: recovered panic: %v", k, r) } }() return apply(k) // may panic deep inside } ``` If `apply` panics, the deferred closure runs while the stack unwinds, `recover` hands back the panic value, `err` is assigned, unwinding stops at this frame, and `Reconcile` returns to its caller like any other failing function. If `apply` returns an error normally, `recover` returns `nil`, the closure leaves `err` alone, and the caller gets exactly what `apply` produced. ## Summary Named result, defer first, check `recover()` for non-nil, assign a self-describing error. Place it where a failure must be contained, say in the message that it was a panic, and resist the urge to install one everywhere.

  • Why does the recovery closure have to be deferred before the code that panics, not after it?
    Only defers that were already registered when the panic started run during unwinding. A `defer` statement placed after the failing call is never reached, so it never registers, and the panic continues past the frame. Registering at the top of the function body is the habit that makes the boundary reliable.
  • If the function also defers cleanup that sets err, does the ordering between the two defers matter?
    Yes. Deferred calls run last-in-first-out, so the closure you defer first runs last. Registering the recovery closure first lets it observe and overwrite whatever the cleanup assigned to the named error, which is normally what you want: a recovered panic should not be masked by a benign close error.
  • What is wrong with `defer func() { recover() }()` at the top of a function?
    It stops the panic and discards it. The function returns its zero values, so the caller is told the work succeeded when it actually crashed mid-way. A recovery boundary must always turn the panic into a reported failure — an assigned error, a status update, or at minimum a logged record — never into silence.

saying these in an interview costs you the question

  • Calling recover in the function body instead of inside a deferred closure
  • Using an anonymous error result and expecting the assignment to reach the caller
  • Deferring the recovery closure after the code that panics
  • Recovering and returning nil, reporting success after a crash
  • Wrapping every internal helper in its own recovery boundary
open as a page

In Go, when should a function panic instead of returning an error?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Panic only for bugs the program cannot continue past: impossible internal states, broken invariants, or setup at package init. Anything a caller could reasonably hit at runtime, including bad input, missing files and I/O failure, returns an error.

open as a page

What happens when two goroutines write to the same Go map without a lock?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Go's built-in map is not safe for concurrent writes. The runtime notices the overlapping write and aborts the whole process with 'fatal error: concurrent map writes'. It is not a panic and the program does not continue past it.

open as a page

What does Go's built-in recover() do, and where must it be called to stop a panic?

level: juniorimportance: must knowfreq 80%

basics

~10 s

recover() stops a panicking goroutine from dying and returns the value that was passed to panic. It works only when called directly inside a function that was deferred; called anywhere else it returns nil.

open as a page

Where must the deferred recover live when a controller runs each reconcile in its own goroutine?

level: middleimportance: must knowfreq 60%

basics

~20 s

Inside the goroutine's own function, at the top, not in the code that started it. A recovery only takes effect in a function deferred by the goroutine that panicked, and a panic that escapes any goroutine terminates the whole process.

open as a page

What does regexp.MustCompile do differently from regexp.Compile, and where should it be called?

level: middleimportance: should knowfreq 62%

basics

~20 s

regexp.Compile returns (*Regexp, error); regexp.MustCompile returns only the *Regexp and panics if the pattern is malformed. Use it where the pattern is a literal in your source, typically a package-level var or init, never on caller or config input.

open as a page

When a Go process dies with 'fatal error: concurrent map writes', does deferred cleanup run?

level: middleimportance: should knowfreq 46%

basics

~20 s

No. A fatal runtime error is not a panic: the runtime prints its message and the goroutine stacks and terminates the process immediately, without unwinding any stack. No deferred function anywhere in the program gets to run.

open as a page

What causes 'fatal error: stack overflow' in Go if goroutine stacks grow on demand?

level: middleimportance: should knowfreq 34%

basics

~20 s

Goroutine stacks grow by copying to a larger stack, but only up to a ceiling - 1 GB on 64-bit by default. Unbounded recursion reaches that ceiling and the runtime aborts the whole process with 'fatal error: stack overflow'.

open as a page

A deferred closure calls a helper, and the helper calls recover() — why does the panic keep unwinding?

level: middleimportance: should knowfreq 44%

basics

~20 s

recover() returns the panic value only when it is called directly by a deferred function. A helper invoked from that deferred function is one level too deep, so its recover() returns nil and the panic carries on unwinding.

open as a page

When a deferred function in an outer frame recovers a panic raised three calls deeper, which deferred calls have already run and where does execution continue?

level: middleimportance: should knowfreq 58%

basics

~20 s

Every deferred call in every frame between the panic and the recovery point has already run, innermost frame first. Once one of them recovers, unwinding stops and that function returns normally to its caller; the frames below it are gone.

open as a page

How do you make sure a panic never escapes your Go package's exported functions?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Know the operations that panic on caller-controlled data, guard them at the entry of every exported function and return errors instead, then prove it with adversarial tests over zero values and fuzzing. Document any panic you deliberately keep.

open as a page

Why doesn't Go's 'all goroutines are asleep - deadlock!' error appear when a live server deadlocks?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Go's deadlock report fires only when every goroutine is parked with no possible wakeup. A live server always has something wakeable - a goroutine in the network poller, a blocked syscall, a pending timer - so a real deadlock hangs silently instead.

open as a page

Your controller recovers every reconcile panic and keeps running; the on-call SRE wants it to crash instead. How do you decide?

level: principalimportance: should knowfreq 30%

basics

~20 s

Decide by what a panic can damage. If the failure is confined to one item and shared state is untouched, recover and fail that item. If shared in-memory state can be left half-updated, crashing and restarting clean is safer.

open as a page

Why capture debug.Stack() inside the deferred function that recovers, rather than logging the panic later?

level: seniorimportance: nice to knowfreq 35%

basics

~20 s

The value returned by recover carries no stack. While the deferred function is running, the goroutine's stack still contains the frames that panicked, so debug.Stack captures the crash site. After that function returns, only the boundary's own callers remain.

open as a page

How do you tell a Go 'fatal error: out of memory' from the kernel OOM-killing the process?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Look for output. Go's abort writes 'fatal error: out of memory' plus goroutine stacks to stderr and exits with status 2. A kernel OOM kill sends SIGKILL, so there is no Go output and the container reports exit code 137.

open as a page

A Go test harness defers a recover around each case so one panic cannot kill the run — what does recovering fail to undo?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Recovering stops the unwinding and undoes nothing else. Statements the panicking code never reached still have not run, so anything left half-done stays half-done unless a deferred call cleaned it up, and shared state carries the damage into later cases.

open as a page

Your team owns a Go library other teams import. How do you decide where panics are allowed in it?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Set one strict default: no exported function panics on caller-controlled data. Allow a short, closed list of exceptions, such as package-level Must calls on source literals. Make the reviewer own exceptions, and back the rule with adversarial tests.

open as a page