Why capture debug.Stack() inside the deferred function that recovers, rather than logging the panic later?
answer
- the recovered value carries no location
- deferred calls run during unwinding
- the crash frames are still there, briefly
- one layer up is the wrong stack
- attach the bytes to the error you build
basics
~20 sThe value returned by recover carries no stack. While the deferred function is running, the goroutine's stack still contains the frames that panicked, so debug.Stack captures the crash site. After that function returns, only the boundary's own callers remain.
solid answer
~50 s`recover` hands you the value passed to `panic` and nothing else — no location, no frames. The only moment the crash site is still on the stack is while the deferred function is running, because deferred calls execute during unwinding, so `debug.Stack()` called there returns a traceback that runs from the handler down through `runtime.gopanic` into the function that actually panicked. Log the same panic value one layer up, after the boundary returned an error, and the traceback you print describes the caller instead — the interesting frames are gone. So the handler captures the bytes immediately and attaches them to the failure: a small error type holding the recovered value plus the captured stack, logged once at the boundary alongside the key of the item that failed. What you never do is put that stack in a message that leaves your process.
code
go · 15 linestype panicError struct {
value any
stack []byte
}
func (e *panicError) Error() string { return fmt.Sprintf("recovered panic: %v", e.value) }
func Reconcile(k Key) (err error) {
defer func() {
if r := recover(); r != nil {
err = &panicError{value: r, stack: debug.Stack()}
}
}()
return apply(k)
}go deeper
Know the fact behind the pattern: a panic value is just a value, with no location attached, so something has to record where the crash happened while it is still knowable.
Explain the timing — deferred calls run during unwinding, so the panicking frames are still present — and name runtime/debug.Stack as what returns that traceback for the current goroutine.
Show the production shape: an error type carrying the value and the captured bytes, logged once at the boundary with the identity of the failed item, never returned outside the process.
Decide who owns capture when boundaries nest, what the retention and redaction rules are for tracebacks in logs, and how a recovered-panic failure is classified differently from an expected error by retry and alerting policy.
## The problem: a panic value is not an exception object In languages with exceptions, the thrown object usually carries its own stack trace, captured at construction. Go's `panic` does not work that way. `recover()` returns exactly the value that was passed to `panic` — often a plain string or a small error — with no location information attached. If the boundary converts that value into an error and returns, the only thing anybody downstream ever learns is *what* the panic said, never *where* it came from. ## The narrow window Deferred functions run **while the goroutine is unwinding**, before the panicking frames are gone. That means a traceback taken inside the deferred handler still contains the whole chain: the handler itself, the runtime's panic machinery (`runtime.gopanic`), and beneath it the function that panicked and everyone who called it. `runtime/debug.Stack()` returns that traceback for the calling goroutine as a `[]byte`. Called in the handler, it is the crash site. Called anywhere after the recovering function has returned, it is a completely different and useless stack — the boundary's caller, the reconcile loop, the scheduler entry point. This is why "we log it at the top level where all errors are logged" quietly destroys the evidence. Centralised error logging is right for ordinary errors, which are created where they occur and can be wrapped with context on the way up. A recovered panic must be *decorated at the boundary* because the boundary is the last place the information exists. ## Attaching the capture to the failure The cleanest shape is a small error type that carries both halves: ```go type panicError struct { value any stack []byte } func (e *panicError) Error() string { return fmt.Sprintf("recovered panic: %v", e.value) } ``` The handler builds it with `&panicError{value: r, stack: debug.Stack()}` and assigns it to the named error result. Now the failure is an ordinary `error` that flows through the code like any other, and code that cares can type-assert to reach the stack. Two properties fall out of that design: - **Callers can tell a bug from an expected failure.** A type assertion (or an `errors.As` against `*panicError`) distinguishes "this item is invalid" from "this code has a bug", which is exactly the distinction a retry policy needs. - **The stack travels with the item's identity.** Log the error together with the key of the object being reconciled and you have the two facts an on-call engineer needs at once: which input triggered it, and which line blew up. ## Cost and discipline `debug.Stack()` formats the traceback into a fresh byte slice; it is not free, and it is not something to call on a hot path. At a panic boundary that is irrelevant — a panic is rare by construction, and if it is not rare you have a bigger problem. What does need discipline: - **Capture once.** If an outer boundary also captures, you get the same crash twice with different, misleading stacks. Decide which layer owns the capture. - **Log it once, at the boundary.** Do not attach a multi-kilobyte stack to an error that will be wrapped and re-logged by every layer above. - **Never return it to an external caller.** A stack names your packages, file paths and internal structure. Log it; return a generic failure to whoever is outside the process. - **Keep it out of metrics labels.** A stack as a label value will blow up whatever aggregates it. ## A note on scope `debug.Stack()` is the *calling goroutine's* stack, which is precisely what you want here: the goroutine that panicked. It tells you nothing about what the other goroutines were doing at that moment. If the panic is a symptom of a wider problem — a deadlock, a leak, a shared structure being mutated concurrently — a single-goroutine capture will not show it, and you are into whole-process dump territory instead. ## Summary The recovered value has no stack; the stack only exists while the deferred handler runs. Capture it there, attach it to the error you construct, log it once with the identity of the failed work, and keep it inside the process.
- What would you actually see if you captured the traceback one frame up, in the caller that received the error?The caller's own stack: the reconcile loop, its caller, and whatever runtime entry point started the goroutine. The panicking function and everything it called have already been unwound, so the capture points at code that did nothing wrong. That is why the boundary, not the logging layer, has to do it.
- Is it safe to include the captured stack in the error message returned across a process boundary?No. A traceback exposes package layout, file paths and internal structure, and it is far too large for a response or a metrics label. Log it once inside the process, keyed to the failed item, and return a generic failure outward — the detail belongs in your logs, not in someone else's client.
- Why not capture a stack for every error, not just recovered panics?Ordinary errors are created at the point of failure and gain context by wrapping as they travel, so the call chain is already recoverable from the message. Capturing and formatting a traceback per error costs real work on paths that are not rare, and buries the genuine crashes in noise.
saying these in an interview costs you the question
- Assuming the value returned by recover already contains a stack trace
- Logging the recovered panic in a central handler one or more layers up
- Capturing the stack before the risky work instead of in the handler
- Returning the captured traceback to an external caller
- Capturing at several nested boundaries and reporting the same crash twice