A Go test harness defers a recover around each case so one panic cannot kill the run — what does recovering fail to undo?
answer
- stopping is not rolling back
- what never ran never runs later
- only deferred cleanup fired
- a lock taken without defer stays taken
- the value survives, the stack does not
basics
~20 sRecovering stops the unwinding and undoes nothing else. Statements the panicking code never reached still have not run, so anything left half-done stays half-done unless a deferred call cleaned it up, and shared state carries the damage into later cases.
solid answer
~50 s`recover()` is a brake, not a rollback. When a case panics on a nil pointer three calls deep, the only work that ran on the way out was the deferred calls in the unwound frames. A lock taken with a plain `Lock()` and released at the bottom of the body is still held, so the next case that touches it hangs; a shared fixture the case was halfway through mutating stays halfway mutated. The harness also learns little: `recover()` hands back only the panic value, and once the handler returns, the frames that would say where it happened are gone. So: recover per case, but treat a recovered panic as a hard failure of that case, give every case its own state, and record what you need at the moment you catch it — never let recovery mean carry on as if nothing happened.
code
go · 11 linesfunc withDefer(mu *sync.Mutex) {
mu.Lock()
defer mu.Unlock() // runs even when mayPanic panics
mayPanic()
}
func withoutDefer(mu *sync.Mutex) {
mu.Lock()
mayPanic()
mu.Unlock() // never reached on the panic path: the lock is held forever
}go deeper
Take away one rule: put cleanup in a defer right after you acquire the thing. That is what makes it run when a panic passes through, and nothing else in the function will.
Explain the difference between stopping the unwinding and undoing work: deferred calls in the unwound frames ran, everything else was skipped, and no state was restored.
Show that you would design for it — per-case isolated state, a recovered panic counted as a failure, and the panic value captured where it is caught rather than reconstructed later.
Own the standard for when recovery is allowed at all: it is a claim that the guarded region owned everything it touched, and it should be reviewed as such rather than adopted as a general safety net.
## What recovery actually buys, and what it does not A harness that runs many independent cases has a legitimate reason to recover: one bad case should not take the whole run down before the other cases have been tried. The wrapper is the familiar shape — a deferred closure around each case that calls `recover()` and records the result. That is the entire benefit: the process survives, and the run continues. It is worth being blunt about the rest, because the word "recover" invites the wrong expectation. Recovery does not: - undo assignments the panicking code already made; - run any statement the panicking code had not reached; - release anything that was not released by a deferred call; - restore the frames that were unwound; - tell you anything about the failure beyond the value the panic carried. ## The state left behind Work through what a nil-pointer panic in the middle of a case leaves behind. **Locks.** `mu.Lock()` followed later by a plain `mu.Unlock()` at the bottom of the function is a lock that survives the panic. The next case that acquires that mutex blocks forever, and the run reports a hang rather than the original panic. `defer mu.Unlock()` immediately after the acquire is what makes the panic path safe — which is the practical reason that idiom is worth insisting on. **Files, connections, temporary directories.** Same rule. If cleanup was deferred it ran during unwinding; if it was a statement at the end of the body it did not. **Shared fixtures.** A slice being grown, a map being populated, a package-level registry being reindexed — any of these can be observed mid-update by the next case if the harness hands out the same object to everybody. The panic did not corrupt the data; the abandoned half of the update did. **Counters and bookkeeping.** An increment whose matching decrement lived after the panicking call is now permanently skewed, and the number it feeds is quietly wrong for the rest of the run. ## The diagnostic cost The value `recover()` returns is just the argument the panic carried, an `any`. For a nil dereference that is a `runtime.Error` whose text is `runtime error: invalid memory address or nil pointer dereference` — which says what happened and nothing at all about where. The stack that would have answered that was being unwound as the handler ran, and once the handler returns it is gone for good. If the harness records only the value, every nil-pointer failure in the suite looks identical. The practical consequence is that whatever you want to know about the failure has to be captured inside the deferred function, at the moment of recovery, not afterwards. ## Designing the harness so recovery is honest Four rules cover it: 1. **A recovered panic fails its case.** Logging it and moving on turns a crash into a silently passing run, which is strictly worse than the crash. 2. **Give each case its own state.** Fresh fixtures per case are what make "one case died" a bounded statement instead of a hopeful one. Shared mutable state plus recovery equals cross-contamination you will chase for days. 3. **Capture what you need where you catch it.** The panic value, and whatever context identifies the case, recorded in the deferred function. 4. **Keep the boundary at the harness.** Recovery belongs at the place that owns the isolation decision. Sprinkling it deeper turns real bugs into shrugs. ## The judgment behind it The reason a panic kills the program by default is that a panic means the program reached a state its author believed impossible. Every recovery is a claim that you can carry on anyway, and that claim is only true when the region you recovered around owned everything it touched. A test harness usually can make that claim honestly for one case; the same wrapper applied to a long-lived process mutating shared state usually cannot. For someone new to a codebase this is the most useful thing to know about a `recover()` they encounter: read it as a statement about what is isolated, and check whether that isolation is real.
- Your harness recovers every case and the run finishes green. What is wrong with that?A recovered panic that is only logged is a failure the run reported as success. Recovery should decide *where* a failure is contained, never *whether* it counts. The wrapper must mark the case failed, and the run must be non-zero because of it.
- Why does every failure look the same when the harness stores only the recovered value?`recover()` returns the panic's argument and nothing else. For a nil dereference that argument is a fixed `runtime.Error` message with no location in it, and the frames that carried the location were unwound before the handler finished. Anything you need must be captured at the moment of recovery.
- When is recovering around a region genuinely safe?When the region owns everything it mutated, so abandoning it half-done cannot be observed by anything that runs afterwards. One case with its own fixtures usually qualifies. The same wrapper around code that updates process-wide state does not, however tidy it looks.
Catching a falling plate stops the crash. It does not put the food back on the plate, and it does not tell you who knocked it off the table.
saying these in an interview costs you the question
- Treats recover as a rollback that restores state
- Recovers, logs, and lets the run report success
- Assumes cleanup ran even where it was not deferred
- Shares one mutable fixture across recovered cases
- Expects the panic's location to be in the recovered value