When the Go runtime throws `fatal error: concurrent map writes`, why does no deferred function run?
answer
- two different failure paths, not one
- a panic unwinds; this one does not
- no unwinding means no deferred calls
- the map may already be inconsistent
- recover is for your bugs, not broken invariants
basics
~20 sBecause it is not a panic. A fatal runtime error skips stack unwinding entirely: the runtime freezes the other goroutines, prints the banner and every goroutine's stack, and exits. Deferred calls and recover never enter the picture.
solid answer
~50 sGo has two different failure paths. A panic is a value travelling up one goroutine's stack: the runtime walks that goroutine's deferred calls, runs each one, and a `recover` inside a directly deferred function stops the unwinding. A fatal error takes the other path — the runtime does not create a panic at all. It freezes every goroutine, prints `fatal error: <message>` followed by a stack dump of the whole program, and exits with status 2. Nothing is unwound, so nothing deferred runs: no `recover`, no cleanup, no flushing a buffered writer, no cleanup functions registered with `runtime.AddCleanup`. That is deliberate rather than an oversight. `concurrent map writes` means unsynchronised writers may have left the map structurally inconsistent, and the deadlock and out-of-memory throws are properties of the whole process; continuing to run user code after any of them would be unsound. The rule of thumb is that `recover` handles your program's mistakes, not the runtime's broken invariants.
code
go · 8 linesfunc writeNilMap() {
defer func() {
fmt.Println("recovered:", recover())
}()
var counts map[string]int
counts["rows"] = 1
}
// prints: recovered: assignment to entry in nil mapgo deeper
Remember the hard rule: some runtime failures are panics you can recover from, and some are fatal errors you cannot. Concurrent map writes are in the second group, and no deferred function runs for them.
Explain the mechanism, not just the rule. A panic unwinds one goroutine's stack and runs its defers; a fatal error skips unwinding entirely, freezes the program, dumps all stacks and exits, so there is no point at which deferred code could run.
Show the operational consequence: unflushed buffers and unclosed resources on crash, and detection that is best effort, so the same racy map can pass for weeks. Talk about the race detector in CI, ownership of shared state, and surviving crashes by restart rather than by rescue.
Own the policy angle. Decide where shared mutable maps are allowed at all, what must be durable before a crash rather than flushed in a defer, and whether concurrency-heavy packages carry a race-enabled test job as a release requirement.
## Two failure paths that look similar from the outside Both a crashed panic and a fatal error kill the process, print a wall of stack traces, and exit with status 2. Underneath they are completely different mechanisms, and the difference is exactly what an interviewer is probing. ### The panic path `panic(v)` — and the runtime-generated panics such as a nil pointer dereference, an index out of range, a failed type assertion, or `assignment to entry in nil map` — creates a panic record attached to **one goroutine**. The runtime then walks that goroutine's deferred calls in last-in-first-out order and runs each one. If a deferred function calls `recover` directly, the panic stops there, `recover` returns the panic value, and the function that deferred it returns normally. If no deferred function recovers, the goroutine's stack is exhausted, the runtime prints `panic: <value>`, dumps the goroutines, and exits 2. The important properties: it is per-goroutine, it runs your deferred code, and it is interceptable. ### The fatal path A fatal error is raised by the runtime itself for a condition it considers unrecoverable. It does **not** create a panic record and does **not** unwind. The runtime freezes the other goroutines so the dump is coherent, writes `fatal error: <message>` to standard error, prints a stack trace for every goroutine, and terminates the process. Because there is no unwinding, there is nowhere for deferred calls to run: - `defer f.Close()` does not run, so files are closed only by the operating system on exit. - `defer w.Flush()` does not run, so buffered output written through a `bufio.Writer` is lost. - `defer func() { recover() }()` does not run, so `recover` never gets a chance. - Cleanup callbacks registered with `runtime.AddCleanup` (or the older `runtime.SetFinalizer`) never fire. ## Why the runtime refuses to let you continue Each fatal condition is one the runtime cannot honestly hand back to your code: - **`concurrent map writes`** (and `concurrent map read and map write`) is detected by a best-effort flag the map sets while a write is in flight. By the time you see it, two goroutines have been mutating the same internal structure without synchronisation, so the map may already be inconsistent. Recovering would mean continuing to read a data structure whose invariants are broken. - **`all goroutines are asleep - deadlock!`** is a property of the whole program: there is no goroutine left that could run a handler. - **`out of memory`** means the runtime could not obtain memory from the operating system; running more Go code, including a deferred function, may need memory it cannot get. - **`stack overflow`** means a goroutine's stack grew past the maximum, so there is no room to push another frame — including the frame of a deferred call. The detection is also best effort, not a guarantee. Concurrent map access without synchronisation is a data race; the fatal error is a courtesy the runtime extends when it happens to notice, not a check you can rely on. The tool that systematically finds the underlying bug is a `-race` build. ## What to do instead of trying to catch it You cannot catch it, so the engineering answer is prevention plus supervision: - Protect a shared map with a `sync.Mutex` or `sync.RWMutex`, or use `sync.Map` for its specific access patterns, or hand ownership of the map to a single goroutine and communicate with it. - Run the race detector in CI on the tests that exercise concurrency; it finds the unsynchronised access even on runs where the fatal error does not trigger. - Accept that the process dies, and make that survivable: a supervisor restarts it, and anything that must be durable is committed before the crash rather than flushed in a defer. ## The line to remember `recover` is for errors your program made in one goroutine while the runtime is still healthy. A fatal error means the runtime itself has decided the program is no longer sound, and it does not negotiate.
- Is there any hook at all that runs before the process dies on a fatal error?No. There are no deferred calls, no `recover`, no cleanup functions from `runtime.AddCleanup` or `runtime.SetFinalizer`, and no exit hooks. The only output is the runtime's own banner and goroutine dump. Anything you need to survive a crash has to be durable before the crash, not flushed on the way out.
- If it cannot be recovered, how do you stop one bad request from taking down the whole service?You prevent the condition rather than catching it: guard the shared map with a mutex, use `sync.Map` where its access pattern fits, or confine the map to a single owning goroutine. Then run the race detector in CI so the unsynchronised access is caught on runs where the runtime happens not to notice, and run under a supervisor so a crash is a restart rather than an outage.
- Does it matter which goroutine triggered the fatal error?Only for reading the dump. The abort is process-wide: every goroutine is frozen and its stack printed, regardless of which one hit the condition. There is no per-goroutine containment the way there is with a panic that a deferred recover stops.
- Why is the concurrent-map detection described as best effort?The runtime sets a flag while a map write is in progress and complains if another operation sees it. That catches many interleavings but not all, so the same racy code can run cleanly for a long time and then abort in production. Treat the fatal error as a symptom and the race detector as the systematic check.
A panic is an evacuation: you walk out, closing doors behind you as you go. A fatal error is the building being condemned mid-shift — the power is cut and nobody finishes anything on the way out.
saying these in an interview costs you the question
- Wraps the code in recover and calls it handled
- Thinks defers still run before the process exits
- Claims sync.Map removes the need for any synchronisation reasoning
- Believes the runtime always detects concurrent map writes
- Treats the fatal error as a normal recoverable error value