skip to content

When a Go process dies with 'fatal error: concurrent map writes', does deferred cleanup run?

level: middleimportance: should knowfreq 46%

answer

  1. two different ways a program dies
  2. one unwinds a stack, the other does not
  3. the process ends with nothing unwound
  4. buffered bytes never reach the file

basics

~20 s

No. A fatal runtime error is not a panic: the runtime prints its message and the goroutine stacks and terminates the process immediately, without unwinding any stack. No deferred function anywhere in the program gets to run.

solid answer

~50 s

Go has two distinct failure paths and they behave differently. A panic unwinds the panicking goroutine's stack, running each deferred call on the way out. A fatal runtime error - concurrent map writes, all goroutines are asleep, stack overflow, out of memory - is a runtime throw: it prints a diagnostic plus goroutine stacks and exits the process with status 2 straight away. Nothing is unwound, so a deferred `Close`, a deferred `Flush` and every other deferred call, in every goroutine, is skipped. The design consequence is what matters in a review: anything you cannot afford to lose must not live only in a deferred call. A daemon that accumulates counters in memory and writes them out in a deferred flush loses the whole window when the runtime aborts. Flush on an interval instead and let a supervisor restart the process.

code

go · 12 lines
go
func main() {
	f, err := os.Create("counters.txt")
	if err != nil {
		log.Fatal(err)
	}
	defer f.Close()

	w := bufio.NewWriter(f)
	defer w.Flush() // never runs on a fatal runtime error

	run(w) // if this aborts fatally, the file stays empty
}

go deeper

for a junior

Remember the rule as a pair: defer is reliable on a normal return and on a panic, and useless when the runtime aborts. Being able to say which cleanup was skipped after a crash is enough at this level.

for a middle

Be ready to contrast the two failure paths precisely - one unwinds a goroutine's stack running defers, the other terminates the process with no unwinding - and to name what is lost from userspace buffers as a result.

for a senior

An interviewer expects the design conclusion: durability cannot live on the exit path. Show how you would restructure a service so that a fatal abort costs a bounded window rather than everything since startup.

for a principal

Own the standard. Decide what the platform guarantees on abnormal exit, what supervisors and consumers must tolerate, and make crash-safety a review criterion rather than something each service invents for itself.

## Two failure paths, not one When a Go program stops badly, it stopped in one of two very different ways. **A panic** is a language mechanism. It starts in one goroutine and unwinds that goroutine's stack, running each deferred call it passes, in last-in-first-out order. That is why `defer f.Close()` and `defer mu.Unlock()` are reliable in normal code: they survive an early return and they survive a panic. **A fatal runtime error** is the runtime giving up. It prints a line starting `fatal error:`, dumps goroutine stacks to standard error, and exits with status 2. It does not unwind anything, in any goroutine. The program's own code does not execute another instruction. The fatal class includes: - `fatal error: concurrent map writes` (and the read-and-write and iteration variants) - `fatal error: all goroutines are asleep - deadlock!` - `fatal error: stack overflow` - `fatal error: out of memory` ## Why the runtime refuses to unwind Each of these conditions means the runtime cannot safely keep going. A map whose internal structure was being mutated by two goroutines may already be inconsistent. A goroutine that hit the stack ceiling has nowhere to put another frame - and running a deferred call requires a frame. An allocation failure means the next thing your cleanup code allocates will fail too. Running arbitrary user code on the way out would risk deepening the corruption, hanging, or failing again in a way that destroys the diagnostic the runtime is trying to print. So it prints what it knows and stops. ## What this costs you Everything the process was holding in userspace vanishes: - A `bufio.Writer` with unflushed bytes writes nothing more. Only bytes already handed to the kernel by a successful `Write` survive. - An `encoding/json` encoder mid-document leaves a truncated file. - Temporary files scheduled for removal by `defer os.Remove(name)` stay on disk. - In-flight HTTP responses are cut off at the socket. - Aggregated in-memory state - counters, batches, a queue of pending work - is gone. What *does* still happen is whatever the operating system does for any dying process: file descriptors and sockets are closed, memory is reclaimed, and locks held only inside this process become irrelevant because the process is gone. ## The design rule this implies Deferred cleanup is for the normal path and for panics. It is not a durability mechanism. If a piece of state must survive a crash, the crash-safe version of it has to exist before the crash: - **Flush on an interval, not at exit.** A metrics daemon that writes every few seconds loses at most a few seconds. One that flushes in a deferred call loses everything since process start. - **Make partial output detectable.** Append-only records with a length or a checksum let the next reader discard a torn tail instead of misreading it. - **Push durability outward.** A supervisor that restarts the process, and a consumer that tolerates a gap or a duplicate, is far more robust than any in-process shutdown path. - **Treat graceful shutdown as an optimisation.** If correctness depends on your shutdown handler running, the system is one fatal error away from being wrong. The same reasoning applies to `os.Exit`, which also terminates without running defers, and to the process simply being killed by the operating system. A crash-tolerant design does not care which of the three happened. ## Reading the postmortem Because nothing unwinds, the goroutine dump is a snapshot of the program exactly as it was. That is genuinely useful: every goroutine's stack is where it was parked or running at the instant of death, so you can see the two goroutines in the same helper, or the handler that was still mid-request. `GOTRACEBACK` controls how much of that dump you get - `none` suppresses the goroutine stacks entirely, `all` includes every user goroutine, `system` adds runtime frames and runtime-created goroutines, and `crash` makes the process die in an operating-system-specific way so a core dump can be collected. It is read by the process that crashes, so it shapes the next crash, not the one already in your hands. ## The trap in review The common defect is a service whose only write of accumulated state is in a deferred call at the bottom of `main`, or a shutdown handler wired to a signal. It works in every test, because tests end normally. It works in every graceful redeploy. It fails precisely on the day something goes wrong - which is the day the data mattered most.

  • What cleanup does still happen when the Go runtime aborts the process?
    Only what the operating system does for any dead process: descriptors and sockets are closed, and memory is reclaimed. Bytes already accepted by a successful write to the kernel are safe. Anything still in a userspace buffer - a bufio.Writer, a partially encoded document, accumulated in-memory state - is lost.
  • How would you make a counter-aggregating daemon survive this kind of abort?
    Move durability off the exit path. Flush on a short timer so the worst case is a few seconds of loss, write append-only records that a reader can truncate safely at a torn tail, and let a supervisor restart the process. Graceful shutdown then becomes an optimisation rather than the only correct path.
  • Why does the runtime refuse to unwind on a fatal error instead of raising a panic?
    Because the conditions that cause one mean the runtime cannot safely execute more user code - a possibly corrupted map, no stack room left for another frame, or an allocator that just failed. Running cleanup could deepen the damage or destroy the diagnostic, so the runtime prints what it has and stops.

saying these in an interview costs you the question

  • Believes defer runs no matter how the process ends
  • Expects a deferred Flush to save buffered output here
  • Confuses a runtime abort with a panic that unwinds
  • Thinks only the offending goroutine is stopped
  • Relies on a shutdown handler for data that must not be lost