Why does a controller's shared status map crash the process with `fatal error: concurrent map writes` despite a deferred recover?
answer
- the runtime catches this one itself
- it is not a panic
- no deferred function ever runs
- the table may already be inconsistent
- encapsulate the map with its lock
basics
~20 sWriting one Go map from two goroutines is a data race the runtime detects itself. It reports a fatal error, not a panic, so no deferred function or recover runs. Guard the map with a mutex or one owner goroutine.
solid answer
~50 sGo's map implementation sets an internal flag while a write is in progress; if another goroutine enters a write, or a read sees that flag, the runtime aborts the program. It does so with a fatal error, not a panic, because the map's internal structure may already be inconsistent — there is nothing safe to unwind into, so deferred functions do not run and `recover` cannot catch it. The detection is **best effort**: a race that slips past the flag corrupts the map silently, so no crash is not evidence of safety. The fix is structural: put the map and a `sync.Mutex` in a struct with unexported fields so every access locks, or confine the map to one owner goroutine that others send updates to. A weekly crash means a narrow interleaving, not a rare bug.
code
go · 26 linestype resourceID struct {
namespace string
name string
}
type tracker struct {
mu sync.Mutex
status map[resourceID]string
}
func newTracker() *tracker {
return &tracker{status: make(map[resourceID]string)}
}
func (t *tracker) set(id resourceID, s string) {
t.mu.Lock()
defer t.mu.Unlock()
t.status[id] = s
}
func (t *tracker) get(id resourceID) (string, bool) {
t.mu.Lock()
defer t.mu.Unlock()
s, ok := t.status[id]
return s, ok
}go deeper
Remember that a Go map is not safe for concurrent use, that writing one from two goroutines aborts the program, and that the standard fix is a mutex held around every access, reads included.
Explain the difference between a panic and a fatal runtime error, and why this one runs no deferred functions. Be able to say why distinct keys and pre-sizing do not help.
Walk the incident: read the crash to find the two code paths that both touch the map, explain why a weekly crash means an everyday bug, and argue that detection is best effort so silence proves nothing. Then propose encapsulation or single-goroutine ownership rather than a defensive recover.
Own the pattern across services: which shared state is allowed at all, whether reconcilers keep state behind an owner goroutine by default, and what the review and on-call posture is for a class of bug where a clean run is not evidence.
## What the runtime is telling you A Go map is a hash table with internal bookkeeping — buckets, an element count, and, while it grows, a partially migrated old table. A write can be in the middle of relocating entries. If a second goroutine writes at the same moment, the two writers can leave the structure permanently inconsistent: entries lost, a bucket chain pointing at itself, a length that does not match the contents. Rather than let that happen silently, the map implementation sets a flag for the duration of a write. If another goroutine starts a write while that flag is set, the runtime aborts with `fatal error: concurrent map writes`. The read path checks the same flag, which produces the sibling message `fatal error: concurrent map read and map write`. ## Why `recover` cannot save you This is the part that surprises people, and it is the crux of the question. Go has two distinct abort paths: - A **panic** unwinds the goroutine's stack, running deferred functions on the way, and a `recover` inside one of those deferred functions can stop it. - A **fatal error** (an internal runtime *throw*) does not unwind anything. No deferred function runs, `recover` is never reached, the runtime dumps goroutine stacks and the process exits. Concurrent map access takes the second path, on purpose. The runtime has just detected that one of its own invariants is broken; continuing to execute application code over a possibly corrupt map would turn a detectable bug into silent wrong answers. So the well-meaning `defer func() { if r := recover(); r != nil { log.Println(r) } }()` at the top of the reconcile loop is irrelevant — there is nothing to recover. A second reason the defensive habit fails, worth stating even though it is not what happens here: `recover` only stops a panic in the goroutine that deferred it. A genuine panic in a worker goroutine cannot be caught by a handler in the goroutine that started it; it takes the whole program down too. ## Reading the crash, and what "once a week" means The crash output names the fatal error and then dumps the running goroutines with their stacks. The two you care about are the ones inside the map write — typically two different code paths that both touch the same map, which is often the real discovery: one is the reconcile loop updating status, and the other is a metrics endpoint or an event handler nobody thought of as a writer. Once a week is a scheduling statistic, not a rarity in the code. The unsynchronized write happens on every single update; what is rare is two of them overlapping in exactly the window the flag covers. Load, a new node with more cores, or a burst of events changes the frequency without changing the bug. This is also why the detection is described as best effort: overlaps that do not land inside the checked window are not reported at all, so a service that has never crashed may have been quietly corrupting its map for months. ## The fix The wrong fixes come first, because they are what people try: - Wrapping writes in `recover` — impossible, as above. - Pre-sizing the map with `make(map[K]V, n)` so it never grows — growth is not the trigger; any two concurrent writes are. - Writing only distinct keys from each goroutine — different keys still share one table, one length and one growth state. There is no per-key safety. - Restarting the process on crash — the map may already have been corrupted in the runs that did *not* crash. The real fixes make unsynchronized access impossible to express: **Encapsulate the map with its lock.** Put both in a struct with unexported fields so no caller can reach the map without going through a method that locks: ```go type tracker struct { mu sync.Mutex status map[resourceID]string } func (t *tracker) set(id resourceID, s string) { t.mu.Lock() defer t.mu.Unlock() t.status[id] = s } ``` An exported map field is an invitation to write it without the lock; an unexported one plus methods is enforceable in review. If reads dominate writes, a `sync.RWMutex` lets readers run concurrently at the cost of more bookkeeping per operation. **Confine the map to one owner.** A single goroutine owns the map as a plain local and serves reads and writes it receives on a channel. Nothing is shared, so nothing can race, and the ownership is obvious to the next reader of the code. This suits a controller well, since a reconcile loop is already a serialized loop. **Shard if the lock becomes hot.** Only after measuring: several maps each with their own mutex, chosen by a hash of the key, so writers to different shards do not contend. Whichever you choose, the review rule is the one to state out loud: a map reachable from more than one goroutine has exactly one legal access path, and it is not the map itself.
- Is one goroutine reading the map while another writes it safe?No. That is a data race too, and the runtime reports it as `fatal error: concurrent map read and map write` when it catches it. Only concurrent reads with no writer at all are safe on a Go map. So the lock has to cover the read path as well — a common half-fix is locking writes and leaving lookups bare, which still crashes.
- The crash happens about once a week. Why not just let the supervisor restart the process?Because the crash is the good case. Detection is best effort: overlaps the runtime does not catch leave the map silently corrupted, so between crashes the controller may be serving wrong status. Restarting also loses in-flight work and hides a defect that will grow with traffic. The interleaving is rare; the unsynchronized access is on every update.
- What does putting the map and the mutex in a struct with unexported fields buy over a package-level map and a package-level mutex?Enforcement. With unexported fields the only way to touch the map is through methods that lock, so a new caller cannot forget — the compiler stops them at the package boundary. Two package-level variables rely on everyone remembering the convention, and the one path that forgets is exactly the one that crashes you at 3am.
saying these in an interview costs you the question
- Wrapping the map write in recover stops the crash
- It is a panic, so a deferred handler catches it
- No crash so far means the map is used safely
- Different goroutines writing different keys is safe
- Pre-sizing the map so it never grows fixes it
- Restarting the process is an acceptable mitigation
- Only writes need the lock, lookups can stay bare