skip to content

Races, Visibility and Detection

What Go guarantees about one goroutine seeing another's writes, the two ways concurrent code breaks anyway - races and deadlocks - and the tooling that catches both. A racy program passes its tests.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

explore

questions

20

Two goroutines run `counter++` on the same shared variable with no lock — why is that a data race?

level: juniorimportance: must knowfreq 82%

answer

  1. one source line, several machine steps
  2. both goroutines can load the same value
  3. one writer plus no ordering
  4. lost update is the mild outcome
  5. lock it, make it atomic, or own it

basics

~20 s

counter++ is a load, an add and a store, not one step. Two goroutines can load the same old value and store the same result, losing an update. Unsynchronized concurrent access with at least one writer is a data race.

solid answer

~50 s

A data race is two goroutines touching the same memory location at the same time, with at least one of them writing, and nothing ordering the two accesses. `counter++` looks atomic in the source but compiles to a read, an add and a write, so the runtime can switch goroutines between the read and the write and one increment silently overwrites the other. The damage is not limited to a wrong count: Go's memory model leaves a racy program's behaviour undefined, and the compiler is free to keep the variable in a register or reorder the accesses, so a racy flag can be missed entirely. The fixes are to guard the variable with a `sync.Mutex`, to make the whole update one atomic operation with `atomic.Int64.Add`, or to give one goroutine sole ownership of the counter and have the others send it deltas over a channel.

code

go · 13 lines
go
var counter int
done := make(chan struct{})

for i := 0; i < 100; i++ {
	go func() {
		counter++ // load, add, store: three steps, not one
		done <- struct{}{}
	}()
}
for i := 0; i < 100; i++ {
	<-done
}
fmt.Println(counter) // usually less than 100, and varies per run

go deeper

for a junior

Be ready to define a data race in one sentence and to say why counter++ is three steps rather than one. Naming a mutex, an atomic and single-goroutine ownership as the three fixes is enough at this level.

for a middle

Explain the load-add-store interleaving concretely and say why the outcome is undefined rather than merely wrong by a few. Justify choosing an atomic for a single word versus a mutex when several fields share an invariant.

for a senior

Show that you spot shared mutable state during review before it ships, and that you treat a passing test run as no evidence of safety. Be able to argue for redesigning toward ownership instead of scattering locks over a growing struct.

for a principal

Own the discipline rather than the incident: which state is allowed to be shared at all, what the codebase's default concurrency pattern is, and what you require in review so the next such bug never reaches production.

## What a data race is A **data race** in Go is precisely defined: two goroutines access the same memory location concurrently, at least one of the accesses is a **write**, and nothing in the program orders one access before the other. All three conditions are required. Two goroutines reading the same variable is not a race. A goroutine writing a variable that no one else touches is not a race. A write plus any other access, unordered, is. ## Why `counter++` is not one operation The source line is one token, but the machine does three things: 1. **Load** the current value of `counter` from memory into a register. 2. **Add** one to the register. 3. **Store** the register back into `counter`. The Go scheduler can switch goroutines between any two of those steps, and on a multi-core machine two goroutines are genuinely running at the same instant. The classic interleaving: | step | goroutine A | goroutine B | counter | |---|---|---|---| | 1 | load 7 | | 7 | | 2 | | load 7 | 7 | | 3 | store 8 | | 8 | | 4 | | store 8 | 8 | Two increments happened and the value went up by one. That is a **lost update**. Run a hundred goroutines each incrementing a shared counter and the final value is reliably *less* than a hundred, and different on every run. ## It is worse than an off-by-a-few count The tempting response is "so the number is a bit low, who cares". That is the misconception this question exists to break. The Go memory model does not merely say the result is unpredictable; it says a program with a data race is **not defined** by the model at all, beyond the guarantee that a read of a location no larger than a machine word observes *some* value that was actually written there (Go rules out out-of-thin-air values, unlike C++). Everything else is on the table: - The compiler may prove, on the assumption that no other goroutine touches the variable, that a loop's reads can be **hoisted into a register**. A `for !done {}` loop reading a racily-written `bool` can then spin forever even after the flag is set. - The compiler and the CPU may **reorder** the racy accesses relative to surrounding ones, so a value published just before a flag may not be visible when the flag is. - For values wider than a machine word the read can observe a **mixture** of two different writes — a value nobody ever stored. So "it is just a counter" is not a defence: the surrounding code can behave in ways that no reading of the source explains. ## The three ways to fix it **A mutex.** Wrap the read-modify-write so only one goroutine is inside at a time: ```go var mu sync.Mutex mu.Lock() counter++ mu.Unlock() ``` Use this whenever more than one variable must change together — a counter and a timestamp, a map and its length — because a mutex can cover a whole invariant while an atomic covers exactly one word. **An atomic.** `sync/atomic`'s typed values (Go 1.19 and later) make the whole update one indivisible operation: ```go var counter atomic.Int64 counter.Add(1) later := counter.Load() ``` This is the right tool for a single independent number — a request counter, a sequence number. It does **not** help if two fields must stay consistent with each other. **Confinement.** Give the state one owner. One goroutine holds the counter as an ordinary local variable and everyone else sends it increments on a channel; the value is never shared, so there is nothing to race on. This is the design Go's "share memory by communicating" slogan is pointing at, and it scales better in review than sprinkling locks. ## The pattern generalises A counter is only the smallest example. Every shared mutable thing has the same rule: a struct field written by a handler and read by a metrics goroutine, a slice appended to from several workers, a map written by more than one goroutine, a package-level variable initialised lazily by whoever gets there first. If more than one goroutine can touch it and at least one writes, it needs a lock, an atomic, or an owner. One last calibration point: `GOMAXPROCS=1` does not make any of this safe. Goroutines are preemptible, so the switch can still land between the load and the store; a single P removes parallelism, not interleaving.

  • Does setting GOMAXPROCS to 1, or using a narrower type like int32, make the increment safe?
    Neither. `GOMAXPROCS=1` removes parallelism but not interleaving — goroutines are preemptible, so a switch can still land between the load and the store. The width of the variable is irrelevant too: the problem is that a read-modify-write is three separate steps, not that the value is wide. Only synchronization — a mutex, an atomic operation, or single-goroutine ownership — removes the race.
  • When would you reach for atomic.Int64 and when for a sync.Mutex?
    Use an atomic when the shared state is exactly one independent word and every update is a single indivisible operation: a hit counter, a sequence number, a swapped-in pointer. Use a mutex as soon as two or more pieces of state must agree with each other — a counter and the timestamp of the last update, a slice and its index — because a mutex can span the whole invariant while an atomic protects one word at a time.
  • Both goroutines only ever write the same constant value to the variable. Is that still a data race?
    Yes. The definition is about unsynchronized concurrent access with at least one writer, not about whether the values differ. The compiler is entitled to optimize on the assumption that the program is race-free, so a race that looks harmless can still change how surrounding code is compiled and scheduled. Write it once before starting the goroutines, or guard it.

saying these in an interview costs you the question

  • Increment is a single instruction, so it is atomic
  • The worst that happens is a slightly low count
  • Setting GOMAXPROCS to 1 makes shared counters safe
  • Only writes need guarding, a read alongside a write is fine
  • Adding a short sleep between goroutines fixes it
  • It never reproduces on my laptop, so the code is correct
open as a page

Why does one goroutine running `ch := make(chan int); ch <- 1; v := <-ch` block forever on the send?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An unbuffered channel holds nothing, so a send blocks until another goroutine is ready to receive. The only receive here is the next line of the same goroutine, which cannot run until the send finishes.

open as a page

What does go test -race do, and what does its DATA RACE report contain?

level: juniorimportance: must knowfreq 72%

basics

~20 s

go test -race builds an instrumented test binary that watches memory accesses while the tests run. When two goroutines touch the same memory unsynchronized and at least one writes, it prints WARNING: DATA RACE with a stack for each access, and the test fails.

open as a page

In a Go test, what should replace time.Sleep when waiting for a goroutine to finish?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Let the goroutine signal on a channel and have the test receive from it, so the test unblocks exactly when the work is done. Add a select with a time.After case so a hang fails the test instead of stalling it.

open as a page

Why does sync.WaitGroup.Wait hang forever when a worker goroutine returns early on an error?

level: middleimportance: must knowfreq 66%

basics

~20 s

A WaitGroup is only a counter: Add raises it, Done lowers it, Wait blocks until it reaches zero. A worker that returns before reaching its Done leaves the counter above zero, so Wait never unblocks. Put Done in a defer.

open as a page

Which happens-before edges does an unbuffered channel operation create in Go?

level: middleimportance: must knowfreq 58%

basics

~20 s

Two, in opposite directions. A send is synchronized before the matching receive completes, so the receiver sees everything the sender wrote first. On an unbuffered channel the receive is also ordered before the send completes.

open as a page

Why can go test -race pass on a package whose code contains a real data race?

level: middleimportance: must knowfreq 58%

basics

~20 s

Go's race detector is a runtime tool: it only judges memory accesses that actually execute. If no test drives the racy path, or the code is assembly or C reached through cgo, the conflicting accesses never reach the detector and the run is green.

open as a page

Why does a controller's shared status map crash the process with `fatal error: concurrent map writes` despite a deferred recover?

level: seniorimportance: must knowfreq 66%

basics

~20 s

Writing one Go map from two goroutines is a data race the runtime detects itself. It reports a fatal error, not a panic, so no deferred function or recover runs. Guard the map with a mutex or one owner goroutine.

open as a page

What does Go's memory model guarantee about writes made before a `go` statement?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Everything the starting goroutine wrote before the go statement is guaranteed visible to the new goroutine when it begins running. The guarantee is one-way: the starter is not guaranteed to see anything the new goroutine writes.

open as a page

Several goroutines assign to an `err` variable declared in the enclosing function — why is that a race?

level: middleimportance: should knowfreq 56%

basics

~20 s

A Go closure captures the variable itself, not a copy of its value. Every goroutine writes the same storage, so concurrent assignments to that one err variable are an unsynchronized data race exactly like writes to a shared global.

open as a page

Does a goroutine finishing establish any happens-before edge in a Go program?

level: middleimportance: should knowfreq 42%

basics

~20 s

No. Go's memory model states that a goroutine's exit is not synchronized before any event. To observe what it wrote you need an explicit edge: a channel receive, a Wait unblocked by Done, or a mutex handoff.

open as a page

Why is calling t.Fatalf from a goroutine your Go test spawned unsafe?

level: middleimportance: should knowfreq 48%

basics

~20 s

Fatalf records the failure and then ends only the goroutine that called it, through runtime.Goexit. The test function keeps running with the broken state, and if it returns first the late call panics. From a spawned goroutine use Errorf and return, or send the error back.

open as a page

Why is holding a sync.Mutex while sending on a channel a deadlock risk in Go?

level: seniorimportance: should knowfreq 52%

basics

~20 s

A blocked channel send does not release the mutex; Go never drops a lock when a goroutine parks. If the consumer that would drain the channel must first take that same mutex, each side waits for the other and neither moves again.

open as a page

A go test -race failure names a write and a previous read on a shared []byte — how do you get from that report to the fix?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Read both stacks and the goroutine-creation frames to find who shares the backing array and where it was handed over. Then fix ownership: guard every access with the same lock, copy the bytes out, or hand the buffer off — locking only the writer leaves the race intact.

open as a page

A Go concurrency bug fails once in 200 CI runs. How do you build a test that reproduces it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Isolate the suspect component into its own test, drive it with replayed recorded traffic at a chosen concurrency, raise GOMAXPROCS, and repeat the case hundreds of times under go test -race. Then quantify: measure the failure rate per hundred runs before and after the fix.

open as a page

Running the whole suite under go test -race doubled CI time — how do you decide what keeps running with it?

level: principalimportance: should knowfreq 36%

basics

~20 s

Spend the budget where the detector can actually earn it: packages that start goroutines and share state, kept blocking on every pull request, with the full-suite race pass moved to merge or nightly. Then write down which paths nobody runs under -race any more.

open as a page

Two goroutines assign different slices to one shared slice variable; what can a racing reader see?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

A slice value is three words — pointer, length and capacity — stored one at a time. An unsynchronized reader can observe one assignment's pointer next to the other's length and index past the end of the real backing array.

open as a page

How do you pin a specific goroutine interleaving in a Go test without time.Sleep?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Give a test-only fake collaborator two unbuffered channels: it sends on one when it is entered and blocks receiving on the other. The test receives the first signal, so it knows the goroutine is stopped inside that call, does its second operation, then releases the fake.

open as a page

Two goroutines lock two account mutexes in opposite order and wedge — what do you capture from the live process?

level: seniorimportance: nice to knowfreq 42%

basics

~10 s

Capture full goroutine stacks before restarting: /debug/pprof/goroutine?debug=2 if net/http/pprof is registered, or GOTRACEBACK=all plus SIGQUIT. Two goroutines parked in sync.(*Mutex).Lock inside mirrored transfer calls are the cycle.

open as a page

After porting a Go service from amd64 to arm64, workers sometimes read a stale package-level config pointer. Why now?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

The publish was never synchronized, so no happens-before edge orders the write against the workers' reads. amd64's strong hardware ordering hid the bug; arm64 is weakly ordered and lets the stale value be observed. The race was always there.

open as a page