skip to content

Why can go test -race pass on a package whose code contains a real data race?

level: middleimportance: must knowfreq 58%

answer

  1. it judges runs, not programs
  2. both accesses must actually execute
  3. error branches and thresholds hide goroutines
  4. assembly and cgo carry no instrumentation
  5. coverage of concurrent paths is the constraint

basics

~20 s

Go's race detector is a runtime tool: it only judges memory accesses that actually execute. If no test drives the racy path, or the code is assembly or C reached through cgo, the conflicting accesses never reach the detector and the run is green.

solid answer

~50 s

`-race` instruments the build, but the judgement happens at runtime on accesses that really occur. Both halves of a race have to execute in the same run for it to be reported, so a race that lives on an error branch, a retry, a rarely enabled feature flag, or a path no test covers is simply never observed. The same applies to code the Go compiler does not instrument, such as hand-written assembly and C reached through cgo. It is worth being precise about what the detector does *not* need: it does not need the unlucky interleaving. It compares happens-before ordering, so two unsynchronized accesses seconds apart are still flagged. The binding constraint is coverage of concurrent paths, not timing luck — which is why a green `go test -race` means "nothing raced among what ran", not "this package is race-free".

code

go · 15 lines
go
type collector struct {
	buf []byte
}

func (c *collector) add(b []byte) {
	c.buf = append(c.buf, b...)
	if len(c.buf) > 1<<20 {
		go c.flush() // reads c.buf with no synchronization
	}
}

func (c *collector) flush() {
	_ = len(c.buf)
	c.buf = c.buf[:0]
}

go deeper

for a junior

Remember the one-line version: it only sees what actually runs, so passing tests do not prove the code is race-free. Being able to name an uncovered error branch as an example is enough here.

for a middle

Explain the mechanics both ways: which executions are missing when a race goes unreported, and why the detector still does not need an unlucky interleaving, because it compares synchronization ordering rather than timing.

for a senior

Show the practical conclusion: invest in tests that actually start the concurrent composition, run integration tests under -race too, and treat a green job as evidence about covered paths rather than a clearance for the package.

for a principal

Be able to state the residual risk the team is accepting in writing — which concurrent paths no job exercises under -race — and to defend spending test budget on covering them rather than on repeating what is already covered.

## The detector judges executions, not programs `go test -race` produces an instrumented binary, but the instrumentation is only a hook. The actual decision — these two accesses to this memory are not ordered by any synchronization, and one is a write — is made at runtime, about accesses that really happened. Code that does not run in that process is not analysed at all. This is the single most important property of the tool and the source of most misplaced confidence in it. So the question "why did it pass?" almost always resolves to "which of the two conflicting accesses never happened?" ## The usual reasons both accesses did not execute **The racy branch was never taken.** The classic shape is a goroutine started only on an unusual path: an error return, a retry, an overflow, a size threshold, a feature flag that is off in tests. The happy path is covered thoroughly; the branch that spawns the concurrent writer never runs. **The package has no test that starts concurrency at all.** Unit tests frequently drive a type through a single goroutine. Every access is then trivially ordered, because there is only one goroutine, and the detector has nothing to say. **The concurrent use is assembled somewhere else.** A type is race-free in its own package's tests and is shared across goroutines only by a caller. Nothing about `go test -race ./...` forces the composition that races to be exercised. **The code is not instrumented.** Hand-written assembly, and C called through cgo, are not compiled by the Go compiler and carry no instrumentation, so accesses made there are invisible to the detector even in a `-race` build. **Bounded access history.** For each word of memory the runtime keeps a bounded record of recent accesses. A variable hammered by many goroutines can have an older conflicting access evicted before the conflicting one arrives, so a real race can go unreported even when both accesses ran. This is a much rarer cause than the coverage reasons above, but it is why repeated runs sometimes report a race the first run missed. ## What the detector does *not* need A persistent misconception is that you must be lucky — that the two goroutines have to collide in a narrow timing window before the detector can notice. That is false, and it changes how you use the tool. The detector reasons about ordering, not about the clock. It asks whether some chain of synchronization operations — a channel send received elsewhere, a mutex unlocked and relocked, a `sync.WaitGroup.Done` observed by a `Wait`, an atomic operation — orders one access before the other. If no such chain exists, the accesses race, whether they happened microseconds or minutes apart. So you do not need to reproduce the corruption; you need only to make both accesses occur. That is why the payoff from `-race` is concentrated in *coverage of concurrent code paths*. Running the same single-goroutine test a thousand times under `-race` buys nothing. Running one test that actually starts the second goroutine buys everything. ## What a green run licenses you to say Exactly this: no unordered conflicting accesses occurred among the memory accesses these tests executed. It is a strong statement about the covered paths and no statement at all about the rest. Treat it as evidence, not proof — and when a race is suspected in production, the useful question is not "did CI pass?" but "does any test actually execute both sides of this sharing?" ## Turning that into practice The repair is coverage-shaped, not flag-shaped. Write a test that drives the concurrent composition rather than the type in isolation. Force the error and threshold branches that start goroutines, rather than only the happy path. Run integration-style tests under `-race` too, not just unit tests, because that is where composed concurrency lives. And keep `-race` failures blocking: a race that is observed once and merged past is worse than one never seen, because the team now believes the job is noise.

  • Does running the tests under -race with -count=100 make up for the gap?
    Only for the narrow case of history eviction or a path whose execution is itself probabilistic. Repetition does not execute code no test reaches, and the detector does not need repeated attempts to notice unordered accesses. Coverage of the concurrent path is what buys detection; repetition mostly buys runtime.
  • If -race only sees what runs, why is it still worth the cost?
    Because for the paths it does see it is close to definitive: it does not depend on hitting a bad interleaving, so a single execution of both accesses is enough. That turns a class of bug normally found once a month in production into a deterministic test failure.
  • How would you find races on paths your tests never take?
    Not with `-race` alone. You extend coverage so the path runs — force the error and threshold branches, run integration tests under `-race`, and exercise the composition rather than the type in isolation. Review and linting cover some of the remainder, but they answer a different question than a runtime detector.

It is a security camera, not a blueprint review. It records what walked through the door; it has no opinion about the corridors nobody used that night.

saying these in an interview costs you the question

  • Says a green -race run proves the package is race-free
  • Thinks the goroutines must collide in a timing window
  • Believes -count repetition substitutes for coverage
  • Assumes cgo and assembly are instrumented too
  • Treats -race as a static check on the whole tree