skip to content

Several Go subtests calling t.Parallel mutate one map the parent test created. How do you diagnose and fix it?

level: seniorimportance: should knowfreq 50%

answer

  1. the fixture was never copied
  2. closures capture, they do not clone
  3. build it under the race detector
  4. concurrent map writes can abort the binary

basics

~20 s

Subtests calling t.Parallel run concurrently after the parent returns, so a map their closures captured is shared, not copied. Confirm with go test -race, then give each subtest its own fixture or guard the shared one with a sync.Mutex.

solid answer

~50 s

The parent builds one map and every subtest closure captures the same reference; once the parent's body returns, the paused subtests resume together and write to it concurrently. That is an ordinary data race, and test code is not exempt. Run the package with `go test -race`: the report names two goroutines with stacks inside the subtest closure, both touching the same map. A concurrent map write can also abort the binary outright with a fatal runtime error instead of failing cleanly. The fix is to stop sharing -- build the mutable fixture inside each subtest so every case owns its own, and treat anything inherited from the parent as read-only. If the cases genuinely must accumulate into one structure, a `sync.Mutex` removes the race, but that usually means the aggregation belongs in the parent rather than in the parallel cases.

code

go · 12 lines
go
func TestShardTotals(t *testing.T) {
	totals := map[string]int{} // one map shared by all three closures
	for _, shard := range []string{"a", "b", "c"} {
		t.Run(shard, func(t *testing.T) {
			t.Parallel()
			totals[shard] = rowsIn(shard) // concurrent write to a shared map
		})
	}
	if len(totals) != 3 { // still empty: the subtests have not resumed yet
		t.Errorf("totals = %d shards, want 3", len(totals))
	}
}

go deeper

for a junior

Recall that a closure captures a variable rather than copying it, so a map created in the parent test is the same map inside every subtest. Concurrent writes to one map are not safe.

for a middle

Explain the two separate defects: the concurrent writes, and the parent's assertion running before any parallel subtest resumes. Be able to say what a race report showing two goroutines at one line means.

for a senior

Walk the whole diagnosis: reproduce under the race detector, read the two stacks, distinguish a genuine race from the runtime's concurrent-map-write abort, and pick the fix that removes the sharing rather than the one that locks around it.

for a principal

Own the standard for the suite: what a parallel case is allowed to touch, how parallelism is introduced to existing tests without importing intermittent failures, and when the honest answer is that a package's tests are too coupled to parallelise at all.

## The shape of the bug A batch job splits CSV input into shards, and the test drives one subtest per shard so the shards can be checked concurrently. The parent allocates a `map[string]int` of row totals, each subtest fills in its own key, and the parent means to assert the aggregate afterwards. Every one of those steps is wrong under `t.Parallel()`, and for two independent reasons. **Reason one: the closures share, they do not copy.** A subtest body is a closure. Capturing a map, a slice, a pointer, or a struct containing any of those captures the *same* underlying object; there is no per-subtest copy. Three parallel subtests writing to that map are three goroutines writing to one map with no synchronisation. **Reason two: the parent's aggregate assertion runs too early.** `t.Parallel()` pauses each subtest until the parent's function body returns. So the statement after the loop — the one that reads the totals — executes while every subtest is still parked. The map is empty there, and no amount of fixing the race changes that. ## Diagnosing it Build and run the package with the race detector: `go test -race ./...`. The detector instruments memory accesses and reports races that actually occur during the run. A report for this bug has a recognisable shape: - a **Write** stack ending inside the subtest closure, in a goroutine created by the testing runner; - a **Previous write** stack in a *different* goroutine, ending at the same line of the same closure; - the goroutine-creation frames pointing at the testing package, and often the allocation site in the parent test function. Two stacks that are the same source line in different goroutines is the signature of a shared fixture: it is not two different pieces of code disagreeing, it is one piece of code running several times over one object. Two practical notes. First, a concurrent map write may not even reach the detector: the map implementation itself detects concurrent writes and throws `fatal error: concurrent map writes`, which kills the whole test binary and takes unrelated tests down with it — an abrupt failure that people often misread as an infrastructure problem. Second, the race detector only sees what actually executed; a green run on a lightly loaded machine proves nothing about a path the scheduler did not interleave that time. ## Fixing it In order of preference: 1. **Give each case its own state.** Allocate whatever the subtest mutates inside the subtest closure. Each shard's test computes its own total and asserts on it immediately. Nothing crosses between siblings, so there is nothing to race on, and a failure names exactly one shard. 2. **Make the inherited fixture read-only.** A parent may legitimately build expensive, immutable setup — parsed reference data, a prepared input directory listing, a compiled pattern — and let every parallel child read it. The discipline is that after the parent's body returns, nobody writes to it. Concurrent reads with no writes are not a race. 3. **Move aggregation out of the parallel phase.** If the point really is a total across shards, compute it in the parent before spawning subtests, or in a separate sequential test. Assertions that span cases and parallel cases are in tension by design. 4. **Lock, only if you must.** A `sync.Mutex` around the shared map removes the race, and for a counter `sync/atomic` is lighter. But a lock in test fixtures is a smell: it says the cases are not independent, and independence was the premise on which `t.Parallel()` was added. ## Why it survives review Sequentially, the code is correct — the map fills in, the assertion passes. Adding `t.Parallel()` to speed the suite up is a one-line change that looks harmless in a diff, and the failure it introduces is timing-dependent, so it may pass a hundred local runs and fail on a busy build machine. That asymmetry is the reason to run tests under the race detector when introducing parallelism, and the reason a reviewer should look at what each parallel case touches, not just at whether the case passes. ## The rule to carry away Under `t.Parallel()`, the only state a case may mutate is state it created itself. Everything reached through a closure from an enclosing test is shared memory between concurrent goroutines, and it must be either immutable from that point on or explicitly synchronised.

  • The race report shows two goroutines at the same source line. What does that tell you about the bug?
    That one piece of code is running concurrently over one object, rather than two different code paths conflicting. It is the signature of a fixture captured by several parallel cases: the fix is to give each case its own instance, not to reorder or serialise the two call sites.
  • Is it ever fine for parallel subtests to share a fixture the parent built?
    Yes, if it is read-only from the moment the parent's body returns. Expensive immutable setup — parsed reference data, a prepared input listing — can be shared safely because concurrent reads without writes are not a race. The discipline is that nothing in the parallel phase mutates it.
  • Why can this bug pass locally and fail on a build machine?
    The race detector only reports interleavings that actually happen, and an unguarded write is only observed as a failure when two goroutines overlap. A loaded, many-core CI machine interleaves them far more often than an idle laptop, so the same code fails there and passes here.
  • Why is adding a sync.Mutex around the shared map a poor first choice?
    It fixes the memory safety and leaves the design problem. A lock in test fixtures says the cases are not independent, which is the premise t.Parallel rests on; it also serialises exactly the part you parallelised. Giving each case its own state is both safer and faster.

It is like three people asked to fill in different rows of one paper form at the same desk at the same time: nothing about having separate rows stops them grabbing the same sheet.

saying these in an interview costs you the question

  • Thinks each subtest closure gets its own copy of the map
  • Believes test code is exempt from data races
  • Expects the parent's aggregate assertion to see subtest results
  • Treats a green run without -race as proof of safety
  • Reaches for a mutex before questioning the shared fixture