skip to content

A Go concurrency bug fails once in 200 CI runs. How do you build a test that reproduces it?

level: seniorimportance: should knowfreq 42%

answer

  1. one in two hundred is a sampling problem
  2. shrink the target, then raise the pressure
  3. replay recorded traffic at a chosen concurrency
  4. count failures per 500 runs, before and after
  5. silence proves nothing until you ran enough

basics

~20 s

Isolate the suspect component into its own test, drive it with replayed recorded traffic at a chosen concurrency, raise GOMAXPROCS, and repeat the case hundreds of times under go test -race. Then quantify: measure the failure rate per hundred runs before and after the fix.

solid answer

~50 s

Stop reproducing the whole pipeline and build a rig around the suspect component: replay a recorded request log against it at a concurrency you control, so the input is realistic and the pressure is a knob rather than an accident. Then multiply the odds — loop the case hundreds of times inside one test, make sure `GOMAXPROCS` is above one so goroutines really run simultaneously, and run it under `go test -race`, which reports a conflicting pair of accesses on the run where it happens rather than waiting for the corruption to surface. The output you want is a number, not an anecdote: failures per 500 runs, recorded before the change and after it. That is what makes "fixed" a claim you can defend — a bug that shows up in 1 run of 200 will survive 600 clean runs about 5% of the time, so a single green run proves nothing. Keep the rig as a soak test behind `testing.Short()` and run it nightly.

code

go · 9 lines
go
func TestReplaySoak(t *testing.T) {
	if testing.Short() {
		t.Skip("soak test; runs in the nightly job")
	}
	reqs := loadRecordedLog(t)
	for i := 0; i < 500; i++ {
		replay(t, reqs, 16) // 16 concurrent replayers per round
	}
}

go deeper

for a junior

Know the first move: repeat the failing case many times rather than trusting a single run, and run it under the race detector so a conflict is reported when it happens.

for a middle

Explain the levers that raise the odds — more concurrent callers, more than one processor, a smaller unit under test — and why a corrupted value usually surfaces far from the code that corrupted it.

for a senior

Show the discipline: a rig fed by replayed traffic, a measured failure rate before and after, and a stated number of clean runs that makes the fix defensible. Then place the soak nightly and keep a cheap deterministic regression test on every run.

for a principal

Own the standard of evidence the team uses to close a rare concurrency bug, and the budget it costs. Be ready to say who runs the soak, what its failure rate over time triggers, and when a chronically flaky area justifies redesigning the ownership of the state rather than hunting one more race.

## The shape of the problem A failure that appears in one CI run out of two hundred is not a mystery, it is a sampling problem. Every technique below does one of three things: raise the probability per attempt, raise the number of attempts, or make the failure visible on the attempt where it happens instead of several stages later. ## 1. Shrink the thing under test A whole-pipeline test that fails 0.5% of the time is useless as a tool: each attempt is slow, and the failure is far from its cause. Pull the suspect component out and give it its own test. This is also where you decide what the test *drives*, and a recorded request log replayed against the component beats a synthetic loop: real inputs bring the shapes, sizes and repeat keys the bug actually needs, and the replay's concurrency becomes a dial you set rather than a property of production traffic. ```go func TestReplaySoak(t *testing.T) { if testing.Short() { t.Skip("soak test; runs in the nightly job") } reqs := loadRecordedLog(t) for i := 0; i < 500; i++ { replay(t, reqs, 16) // 16 concurrent replayers per round } } ``` ## 2. Raise the probability per attempt - **Concurrency.** Two goroutines rarely collide; sixteen collide often. The replay's worker count is the cheapest dial you have. - **`GOMAXPROCS` above one.** With a single processor, two goroutines never execute at literally the same instant, so genuinely parallel interleavings become far rarer. A reproducer pinned to one core that stays green has proved nothing; make sure the rig runs with several. - **Widen the window during triage.** A `runtime.Gosched()` at a suspicious point, in a scratch branch, makes a narrow race enormous. It is a diagnostic instrument, not a change you commit. - **Remove anything that serialises accidentally.** A shared logger with a mutex, a single-connection fake, or a debug print inside the hot path can hide the very overlap you are hunting. ## 3. Make the failure visible where it happens Run the loop under `go test -race`. Without it you are waiting for a corrupted value to travel far enough to break an assertion, which may take many more iterations and points at the wrong code; with it, a conflicting pair of accesses is reported with both stacks on the run where it occurs. Add a goroutine-leak check at the end of the run too — recent Go exposes a `goroutineleak` profile in `runtime/pprof` that reports goroutines the runtime can prove will never resume — because the same defect that races often also parks a goroutine forever, and a leak is deterministic evidence where a race is probabilistic. ## 4. Turn the anecdote into a number This is the part teams skip and the reason "fixed" is so often wrong. Build a small table before you change any code: | build | runs | failures | rate | |---|---|---|---| | main | 500 | 7 | 1.4% | | candidate fix | 500 | 0 | ? | The question mark is honest. If the true rate were unchanged at 0.5% per run, 600 clean runs would happen by luck about 5% of the time (0.995^600 is roughly 0.05). So the number of clean runs you need is a function of the rate you measured: with a rate you have *raised* to a few percent by the steps above, a few hundred clean runs is strong evidence; with the original 1-in-200 you would need thousands. That arithmetic is the whole reason step 2 matters — you are not just reproducing faster, you are making the eventual proof affordable. ## 5. Decide where the rig lives afterwards A five-minute soak has no place on every pull request. Guard it with `testing.Short()` so `go test -short` skips it, and run the full version in a nightly job that records the failure count over time — a rate creeping up from zero is an early warning that something else has changed. Meanwhile, once you understand the bug, write the cheap deterministic test that reproduces it by construction and keep *that* on every run; the soak is for discovery, the deterministic test is for regression. ## What not to do Do not lengthen a sleep until CI goes quiet — that hides the bug and slows every run. Do not delete or skip the test that caught it. Do not change the production code before you have a reproducer, because then you have no way to know whether the silence afterwards is the fix or the sampling. And do not report the bug as fixed on the strength of one green pipeline; with a rare failure, one green run is the single weakest piece of evidence available.

  • How many clean runs justify calling a 1-in-200 failure fixed?
    Do the arithmetic rather than guess: at an unchanged 0.5% per run, 600 clean runs still happen about 5% of the time by luck, so that is weak evidence. The practical move is to raise the failure rate first — more concurrency, more parallelism, a tighter test — so a few hundred clean runs at the new, higher rate becomes convincing.
  • Why does the reproducer need GOMAXPROCS above one?
    With one processor, goroutines interleave but never execute simultaneously, so failures that need two goroutines genuinely running at once become very rare. A rig that stays green on a single processor has not exonerated the code; set the concurrency and the processor count deliberately and record both alongside the failure rate.
  • What belongs in the suite afterwards: the soak rig or a deterministic test?
    Both, in different places. The deterministic test that reproduces the understood bug by construction runs on every change; the soak stays behind testing.Short() and runs nightly, where its failure count over time is a signal in its own right. Keeping only the soak makes every pull request slow, keeping only the deterministic test stops the search for the next bug.

saying these in an interview costs you the question

  • Lengthens a sleep until CI stops complaining
  • Declares the bug fixed after one green run
  • Skips or deletes the test that caught it
  • Runs the reproducer on a single processor and concludes nothing is wrong
  • Changes production code before having any reproducer