skip to content

A worker-pool test sleeps 50ms before asserting and goes red in CI weekly — how do you prove the sleep is the cause?

level: seniorimportance: must knowfreq 62%

answer

  1. one a week is not evidence, one a minute is
  2. narrow, repeat, and never let it report cached
  3. the detector is being used to change timing here
  4. your idle laptop is the worst place to look
  5. move the constant and watch the rate move

basics

~20 s

Make it fail on demand rather than reasoning about it: narrow with -run, repeat with -count, build with -race, and squeeze the parallelism the way CI does. Then vary the sleep constant — if the failure rate tracks it, the sleep is the synchronisation.

solid answer

~50 s

The goal is to turn one red run a week into a red run a minute on my own machine. I narrow to the single test with `-run`, repeat it with `-count=200` and `-failfast` so the first failure keeps its output, and build with `-race`, which both reports genuine data races and perturbs timing enough to shake out marginal ones. If it stays green I load the box the way CI is loaded: pin a low parallelism with `GOMAXPROCS=1`, or run several copies of the loop at once. The decisive experiment is to vary the constant: cut the sleep to 1ms and the failure rate should climb sharply, remove it and the test should fail every time. A result that tracks a sleep duration is proof that the sleep is doing the synchronising. The repair is to wait for the workers rather than for the clock — never a longer sleep, which only lowers the rate and slows the suite.

code

go · 23 lines
go
func TestWorkerPoolResults(t *testing.T) {
	jobs := make(chan int, 3)
	results := make(chan int, 3)
	var wg sync.WaitGroup
	for range 3 {
		wg.Add(1)
		go func() {
			defer wg.Done()
			for j := range jobs {
				results <- j * 2
			}
		}()
	}
	jobs <- 1
	jobs <- 2
	jobs <- 3
	close(jobs)

	time.Sleep(50 * time.Millisecond) // a bet, not synchronisation
	if len(results) != 3 {
		t.Fatalf("got %d results, want 3", len(results))
	}
}

go deeper

for a junior

Know that a time.Sleep before an assertion is a guess about how long other goroutines take, and that a test failing only sometimes usually means it is waiting for the clock instead of for the work.

for a middle

Explain the reproduction toolkit and why each part helps: -run to narrow, -count to repeat, -failfast to keep the failing output, -race to stretch timing windows, and a low GOMAXPROCS to imitate a constrained CI machine.

for a senior

Show that you make the hypothesis falsifiable rather than plausible. Vary the sleep constant and demonstrate that the failure rate tracks it, and be able to argue to a team why a longer sleep is the same bug at a lower probability.

for a principal

Own the policy: what happens to a flaky test tonight — quarantine with an owner and a deadline versus retry-until-green — and what the suite's standing rule is about sleeps, because the cost of the reflex is a team that no longer believes a red build.

## The shape of the bug A fan-out test starts a few worker goroutines reading from a jobs channel and writing to a results channel, feeds the jobs in, closes the jobs channel, sleeps for a constant, and then asserts on what arrived. The sleep is not synchronisation — it is a bet that 50 milliseconds is longer than the workers need. On an idle developer machine with ten processors that bet wins hundreds of times in a row. On a loaded CI box with a fraction of a CPU available to the job, it loses occasionally, and the loss looks random because from the test's point of view nothing else changed. The interviewer is not asking for the fix — everyone knows the fix is to wait for the workers. They are asking how you establish, in front of a team whose main branch has been red three times this week, that this particular sleep is the cause rather than a plausible suspect. ## Step one: get it to fail where you can watch it CI is a terrible debugger. The whole first phase is about moving the failure onto your machine and raising its rate until it is boring. - **Narrow the run.** `-run TestWorkerPoolResults` cuts the noise and makes each iteration cheap. If the test is in a package with an expensive `TestMain`, that cost is paid once per binary, not per iteration of `-count`. - **Repeat.** `-count=200` executes the selected test two hundred times in one binary. This also sidesteps the cached-result problem: a suspect test must never be allowed to report `(cached)`. - **Stop at the first failure.** `-failfast` plus `-v` keeps the output of the run that actually went wrong instead of burying it under the next hundred passes. - **Build with the race detector.** `-race` is used here purely as a reproduction aid: it slows every memory access, which stretches the windows a timing bug needs, and it will separately report if the flake is in fact an unsynchronised read. A green `-race` run is not proof of correctness — it only observes accesses that actually happened on that run — but a red one ends the investigation immediately. - **Match CI's parallelism.** `GOMAXPROCS=1 go test -count=200 ...`, or `-cpu=1,2,8`, because a CPU-limited CI container runs far less real parallelism than your laptop and the main test goroutine can sail past the assertion before a worker ever runs. - **Load the machine.** Run three or four copies of the loop concurrently, or compile something large alongside it. Contention is the CI condition you are missing. ## Step two: the falsifiable experiment Once it fails locally at some rate, stop gathering evidence and start testing the hypothesis. The hypothesis is "the assertion depends on the sleep constant", and it is falsifiable in one move: **change the constant and watch the rate**. - Shorten the sleep to 1ms: the failure rate should rise sharply. - Delete the sleep entirely: the test should fail nearly always. - Lengthen it to 500ms: the failure rate should fall toward zero without ever reaching it. A quantity that moves monotonically with a timing constant is not a coincidence, and this is the argument to bring to the team. It also demonstrates why the tempting fix is wrong: raising the constant is the *same* experiment run in the direction that makes the evidence disappear. You have not removed the dependency, you have only made it rarer and made every run of the suite slower — and CI, being the slowest and most contended machine, is where the remaining probability will land. ## Step three: rule out the neighbours Two cheap checks stop you fixing the wrong thing: - Run the package with `-shuffle=on` a few times. If the failure only shows at particular seeds, the problem is coupling between tests, not this test's timing. - Check whether the assertion is over something whose order the language never fixed — results arriving from several workers have no defined order, and neither does a range over a map used to build the expectation. If the test asserts an exact sequence, the sleep may be masking a second, independent defect: an ordering assumption that would be wrong even with perfect synchronisation. ## What good looks like afterwards The repaired test waits for a condition, not a duration: the producer closes the jobs channel, the workers finish, something closes the results channel, and the test drains it to completion before asserting on a set rather than a sequence. Its runtime drops from a guaranteed 50 milliseconds to whatever the work actually takes, and its failure rate drops to zero rather than to a smaller number. That is the difference worth articulating: synchronisation makes a test both faster and correct, while a bigger sleep trades correctness you never had for minutes you do have.

  • Why is raising the sleep from 50ms to 500ms a bad fix?
    It changes the failure probability, not the failure mode: the test still asserts that a duration is longer than an unbounded amount of work. It also adds nearly half a second to every run of a suite that may contain dozens of such waits, and the residual probability lands on CI, the slowest and most contended machine you own.
  • The race detector reports nothing across two hundred runs. What have you learned?
    Only that no unsynchronised access happened on the paths those runs exercised. The detector observes real accesses at run time rather than proving anything statically, so a clean sweep is weak evidence — and it is entirely consistent with this bug, which is a missing wait rather than a data race.
  • How do you decide whether the flake is this test's timing or coupling with another test?
    Run the package shuffled a few times and separately run the single test in a long -count loop. Coupling shows up as a failure tied to a particular shuffle seed and disappears when the test runs alone; timing shows up under repetition of the test on its own and ignores the seed entirely.
  • What would you tell the on-call engineer to do with the red build tonight?
    Skip or quarantine the single test with a linked issue rather than retrying the pipeline, so the branch is trustworthy again within minutes and the evidence is not lost. Retrying until green trains everyone to ignore the signal, and the next real failure hides behind the same reflex.

saying these in an interview costs you the question

  • Increases the sleep and calls it fixed
  • Retries CI until it passes and closes the ticket
  • Concludes it is a CI infrastructure problem without reproducing
  • Treats a clean race-detector run as proof of correctness
  • Debugs only on an idle laptop at full parallelism
  • Never varies the sleep constant to test the hypothesis