Running the whole suite under go test -race doubled CI time — how do you decide what keeps running with it?
answer
- the cost is real, so is the coverage limit
- value tracks executed concurrent code
- not every package has goroutines
- pull request versus merge versus nightly
- advisory jobs get merged past
basics
~20 sSpend the budget where the detector can actually earn it: packages that start goroutines and share state, kept blocking on every pull request, with the full-suite race pass moved to merge or nightly. Then write down which paths nobody runs under -race any more.
solid answer
~50 sA `-race` build runs 2–20x slower and takes 5–10x the memory, so this is a real capacity decision, not a tuning knob. The detector only ever reports races on code that executes, so uniform coverage of the whole tree is the wrong shape: most packages have no goroutines and buy nothing. I would keep `-race` mandatory and blocking on the packages that actually start goroutines or hold shared state, run the full-suite `-race` pass on merge to the main branch or nightly, and give it its own job with its own timeouts and memory limits rather than folding it into the existing matrix — otherwise the first symptom is timeouts, and the team learns to distrust the job. What I would not do is make it advisory: a race job nobody must fix is spend with no return. Finally I would record the residual risk explicitly, because the paths dropped from `-race` are now paths whose races arrive as production incidents.
go deeper
You are not expected to set this policy. Know that a race build is much slower and heavier than a normal one, and that where and when your team runs it is a deliberate choice rather than a default.
Be able to explain why running the instrumented suite everywhere is expensive and why the packages without goroutines gain little, and to spot that a suddenly timing-out job is the instrumented build rather than a new bug.
Argue the split concretely: which packages stay blocking, what moves to nightly, and why the job needs its own timeout and memory budget rather than sharing one with the existing matrix.
Own the tradeoff end to end. State what coverage you are buying with the runner budget, refuse the advisory-job compromise, write down the residual exposure the cut creates, and name the conditions that reopen the decision.
## What you are actually buying The decision is a capacity decision with a safety consequence, so start from what the money buys. An instrumented build typically runs 2–20x slower and uses 5–10x more memory. And the detector reports races only on accesses that actually execute. Put those together and the return on a `-race` job is proportional to *how much concurrent code the job executes*, not to how many packages it compiles. That single observation reshapes the usual proposal. "Run `-race` on everything, every time" spends most of its budget instrumenting packages of pure functions where every access is trivially ordered because there is one goroutine. Meanwhile the packages that matter — the cache, the pool, the request pipeline, the background worker — get the same slice of attention as a package of string helpers. ## A defensible split **Keep `-race` blocking where concurrency lives.** Identify the packages that start goroutines or hold state shared across them, and make `-race` mandatory on those for every pull request. This is the highest-value spend and it is usually a small fraction of the tree. **Move the full-suite `-race` pass off the pull-request path.** Run it on merge to the main branch, or nightly. It still catches the case where somebody adds concurrency to a package nobody classified, just with a longer feedback delay — which is the right trade when the alternative is doubling every developer's wait. **Give it its own job with its own limits.** The failure mode of bolting `-race` onto an existing job is not race reports; it is timeouts, out-of-memory kills in constrained runners, and tests that were tuned around real timing suddenly failing. Those look like flakes, and flakes are how a team learns to ignore a signal. A separate job with its own timeout and memory allowance keeps the signal clean. **Extend `-race` to integration-style tests, not only unit tests.** Composed concurrency — a handler, a pool, and a cache used together — is where the interesting races live, and unit tests frequently drive one type through one goroutine, where the detector has nothing to say. ## The options worth arguing against **Advisory-only.** Making the race job non-blocking is the tempting compromise and the worst one. A failure that never stops a merge is read, at best, once. You keep paying for the runners and stop collecting the benefit. **Cutting the timeout instead of the scope.** Shortening timeouts so the instrumented job fits produces failures that are indistinguishable from real problems and teaches everyone to retry. **Repetition as a substitute for coverage.** Running the same tests many times under `-race` mostly buys runtime. The detector does not need repeated attempts to notice unordered accesses; it needs both accesses to run once. Repetition is worth budgeting only for tests whose concurrent path itself executes probabilistically. ## Make the residual risk explicit Whatever you cut, cut it out loud. Write down which packages and which test suites no longer execute under `-race` on the pull-request path, and what the exposure is: races on those paths will now be found later — on the nightly job at best, in production at worst — and the person on call inherits them. That sentence is what makes this a decision rather than a quiet erosion, and it is what lets you revisit it when the suite changes shape or an incident lands. ## Revisit triggers Set the conditions that reopen the decision rather than promising to "review it periodically": a race reaching production from a path you dropped, a package acquiring goroutines that was not on the list, the nightly job routinely red, or CI hardware changing so that the full pass fits again. Cheap to state up front, and it converts a one-off compromise into something the team can own.
- Why not simply make the race job advisory so nothing is blocked?Because you keep the cost and lose the benefit. An advisory failure gets merged past routinely, so the race is observed and shipped anyway, and the team learns the job is noise. If a signal is not worth blocking on, the honest move is to stop paying for it, not to demote it.
- How do you choose which packages must stay under -race?By where concurrency actually is: packages that contain `go` statements, hold state reachable from more than one goroutine, or wrap shared resources such as caches, pools and connection managers. A grep for `go ` plus the packages named in past concurrency incidents gets you most of the list; keep it in the CI config rather than in someone's memory.
- Your team wants -race on a production canary instead of in CI. What is your answer?It exercises real paths, which is genuinely what CI cannot do, but it costs 2–20x throughput and 5–10x memory on live traffic, and a race there is discovered after shipping. It is a supplement for a service you can afford to run degraded, never a replacement for a blocking test job.
saying these in an interview costs you the question
- Makes the race job advisory to keep merges moving
- Shortens timeouts instead of narrowing scope
- Assumes uniform coverage of every package is best
- Buys repetition instead of concurrent-path coverage
- Drops coverage without recording what was dropped