skip to content

How do you decide whether go test -race runs on every pull request, only nightly, or only over some packages?

level: principalimportance: should knowfreq 40%

answer

  1. cheapest place to catch, against minutes spent
  2. attribution is what nightly loses
  3. races live where goroutines live
  4. fix cost by sharding before cutting scope
  5. an exclusion needs an owner and a date

basics

~20 s

Weigh where a race is cheapest to catch against what the instrumented run costs. A workable split runs the detector over the concurrent packages on every pull request and the whole tree nightly and before merge, with a named owner and a review date for every exclusion.

solid answer

~50 s

Start from what each posture buys. Per-pull-request is where a race is cheapest to fix, because the author still has the change in their head and the report points at their diff; the price is instrumented minutes on every push. Nightly moves the cost off the critical path but destroys attribution — a failure covers a batch of commits and lands on whoever reads the dashboard. A package subset gets most of the value cheaply, because races live where goroutines do, but someone has to keep the list honest as the code grows. I would measure the instrumented job first, then default to the detector on packages that start goroutines or share state per pull request, the full tree under `-race` nightly and on the merge queue, and treat cost as an engineering problem — sharding, persisted build caches, bigger runners — before narrowing coverage. Any exclusion gets an owner, a reason and a date.

go deeper

for a junior

Know that running the detector on every change costs real CI time and that teams make a deliberate choice about where it runs, rather than it being on everywhere by default.

for a middle

Be able to compare per-pull-request, nightly and per-package postures on cost and on how quickly a race gets attributed to the change that introduced it.

for a senior

Bring measurements: what the instrumented suite costs today, and which engineering levers you would pull before reducing coverage. Show how you verify the step can actually fail the build.

for a principal

Own the trade and its escalation. Say who pays, what the exit criteria are, how exclusions are recorded and reviewed, and what you would change the day a race reaches production.

## The decision, stated honestly Running the race detector is not free and its value is not uniform, so somebody has to decide where the instrumented minutes go. That person is usually whoever owns the CI budget, and they can be overruled — by an engineering lead after a race reaches production, or by the same budget owner after the bill arrives. Treating it as a purely technical question is the mistake; it is a spend allocation with a risk attached, and it should be argued as one. ## What each posture actually costs **Every pull request.** The strongest posture, and the one with the best economics per race found, because the report arrives while the author still has context and the change is small enough that the diff and the stack line up. Cost: every push pays for an instrumented build and run, which is several times the plain suite. On a busy repository that is a real line item, and it is also latency — developers wait on it. **Nightly, or on a schedule.** The cost moves off the critical path entirely and the developer experience is unchanged. What you lose is attribution: a failure covers everything merged that day, so somebody has to bisect, and that somebody is rarely the author. Races also live on the main branch until the next run, so anything cutting a release inside that window ships them. The practical failure mode is worse than the theoretical one: nightly race failures tend to accumulate an owner of nobody. **A package subset per pull request.** The pragmatic middle. Races require concurrency, so instrumenting the packages that start goroutines, hold caches or share state captures most of the risk for a fraction of the minutes. The weakness is drift: the list is correct on the day it is written, and six months later a package that had no goroutines has three. Mitigate it with a rule — any package that gains a goroutine gains the detector, and any package implicated in a production race is added permanently — and by running the full tree under the detector on a slower cadence so drift is eventually caught. ## How I would actually decide Measure first. Run the whole suite under `-race` once and record duration and peak memory per package. That single number turns the argument from opinion into arithmetic: if the instrumented run costs a few extra minutes, the debate is over and it goes on every pull request. If it costs half an hour and OOMs the runner, you have a scaling problem to solve before you have a policy question. Then exhaust the engineering levers before cutting scope. Shard the tree across parallel jobs. Persist the build cache so each job is not rebuilding the instrumented standard library. Check you are not running the whole suite twice — once plain, once instrumented — when the instrumented run already covers it. Size the runner for the detector's memory rather than accepting an OOM as a fact of life. Reducing what runs under the detector is the last lever, not the first, because it is the only one that removes signal. Then write the policy down with its exit criteria: what runs where, why, who owns it, and when it is next reviewed. ## Protecting the gate Whatever the posture, two failure modes matter more than the scope decision itself. The first is a gate that has stopped gating. The detector's value in CI rests entirely on the non-zero exit failing the step. A step that forces success, or that only greps the log, is a job that costs money and reports nothing. Audit for that specifically; it is the most common way a race check dies. The second is silent removal. `-race` disappearing from a command line in a routine pipeline change is a coverage decision made by whoever was unblocking their build that afternoon. Make it visible: the flag lives in a place with review, exclusions are listed with a reason and a date, and the list is reviewed on a cadence rather than accumulating forever. ## When you are overruled If a race reaches production, the conversation reopens and you should already know what you will say: which posture was in force, whether the racing package was covered, and what the coverage would have cost. If it was excluded to save minutes, the honest answer is that the trade was wrong for that package and it moves to the per-pull-request set permanently. If it was covered and the detector still missed it — because no test exercised the path — then the answer is not more instrumented minutes at all but a test that drives the path, and saying so plainly is more valuable than promising to run more.

  • A team moves the race job from every pull request to nightly. What have they actually given up?
    Attribution and immediacy. A nightly failure spans a day of merges, so someone has to bisect and it is rarely the author, who has moved on. Races also sit on the main branch until the run, so anything released in that window ships them. Coverage is the same; the ability to act on it cheaply is not.
  • How do you keep a per-package race list from going stale?
    Tie it to a rule rather than to a review: any package that gains a goroutine or shared state joins the list, and any package implicated in a production race joins permanently. Back it with a full-tree instrumented run on a slower cadence so drift shows up, and give the list an owner and a review date.
  • Your CI budget owner asks to drop the detector entirely for a quarter to cut spend. What do you offer instead?
    Numbers and alternatives: what the instrumented job costs today, what sharding and a persisted build cache remove from that, and what the narrowest defensible scope would be — the packages with goroutines, per pull request. If the answer is still to drop it, that is a recorded decision with an owner and a review date, not a silent flag removal.

saying these in an interview costs you the question

  • Treats it as purely technical, with no owner or cost named
  • Argues for nightly without acknowledging lost attribution
  • Cuts detector scope before trying sharding or caching
  • Leaves a race step that cannot fail the build
  • Excludes packages with no reason, owner or review date