skip to content

Your go test -race CI job now gets OOM-killed and blows its time budget. How do you get it green again without dropping the detector?

level: seniorimportance: should knowfreq 45%

answer

  1. read the job's own graphs first
  2. one package usually dominates
  3. shadow state is not the Go heap
  4. how many binaries are alive at once
  5. shrink the fixture before shrinking the check

basics

~20 s

Measure per-package duration and peak memory first, then cut concurrency with go test -p and -parallel, shard the tree across jobs, shrink oversized fixtures and goroutine fan-out, and size the runner for the detector's five-to-ten-times footprint. Disabling the flag is the last resort.

solid answer

~50 s

First find out where it goes. The job's duration and peak RSS graphs tell you when it started drifting, and running `-race` package by package tells you which one dominates. In my experience it is usually one package — the in-process caching layer whose tests fill a shared map keyed by a struct and read it from many goroutines — and instrumentation multiplies exactly that memory five to ten times. Then the levers: lower `-p` so fewer instrumented package binaries are alive at once, lower `-parallel` inside each binary, and shard the tree across CI jobs so each has its own memory ceiling. Shrink the fixture: a thousand entries usually reproduces what a hundred thousand did. Cap goroutine fan-out, because the instrumented runtime dies past a few thousand live goroutines. Persist the build cache. And be willing to pay for a bigger runner — CI minutes are cheaper than a race reaching production.

code

text · 6 lines
text
# One instrumented package binary at a time, fewer parallel tests inside it.
go test -race -p 1 -parallel 2 ./...

# Or split the tree across CI jobs, keeping -race on every shard.
go test -race ./cache/...
go test -race ./transport/...

go deeper

for a junior

Know that instrumented runs cost far more memory and time, and that the first step is to measure which package is expensive rather than to guess at flags.

for a middle

Be able to name the concrete levers: -p for package binaries in flight, -parallel inside a binary, smaller fixtures, and a persisted build cache. Explain why collector knobs do not help.

for a senior

Show the whole loop: read the job's duration and RSS graphs, bisect by package, apply the cheapest lever, and set the new numbers as a budget you alert on. Name the destructive shortcuts and why you reject them.

for a principal

Frame the spend. Decide whether the answer is bigger runners, more shards or narrower scope, say who pays, and make sure any reduction in coverage is a recorded decision rather than an edit that nobody reviews.

## Diagnose before you tune The symptom — an OOM kill and a blown timeout — is the job's, not the code's, so start from the job's own telemetry: the duration graph and the peak RSS graph over the last few weeks. A step change points at a specific merge; a steady climb points at a fixture or a test count that has been growing. Then reproduce locally with the same flags and bisect by package, because `go test -race ./...` reports per-package timings and you can run the suspects alone under `/usr/bin/time -v` or the container's own accounting. What you almost always find is that the cost is not spread evenly. One package dominates — typically the concurrent one. In a service with an in-process caching middleware in front of an upstream, the cache's tests populate a shared map keyed by a struct and then read it from dozens of request goroutines. That is a lot of distinct memory touched by a lot of goroutines, which is the worst case for the detector on both axes at once. ## Understand which memory it is The extra footprint is the detector's shadow state, mapped and managed by the race runtime rather than allocated on the Go heap. That single fact rules out the first thing most people try: `GOGC` and `GOMEMLIMIT` govern the Go heap and will not shrink shadow memory. Setting them lower just makes the collector work harder for no benefit. What does move the number: - **Fewer instrumented processes at once.** `go test` builds and runs package test binaries concurrently, governed by `-p`, which defaults to `GOMAXPROCS`. Peak job memory is roughly the per-binary cost times the number in flight. `-p 2` or `-p 1` on the race job trades wall clock for headroom, and is often the difference between an OOM kill and a green run. - **Fewer parallel tests inside a binary.** `-parallel` bounds how many `t.Parallel` tests run simultaneously in one package, which bounds how much live data and how many goroutines exist at the peak. - **Smaller fixtures.** Shadow state is proportional to distinct memory touched. If a cache test seeds a hundred thousand entries, ask what the test is actually proving; a thousand entries exercises the same concurrent read and write paths at a tenth of the cost. This is usually the biggest single win and it costs nothing in coverage. - **Less fan-out.** The instrumented runtime enforces a ceiling on simultaneously live goroutines in the low thousands and dies outright when a test exceeds it. A load-shaped test that spawns fifty thousand goroutines has to be scaled down, not exempted. ## Then attack the wall clock Sharding is the honest answer to duration. Split the tree across parallel CI jobs — by top-level directory or by an explicit list — with `-race` on every one of them. Each shard gets its own memory ceiling and its own runner, and the total elapsed time falls even though total CPU rises. Persisting `GOCACHE` between runs removes the cold-cache rebuild of the instrumented standard library, which on a short suite can be most of the job. And check that the race job is not re-running everything that a non-race job already ran: many teams end up executing the whole suite twice, once each way, when the plain run adds nothing the instrumented one does not already cover. ## What not to do The tempting fixes are the destructive ones, and an interviewer is listening for whether you name them as traps. Raising the job's timeout until it passes hides a trend that will come back. Deleting the slowest concurrency tests removes exactly the tests the detector exists to serve. Appending something to the step that always succeeds turns a failing race report into a green build — which is worse than having no race job at all, because the dashboard now lies. And quietly dropping `-race` from the command is a policy change disguised as a config edit; if the detector really has to come off a path, that belongs in a decision with an owner and a review date, not in a one-line diff. ## Close the loop After the fix, record the new numbers as the budget and alert on them. A race job that has doubled in duration since last quarter is telling you something about how much shared state the service has grown, and that is a signal worth reading rather than absorbing.

  • Why does lowering -p help memory when the tests themselves are unchanged?
    `-p` bounds how many package test binaries the go command runs concurrently, defaulting to `GOMAXPROCS`. Each instrumented binary carries its own shadow state, so the job's peak is roughly the per-binary cost times the number in flight. Running one or two at a time trades elapsed time for a much lower ceiling.
  • A colleague proposes appending a shell construct that forces the race step to succeed so the pipeline unblocks. What is your response?
    That deletes the gate while leaving the dashboard green, which is strictly worse than removing the job, because now nobody knows the check is gone. If the step must stop blocking, make it a visible, owned decision with a date to revisit, and fix the actual cost with sharding and runner size in the meantime.
  • The race job dies with a runtime message about too many live goroutines. Is that a detector bug?
    No. The instrumented runtime tracks per-goroutine state and enforces a ceiling in the low thousands on simultaneously live goroutines. A test that fans out far beyond that has to be scaled down for the race build — the concurrency it is proving is almost never dependent on the exact count.

saying these in an interview costs you the question

  • Reaches for GOMEMLIMIT to shrink the detector's footprint
  • Raises the job timeout instead of finding the growth
  • Deletes concurrency tests because they are the slowest
  • Forces the race step to exit zero so the pipeline unblocks
  • Tunes flags without measuring which package dominates