skip to content

What does the GORACE environment variable control, and when would you set halt_on_error or history_size?

level: seniorimportance: nice to knowfreq 22%

answer

  1. space-separated key=value pairs
  2. read by the race runtime at start
  3. one knob decides stop or continue
  4. another buys the earlier access's stack
  5. history costs memory, one to seven

basics

~20 s

GORACE carries space-separated options read by the race runtime at process start. halt_on_error=1 stops at the first report instead of continuing; history_size, 0 to 7 with a default of 1, enlarges each goroutine's access history so the earlier access's stack can still be printed.

solid answer

~50 s

`GORACE` is a space-separated list of `key=value` options the race runtime parses when an instrumented process starts, so you set it on the environment of the run: `GORACE="halt_on_error=1" go test -race ./cache`. The useful ones are `halt_on_error` (default 0 — the run continues and collects every race; set to 1 to die at the first one, which is what you want when hunting a specific report and not what you want in CI, because you lose the rest of the suite's results); `history_size` (0 to 7, default 1, each step doubling the per-goroutine memory-access history, so the report can still show the stack of the *earlier* access instead of admitting it could not restore it — at the cost of more memory); `exitcode`, the status used when a process exits because races were detected, defaulting to 66; `atexit_sleep_ms`, defaulting to 1000; plus `log_path` and `strip_path_prefix` for where reports go and how paths are printed.

code

text · 1 line
text
GORACE="halt_on_error=1 history_size=4" go test -race -run TestStoreConcurrent ./cache

go deeper

for a junior

Know that the detector has runtime options supplied through the GORACE environment variable and that you do not normally need to set any of them.

for a middle

Be able to name halt_on_error and history_size and say what each changes about the report, and note that the variable is read by the instrumented binary at start-up.

for a senior

Demonstrate the debugging loop: a CI report missing the earlier stack, a targeted local re-run with a larger history, both stacks recovered. Explain why halt_on_error belongs in that re-run and not in the pipeline.

for a principal

Recognise that no option here suppresses a race. Decide what the team does with a race it will not fix now, and make sure the answer is a recorded, revisited decision rather than a quietly narrowed test path.

## What GORACE is The instrumentation compiled into a `-race` binary is configured at process start by a single environment variable, `GORACE`, whose value is a space-separated list of `key=value` pairs. It is read by the race runtime, not by the go command, which means it applies equally to `go test -race`, `go run -race` and a binary you built with `go build -race`. Because `go test` passes the environment through to the test binary it spawns, setting it on the `go test` invocation works as expected. ## The options worth knowing **halt_on_error** (default 0). By default a detected race does not stop the process: the runtime prints the report and execution continues, so one run can surface several distinct races, and at exit you get `Found N data race(s)`. Setting `halt_on_error=1` makes the process exit at the first report. That is genuinely useful when you are hunting one race and want the smallest possible log, or when a race corrupts state so badly that everything after it is noise. It is a poor default for a suite, because the binary dies mid-run and you lose the results of every test that had not finished — you learn about one race and nothing else. **history_size** (default 1, range 0 to 7). The detector keeps a bounded history of recent memory accesses per goroutine so it can reconstruct the stack of the *previous* access when a conflict is found. When that history has been overwritten, the report says as much instead of showing the second stack — and a report with only one stack is far harder to act on, because the whole value of the output is seeing both sides of the conflict. Each increment doubles the history, and the memory cost doubles with it, which matters given that instrumented runs are already memory-heavy. So this is a targeted setting: raise it for the specific run you are debugging, not for the whole CI suite. **exitcode** (default 66). The status a process exits with because races were detected. It exists so a wrapper can distinguish a race from an ordinary failure. Under `go test` you rarely need to change it — the testing package has already failed the racing test and `go test` reports FAIL — but a custom harness running an instrumented binary directly may care. **atexit_sleep_ms** (default 1000). How long the runtime sleeps before exiting after a race, giving in-flight goroutines a chance to finish printing. **log_path** and **strip_path_prefix**. The first controls where reports are written; the second trims a leading path prefix from the file names in the report, which is how you stop every frame reading like a container build path and make reports comparable between a laptop and CI. ## How this shows up in practice The common sequence is: CI reports a race, the log shows one usable stack and a note that the earlier access's stack could not be restored, and you cannot tell which other code path was involved. You re-run the single failing test locally with a larger history and the flag to stop at the first report, get both stacks, and fix it. The second common use is cosmetic but real — a team whose CI builds under a long generated path adds `strip_path_prefix` so that reports name `cache/store.go` rather than a forty-character build directory, which makes them greppable and diffable across runs. ## What GORACE is not for It does not change what counts as a race, and it cannot make the detector see accesses that did not execute. It does not reduce the instrumentation overhead — there is no fast mode. And nothing in it lets you suppress a specific known race in CI; the Go detector deliberately offers no suppression file in the way some other tooling does, so a race you have decided to live with has to be handled by not exercising that path under the detector, which is a decision you should be uncomfortable making.

  • Why is halt_on_error=1 a bad default for a CI suite?
    The process exits at the first report, so every test that had not finished produces no result. You learn about one race and lose the rest of the run's signal, including other races that would have been reported. It is a debugging setting for a targeted re-run, not a pipeline default.
  • A race report shows only one stack and says the other could not be restored. What do you do?
    Re-run the offending test with a larger `history_size` in `GORACE`. The detector keeps a bounded per-goroutine history of recent accesses; when it has been overwritten, the earlier access's stack is gone. Doubling the history a couple of times usually recovers it, at the cost of more memory for that run.
  • Can you use GORACE to suppress a race you have decided not to fix?
    No. The Go race detector offers no suppression list through this variable; the options cover reporting, history, exit status and output paths. Living with a known race means not exercising that path under the detector, which removes coverage rather than the race, and should be a recorded decision.

saying these in an interview costs you the question

  • Thinks GORACE can suppress a known race
  • Sets halt_on_error=1 as the CI default
  • Believes history_size makes detection more sensitive
  • Assumes the variable is read by the go command, not the binary
  • Expects a GORACE option that lowers instrumentation overhead