skip to content

Setting GODEBUG=asyncpreemptoff=1 makes a flaky failure disappear after a Go upgrade. What does that tell you?

level: seniorimportance: nice to knowfreq 22%

answer

  1. the toggle changes exactly one mechanism
  2. what is unusual about being stopped mid-instruction
  3. signals make slow syscalls return early
  4. it narrows the search, it is not a cure
  5. unretried EINTR is the usual culprit

basics

~20 s

That the failure depends on the runtime interrupting goroutines with a signal at arbitrary instructions, since asyncpreemptoff=1 leaves only function-call preemption. Treat the setting as a bisect that narrows the bug, not as a fix to ship.

solid answer

~50 s

`GODEBUG=asyncpreemptoff=1` disables signal-based preemption, so goroutines are stopped only at function-call safe points. If the failure vanishes, something in the process is sensitive to being interrupted mid-instruction by the runtime's preemption signal. The usual suspects are code that makes raw system calls and treats `EINTR` as fatal instead of retrying — asynchronous preemption makes slow syscalls return `EINTR` far more often — hand-written assembly whose frames the runtime cannot describe, and latent `unsafe` pointer bugs that only bite when a goroutine is stopped at an unusual instruction. I would use the toggle as a bisect: confirm the repro with it on and off, capture the failing stack, run under `-race`, then find the specific call site. I would not ship it. With asynchronous preemption off, a loop containing no calls is unpreemptible again, which stretches stop-the-world pauses and starves other goroutines.

go deeper

for a junior

Know that GODEBUG is an environment variable that changes Go runtime behaviour, and that asyncpreemptoff=1 turns off the runtime's signal-based preemption of goroutines.

for a middle

Explain the mechanical link: with the setting on, no preemption signal is delivered, so goroutines are stopped only at function calls and slow system calls stop being interrupted.

for a senior

Demonstrate the investigation, not the workaround: confirm the reproduction rate both ways, name the likely causes such as unretried EINTR or hand-written assembly, chase the call site, and refuse to ship the toggle.

for a principal

Own the upgrade posture: what a canary must exercise before a fleet moves to a new Go release, and the standing rule that a runtime debug setting is a time-boxed diagnostic with an owner, never a permanent deployment knob.

## What the toggle actually changes `GODEBUG=asyncpreemptoff=1` is an environment-variable switch read by the Go runtime at startup. It turns off exactly one mechanism: signal-based asynchronous preemption, added in Go 1.14. With it set, a goroutine can still be preempted — but only where the compiler left a safe point, which in practice means at a function call. The program reverts to Go 1.13 scheduling semantics while keeping everything else about the current toolchain. That narrowness is what makes it a good bisect. It changes one variable, and it changes it in a direction whose consequences are well understood. ## What a positive result means If the failure reproduces with asynchronous preemption on and disappears with it off, the bug is in the class of things that only happen when a thread is interrupted by a signal at an arbitrary instruction. That is a much smaller space than 'something changed in the Go upgrade': **Interrupted system calls.** A signal delivered to a thread parked in a slow system call causes that call to return early with `EINTR`. Asynchronous preemption sends such signals routinely, so any code path that treats `EINTR` as a hard error rather than retrying now fails intermittently under load. The standard library's `os` and `net` packages retry for you; your own `syscall` wrappers, and any C library called through cgo, may not. This is the single most common cause and the first thing to check. **Assembly and non-Go frames.** The runtime will only stop a goroutine where it has a map of the pointers in its registers and frame. Hand-written assembly that does not follow the conventions the runtime expects, or that is marked as not needing a stack check, can behave badly if the surrounding machinery makes assumptions the async path violates. **Latent unsafe bugs.** Code holding a raw pointer that the collector cannot see, or reusing memory in a way that only works if a goroutine is never stopped at a particular moment, can be perfectly stable for years and then fail once the runtime starts interrupting at new instructions. The async preemption did not create the bug; it removed the accident that hid it. **A signal your own process cares about.** A component that installs signal handlers, or that counts on system calls never being interrupted, meets a stream of signals it never saw before. ## Why you must not ship it It is tempting to add the environment variable to the deployment and close the ticket. Do not, for three reasons. 1. **You have re-opened a real defect class.** Without asynchronous preemption, a loop with no function call in it cannot be preempted. Any operation that needs every goroutine to stop waits for that loop. Instead of a rare intermittent failure you now have latency spikes that are harder to attribute — and, in the worst shape, a goroutine that can never be stopped at all. 2. **The bug is still there.** Something in your process is unsafe when interrupted. That is not a scheduling preference; it is a correctness problem that will resurface through another path. 3. **It is a debugging knob, not a supported tuning parameter.** Runtime debug settings are not part of the contract you want a production fleet standing on, and the next person to read the deployment manifest will have no idea why it is there. ## The disciplined way to use it - **Confirm it is a real signal, not noise.** An intermittent failure that goes away can go away for many reasons. Get a reproduction rate with the toggle off and on, over enough runs to be meaningful. - **Keep the toggle scoped.** One environment, one canary, time-boxed, with a tracking issue that has an owner and an expiry. - **Chase the concrete call site.** Grep your own `syscall` usage for `EINTR` handling; wrap those calls in retry loops keyed on `errors.Is(err, syscall.EINTR)`. Run the failing workload under `-race`. Capture the panic's stack trace and see which package is on it. - **Fold it into the upgrade plan.** If you are the one writing the plan for moving a fleet to a new Go release, this belongs in it explicitly: run the suite under load, on realistic shapes, and treat the toggle as the diagnostic that tells you whether a regression is preemption-related — before it becomes a production incident rather than a canary finding. ## What a good answer sounds like The toggle disables signal-based preemption, so the failure needs a goroutine to be stopped at an arbitrary instruction. That points at unretried `EINTR`, at hand-written assembly, or at `unsafe` code that was accidentally safe before. Use it to bisect, fix the actual call site, and never ship it — with it on, call-free loops are unpreemptible again and pauses grow.

  • What does leaving GODEBUG=asyncpreemptoff=1 on in production actually cost you?
    Any loop with no function call in it becomes unpreemptible again. Operations that need every goroutine to stop then wait for that loop, so pauses lengthen and other goroutines are starved of scheduling turns. You trade a rare, diagnosable failure for latency behaviour that is harder to attribute — and the original defect is still in the binary.
  • How do you write a system-call path that survives frequent EINTR?
    Retry rather than fail. Most `os` and `net` operations already retry internally, so the exposure is your own `syscall` calls and any C code you call. Loop while the error matches `syscall.EINTR` — `errors.Is(err, syscall.EINTR)` — and only then treat the failure as real. Interruption is expected, not exceptional.
  • You are writing the plan to move a fleet onto a new Go release. How would you catch this class of problem before production?
    Exercise realistic workloads, not just unit tests: run the suite and a load-test shape under `-race`, on a canary, long enough for intermittent failures to surface. Audit your own syscall and assembly code for interruption assumptions. Keep the toggle in the plan as a diagnostic to classify a regression quickly, with an explicit rule that it never becomes a permanent deployment setting.

saying these in an interview costs you the question

  • Ships GODEBUG=asyncpreemptoff=1 in production as the fix
  • Concludes the Go release is broken and stops investigating
  • Thinks the setting disables the goroutine scheduler entirely
  • Assumes GODEBUG settings only affect debug builds
  • Dismisses EINTR because the standard library handles it