skip to content

What is a preemption safe point in Go, and how is a goroutine that never reaches one stopped?

level: middleimportance: must knowfreq 45%

answer

  1. the runtime must know what is in the registers
  2. a function's first instructions do more than call the body
  3. the stack-growth check doubles as a hook
  4. with no call to hijack, interrupt the thread instead
  5. the runtime signals the thread with SIGURG

basics

~20 s

A safe point is an instruction where the runtime knows the goroutine's stack layout well enough to stop it, in practice a function prologue's stack check. Go 1.14 added asynchronous preemption, interrupting the thread with SIGURG when no safe point is reached.

solid answer

~50 s

To preempt cooperatively, Go's runtime poisons the goroutine's stack limit so the next function prologue's stack-growth check fails and jumps into the runtime instead of growing the stack. That check is the safe point, and it is free when no preemption is pending. Functions marked `//go:nosplit` have no such check, and a loop body with no calls contains none at all — which is why a tight loop was historically unpreemptible. Since Go 1.14 the runtime also preempts asynchronously: it sends SIGURG to the OS thread running the goroutine, and the handler inspects the interrupted program counter. If the runtime has a stack map for that instruction — an asynchronous safe point — it diverts the thread into a stub that saves the registers and calls the scheduler; if not, it resumes and tries again later. `GODEBUG=asyncpreemptoff=1` disables the asynchronous path.

code

go · 11 lines
go
// Loop A: each iteration calls a function, whose prologue carries a
// stack check - a cooperative preemption point.
for i := 0; i < n; i++ {
	total += score(i)
}

// Loop B: no call at all, so no cooperative safe point exists.
// Go 1.14 and later stop this one by signalling the thread.
for i := 0; i < n; i++ {
	total += uint64(data[i])
}

go deeper

for a junior

Know the vocabulary: a safe point is a place the runtime is allowed to stop a goroutine, function calls provide them, and modern Go can also interrupt a loop that has none.

for a middle

Be able to explain both mechanisms concretely: poisoning the stack limit so the next call's stack check traps into the runtime, and the SIGURG handler that checks whether the interrupted instruction has a stack map.

for a senior

Explain why the mechanism is best-effort - an attempt at an unmapped instruction is abandoned and retried - and what that means when you are chasing a pause you cannot account for.

for a principal

Frame the trade the design makes: a nearly free check on every call plus a rare signal, versus guaranteed instant preemption. Be ready to say what that buys a fleet and where it still costs tail latency.

## Why a safe point is needed at all Stopping a goroutine is not simply a matter of parking a thread. Go has a garbage collector that must find every pointer the program can still reach, including pointers sitting in CPU registers and in the current stack frame. The compiler emits *stack maps* describing, for particular instructions, which slots and registers hold pointers. A goroutine can only be safely stopped at an instruction the runtime can describe that way. Those instructions are the safe points. ## Cooperative preemption: the stack check as a hook Goroutine stacks start small and grow on demand, so the compiler puts a check at the top of nearly every function: is there room for this frame? If not, control goes into the runtime, the stack is copied to a larger one, and the function is retried. The runtime reuses that check as its preemption hook. To preempt a goroutine cooperatively it sets the goroutine's stack limit to a poison value that no real stack can satisfy. The very next function call therefore takes the slow path, lands in the runtime, and the runtime — seeing the poison rather than a genuine stack shortage — reschedules the goroutine instead of growing its stack. The elegance is the cost profile: when nothing is pending, preemption costs a comparison the CPU predicts perfectly. The weakness is coverage. There is a preemption opportunity only where there is a function call with a stack check. That excludes: - functions annotated `//go:nosplit`, which deliberately omit the check; - loop bodies that call nothing — a checksum over a byte buffer, an inner compression loop, a spin on an atomic; - loops whose only call is small enough to be inlined, since an inlined call is no longer a call. Go also does not insert preemption checks at loop back-edges, so writing an arithmetic loop really does produce a call-free stretch of machine code. ## Asynchronous preemption: interrupting the thread Go 1.14 closed that hole. When the runtime wants to stop a goroutine that is not cooperating, it asks the operating system to deliver a signal — SIGURG — to the OS thread running it. SIGURG is a good choice: the runtime can install a handler for it, its default disposition is harmless, and applications and C libraries rarely use it. The handler runs on that thread, on top of the interrupted work, and asks a precise question: *is the interrupted program counter an asynchronous safe point?* That means the runtime has a register and stack map for exactly that instruction and is not inside a region it must not be stopped in — parts of the runtime itself, code holding certain runtime locks, and `//go:nosplit` stretches are excluded. - If it is a safe point, the handler rewrites the interrupted context so that when it returns, the thread runs a runtime stub instead of resuming the loop. That stub spills every register into a frame the collector can scan and then calls into the scheduler, which parks the goroutine. When the goroutine is scheduled again the registers are restored and the loop continues from exactly where it was interrupted. - If it is not a safe point, nothing happens: the handler returns, the goroutine resumes, and the request stands. The runtime will try again shortly. Preemption is best-effort per attempt, guaranteed only in aggregate. ## The consequences you can observe Two are worth knowing. **Syscalls return EINTR more often.** A signal delivered to a thread sitting in a slow system call interrupts it. Go's own `os` and `net` packages retry transparently, but hand-written `syscall` code, and C code called through cgo, may see interruptions it never saw before Go 1.14 and must retry rather than treat them as fatal. **Preemption is not free at the moment it happens.** A goroutine stopped asynchronously has to have its registers spilled and its frame scanned, which is more work than a cooperative stop at a call boundary. This is a fine trade for the guarantee it buys, but it is why cooperative preemption remains the primary mechanism and the signal is the fallback. ## The escape hatch `GODEBUG=asyncpreemptoff=1` turns signal-based preemption off, leaving only call-site safe points — that is, Go 1.13 semantics. It exists to bisect problems that appear when a program is interrupted at arbitrary instructions. It is a debugging tool, not a tuning knob: with it set, a call-free loop is unpreemptible again and can delay any operation that needs the world stopped. ## What a good answer sounds like A safe point is an instruction where the runtime has a map of the pointers in the frame and registers; in practice that is the stack-growth check at the top of a function, which the runtime hijacks by poisoning the stack limit. A loop with no calls reaches no such point, so since Go 1.14 the runtime signals the thread with SIGURG and, if the interrupted instruction is an asynchronous safe point, diverts it into the scheduler — otherwise it resumes and retries.

  • Why does the runtime need a stack map at the exact instruction it preempts?
    Because a preempted goroutine's stack and registers must be scannable. The collector has to distinguish pointers from integers to trace live objects, and it can only do that where the compiler recorded a map. Without one the runtime cannot stop safely, so it abandons that attempt and retries at another instruction.
  • Which Go functions have no cooperative preemption point?
    Functions marked `//go:nosplit`, which omit the stack-growth check by design, and any stretch of code with no call in it — an arithmetic loop, a byte-scanning loop, a spin on an atomic. Inlining matters too: a loop whose only call is inlined away has no call left, so it has no safe point either.
  • Why did the runtime pick SIGURG for preemption?
    It needs a signal it can install a handler for, that does not terminate the process by default, that can interrupt a blocking system call, and that real programs and C libraries almost never use. SIGURG fits: it concerns out-of-band socket data, which practically nothing relies on, so hijacking it collides with almost nobody.

saying these in an interview costs you the question

  • Says Go switches goroutines on a fixed OS time slice
  • Claims every function call switches to another goroutine
  • Thinks asynchronous preemption can stop a goroutine at absolutely any instruction
  • Believes the OS scheduler preempts individual goroutines
  • Says //go:nosplit only affects stack size and not preemption