skip to content

Your `go tool trace` shows a stop-the-world pause of several milliseconds instead of the usual short one. What explains it?

level: seniorimportance: nice to knowfreq 26%

answer

  1. stopping is the expensive half
  2. a thread the OS never runs delays everyone
  3. who asked for the stop?
  4. scrapes on a regular cadence are suspicious
  5. ReadMemStats and goroutine profiles freeze the world

basics

~20 s

A stop-the-world pause is usually dominated by the time it takes to stop every logical processor, not by the work done once stopped. Suspect threads the host is not scheduling, or a caller such as runtime.ReadMemStats.

solid answer

~50 s

A stop-the-world pause has two parts: getting every P to yield, then doing the short piece of work that needs a frozen world. The second part is small and bounded; the first is not, because the runtime must reach every running goroutine, and any thread the operating system is slow to schedule delays the whole stop. So a multi-millisecond pause in a trace usually means the host is oversubscribed or the process is CPU-throttled, not that the collector did more work. The second thing to check is who asked for the stop: `runtime.ReadMemStats`, taking a goroutine profile, writing a heap dump and an explicit `runtime.GC()` all stop the world, and a metrics endpoint scraped every few seconds can quietly add pauses nobody attributed to the collector. Prefer `runtime/metrics`, which does not stop the world, over `runtime.ReadMemStats` in a hot metrics path.

go deeper

for a junior

Know that Go's collector does most of its work concurrently and only freezes the program for short phases, and that a trace shows those phases explicitly.

for a middle

Be able to split the pause into stopping every processor and the work done once stopped, and to name calls other than collection that stop the world.

for a senior

Demonstrate an ordered investigation: cadence, correlation with other stalls, correlation with load, and a quiet-window comparison before you change any setting.

for a principal

Decide when a pause is the platform's problem rather than the service's, and set the team's policy on metrics endpoints that stop the world on every scrape.

## Reading a stop-the-world span In a trace opened with `go tool trace`, a stop-the-world phase is visible as a span during which no application goroutine runs, and in the goroutine analysis it appears as GC pause time distributed across the goroutines that were frozen. Go's collector takes two such phases per cycle and they are designed to be very short — comfortably sub-millisecond on a healthy machine. So a span of several milliseconds is a signal, and the useful question is which of the two components grew. **Component one: stopping.** The world is not stopped until every logical processor has yielded. A goroutine that is running must reach a point where it can be preempted; modern Go preempts asynchronously by signalling the thread, so a tight loop with no function calls is no longer the classic culprit it once was. But the runtime still needs each thread to respond, and a thread that the operating system has not scheduled cannot respond. On an oversubscribed host, or a process whose CPU allowance has been exhausted for this period, the stop drags for as long as it takes the last thread to get a slice of a core. This is the single most common cause of a surprisingly long pause, and the giveaway is that it is *uncorrelated with your program*: it grows with host load, appears on unrelated stops, and shows up as gaps where nothing at all ran. **Component two: the work inside the pause.** This piece is bounded by design and rarely the explanation on its own. If it is large you generally have a pathological amount of runtime state, and you would expect it to grow smoothly rather than spike. ## Who else stops the world The second question worth asking is whether the collector requested the stop at all. Several ordinary calls do: - **`runtime.ReadMemStats`** stops the world to take a consistent snapshot of allocator statistics. A metrics exporter that calls it on every scrape adds a stop on a fixed cadence. - **Taking a goroutine profile** — through `runtime.GoroutineProfile` or the `/debug/pprof/goroutine` endpoint — needs a brief stop. - **Writing a heap dump** with `runtime/debug.WriteHeapDump` stops the world for the duration. - **An explicit `runtime.GC()`** forces a full cycle, with its stops, wherever you called it. This is a genuinely common production surprise: pauses appearing at a suspiciously regular interval that matches a scrape schedule, in a service whose allocation rate does not justify them. Reading memory statistics through the `runtime/metrics` package instead of `runtime.ReadMemStats` avoids the stop, and it is the modern interface for those numbers anyway. ## A workable order of investigation 1. **Are the long pauses periodic and regular?** Regular cadence points at a scraper or a cron-like caller — look for `runtime.ReadMemStats`, profile endpoints, or an explicit forced collection. 2. **Do they cluster in time with other stalls?** If every kind of stop is long in the same window, and per-processor rows in the timeline show gaps where nothing ran, the host is the cause: noisy neighbours, exhausted CPU allowance, or a machine with far more runnable threads than cores. 3. **Do they scale with your workload?** Pauses that grow as traffic grows point back at your program — a much larger number of goroutines to freeze, or far more frequent cycles. 4. **Compare a quiet window against a bad one.** The value of a trace here is that it timestamps everything: you can see whether the pause was long because stopping was slow or because a scrape landed exactly there. ## What not to conclude Do not read a long stop-the-world span as evidence that the collector is doing more marking. Marking is concurrent; assists and background mark workers are where the marking cost lands, and neither of those is a pause. Equally, do not tune collector parameters in response to a pause caused by an oversubscribed host — that changes how often you pay a cost whose length is being set by something outside the process. Confirm which component grew before you touch anything.

  • Which ordinary calls stop the world outside a collection cycle?
    `runtime.ReadMemStats` for a consistent allocator snapshot, taking a goroutine profile via `runtime.GoroutineProfile` or the pprof goroutine endpoint, `runtime/debug.WriteHeapDump`, and an explicit `runtime.GC()`. A metrics exporter calling any of these per scrape produces pauses on a fixed cadence that nobody attributes to the collector. Reading from `runtime/metrics` instead avoids the stop entirely.
  • How would you confirm the pause is caused by the host rather than the program?
    Look for the pattern across the trace rather than at one span: pauses long on every kind of stop, per-processor rows with gaps where nothing ran at all, and no correlation with request volume or allocation rate. If the same binary and workload pause briefly on an unloaded machine, the length is being set outside the process.
  • Why is a tight loop with no function calls no longer the standard explanation for a slow stop?
    Modern Go preempts a running goroutine asynchronously by signalling the thread it is on, so a loop that never calls a function can still be interrupted. That removed the classic case where one uncooperative goroutine held the entire stop. What is left is a thread that the operating system has not scheduled, which no amount of signalling can hurry.

saying these in an interview costs you the question

  • Assumes a long pause means the collector marked more
  • Forgets that stopping every processor is itself the cost
  • Does not know ReadMemStats stops the world
  • Tunes GOGC in response to a host-induced pause
  • Believes only the garbage collector ever stops the world