skip to content

In `go tool trace`, what does a goroutine's GC mark assist time mean, and which goroutines pay it?

level: middleimportance: should knowfreq 30%

answer

  1. the allocator pays the collector
  2. charged to the goroutine that allocated
  3. proportional to bytes allocated in the cycle
  4. shows on your row, not a collector row
  5. not a pause: the goroutine is running

basics

~20 s

Mark assist is garbage-collection marking work done by an application goroutine itself, charged in proportion to how much it allocated during a collection cycle. Allocation-heavy goroutines pay it, and it lands directly in their latency.

solid answer

~50 s

While a collection cycle is running, Go's collector marks concurrently with your program. If allocation outruns the marker, a goroutine that allocates is made to do a slice of the marking itself before its allocation is granted — that is a mark assist, and `go tool trace` shows it as assist spans on that goroutine's own row rather than on a collector goroutine. The charge is proportional to bytes allocated, so the goroutines that allocate hardest pay the most, which is why a latency spike can land on exactly your hottest handler. It is not a stop-the-world pause: the goroutine is running on a P, just doing collector work instead of yours. You reduce it by allocating less per unit of work — reusing and presizing buffers — or by giving the collector more headroom so cycles are less frequent and less rushed.

code

go · 5 lines
go
for j := range jobs {
	buf := make([]byte, 0, j.Size) // fresh allocation per job
	buf = append(buf, j.Payload...)
	results <- Result{ID: j.ID, Body: buf}
}

go deeper

for a junior

Know that a goroutine which allocates heavily can be made to do garbage-collection work itself, and that this shows up as time attributed to that goroutine in a trace.

for a middle

Explain the mechanism: assist debt proportional to bytes allocated, work performed by the allocating goroutine while it holds a processor, and why that differs from being paused.

for a senior

Show you can separate an assist problem from a pause problem during an incident, and pick between cutting allocation and widening the collector's headroom based on what the trace and a heap profile say.

for a principal

Own the tradeoff between resident memory and tail latency: more headroom buys lower assist cost at the price of footprint, and that is a capacity decision, not a code fix.

## The problem assists exist to solve Go's collector is a concurrent mark-sweep collector: it traces live objects while your goroutines keep running and keep allocating. That creates a race against your program. If the program allocates faster than the marker can mark, the heap keeps growing during the cycle and the collector may not finish before the heap blows past its target. The runtime's answer is to make allocation itself pay: a goroutine that allocates during the mark phase can be required to perform a proportional amount of marking work before it gets its memory. That work is a **mark assist**. ## What you actually see in the trace Open an execution trace with `go tool trace` and the assists appear in two places. On the timeline they are spans on the **application goroutine's own row**, sitting between your function's work, inside a stretch where the collector is active. In the **goroutine analysis** table they inflate the garbage-collection share of that goroutine's lifetime. The surprising part for most people is the attribution. There is no separate goroutine to blame. The handler, worker or parser that allocated is the one that stopped doing your work and started tracing pointers, so the cost shows up as if that code got slower. It did not; it got taxed. ## Charged in proportion to allocation The runtime tracks an assist debt per goroutine in bytes allocated, and requires marking work proportional to that debt. The consequences follow directly: - A goroutine that allocates a few small structs pays essentially nothing. - A goroutine allocating megabytes per request pays a lot, repeatedly. - The cost is not spread evenly across the program. Two goroutines running the same handler can show very different assist time if one hit a large payload. - Assist time is invisible to a naive read of a CPU profile of your handler: the cycles are real and they are on your goroutine, but the code doing them is collector code. Alongside assists, the runtime dedicates background mark workers to roughly a quarter of the logical processors for the duration of a cycle. That is a second, separate way a collection cycle steals latency: it does not just tax the allocator, it also shrinks the pool of processors available to everyone else, which shows up as scheduler wait on unrelated goroutines. When you see p99 latency rise across every endpoint during collection cycles, both effects are usually present. ## Assist is not a pause This distinction matters in an interview and on call. During a mark assist the world is running: your goroutine holds a processor and executes. Nothing is frozen. That is completely different from a stop-the-world phase, where every goroutine is halted. The practical difference is that assist time scales with your allocation rate and lands on specific goroutines, whereas a stop-the-world pause is short, global, and independent of which goroutine you look at. ## What to do about visible assist time 1. **Allocate less per unit of work.** Presize slices and maps when the final size is known, so growth does not allocate repeatedly. Reuse buffers instead of allocating a fresh one per item. Avoid converting between `[]byte` and `string` in a hot path when the conversion copies. 2. **Allocate less often rather than less in total** where you can — one large buffer reused beats thousands of small short-lived ones, both for assist debt and for the marker's work. 3. **Give the collector headroom.** `GOGC` sets how much the heap may grow before the next cycle starts, and `GOMEMLIMIT` sets a soft ceiling the collector will work to stay under. More headroom means fewer cycles, and fewer cycles means fewer windows in which allocation can be taxed at all. Less headroom means the opposite, which is why a tight memory limit can turn a memory problem into a latency problem. 4. **Confirm with a second signal.** A benchmark run with `-benchmem` gives allocations per operation for the code path you suspect, and a heap profile shows which call sites allocate the bytes. The trace tells you assists are hurting; those tell you where the bytes come from. ## The misconception to avoid The common wrong answer is that mark assist is "the GC pausing my goroutine". It is the reverse: the goroutine is being made to *work*, not stopped, and the amount of work is a direct function of how much that goroutine allocated. Reading it as a pause sends people tuning pause behaviour when the real lever is allocation rate.

  • Why does the assist show up on your worker goroutine instead of on a collector goroutine?
    Because the assist is performed by the allocating goroutine itself. The runtime charges it a debt in bytes allocated and makes it do proportional marking work before handing over the memory, so the cycles are spent on that goroutine's row. There is no separate goroutine to attribute them to, which is why the cost reads as your handler suddenly getting slower.
  • During a collection cycle, unrelated goroutines that barely allocate also slow down. Why?
    Assists are only half the cost. For the duration of a cycle the runtime dedicates background mark workers to roughly a quarter of the logical processors, so everyone else is competing for fewer Ps. That surplus shows up as scheduler wait in the goroutine analysis, on goroutines that allocated nothing at all.
  • What is the difference between mark assist time and GC pause time in a goroutine's breakdown?
    Mark assist is time the goroutine spent running collector work while holding a processor; it scales with that goroutine's allocation rate. GC pause is time it was frozen for a stop-the-world phase, which is short, global and roughly the same for every goroutine in the trace. Assists are a tax on allocating; pauses are a freeze on everyone.
  • You cannot easily reduce allocation this quarter. What else lowers assist time?
    Give the collector more headroom so cycles start later and run less urgently: raise GOGC, or raise a GOMEMLIMIT that is set too tight. Fewer cycles mean fewer windows in which allocation is taxed at all. The tradeoff is a larger resident heap, so it converts a latency problem into a memory-footprint decision.

It is a toll on allocation: the more memory a goroutine asks for while a collection is in flight, the more marking it must do before the request is granted.

saying these in an interview costs you the question

  • Calls mark assist a stop-the-world pause
  • Thinks a background collector goroutine does the assist work
  • Says assist time is spread evenly across all goroutines
  • Claims lowering GOGC always reduces assist time
  • Blames the handler's own code when the cycles are collector work