In Go, what is a garbage-collection mark assist, and which goroutine pays for it?
answer
- the collector does not mark alone
- allocation is not free mid-cycle
- the runtime drafts the allocating goroutine
- you pay in proportion to bytes allocated
basics
~20 sA mark assist is garbage-collection marking work the Go runtime makes an allocating goroutine do itself. During a collection cycle, allocating memory charges your own goroutine some scanning, so the busiest allocators pay for the collector's backlog.
solid answer
~50 sGo's collector marks live objects concurrently with your program, using background mark workers budgeted at about a quarter of the available Ps. If the program allocates faster than those workers can mark, the runtime makes allocation itself carry the load: a goroutine that allocates while a mark phase is running is charged a proportional amount of scanning and performs it inline, before its allocation returns. That inline work is a mark assist. Because the charge is applied at the allocation site, the goroutine that allocates is the one that pays — a goroutine blocked on I/O, or working over memory it already holds, is never drafted into marking. It is deliberate backpressure: the faster you allocate, the more collector work you are made to do, which slows the allocation down. An assist stalls only that goroutine; it is not a stop-the-world pause.
go deeper
Be ready to say plainly that Go's collector runs alongside your program and that allocating during a cycle can make your own goroutine do marking work. Knowing the name and that the allocator pays is enough at this level.
An interviewer expects the reason the mechanism exists: background marking is capped, so without backpressure a fast allocator could outrun the collector. Say that the charge is applied inside the allocation path, in proportion to bytes allocated.
Show you have seen it in production. Assist time is attributed to your own function in a CPU profile, so a hot path looks slower for reasons that have nothing to do with its logic, and you should be able to separate that from stop-the-world time.
Frame it as a cost the whole process shares: allocation rate anywhere in the binary buys latency for every allocating request path. That makes allocating less a service-wide budget question rather than a local optimisation on the slowest endpoint.
### The collector runs while your program does Go's garbage collector is concurrent. Once a cycle starts, the runtime walks the pointer graph from its roots — every goroutine's stack, the package-level variables, and a few internal structures — and *marks* every object it can reach, so that afterwards anything unmarked can be reclaimed. That walk happens while your goroutines keep running. Nothing is moved; the program never sees the collector except as CPU time and, briefly, as two very short pauses that bracket the cycle. Marking is CPU work, and the runtime funds it with background mark workers whose total share is capped at roughly a quarter of the process's Ps (the logical processors counted by GOMAXPROCS). That cap exists so the collector cannot starve the application it is collecting for. ### The race the cap creates The cap also creates a race the collector can lose. While marking proceeds, the program keeps allocating, and every new allocation both consumes heap and, if it holds pointers, adds edges to the graph that must be walked. Allocate fast enough and the background workers will not have finished marking by the time the heap has grown to the size at which this cycle was supposed to complete. There is no comfortable way to lose that race: the collector cannot abandon the cycle, and the heap cannot simply keep growing. Go's answer is not to raise the collector's CPU share. It is to make allocation carry the work. When a goroutine allocates during the mark phase, the runtime charges it a proportional amount of scanning and makes it perform that scanning inline, in the allocation path, before the new object is handed back. That inline marking is a **mark assist**. ### Charged at the allocation site Because the charge is applied inside the allocator, the goroutine that allocates is the goroutine that pays. A goroutine parked on a channel receive, blocked in a system call, or looping over a slice it already owns allocates nothing and is never drafted into marking. A goroutine that decodes a fresh object graph on every iteration pays a great deal. This is deliberate. It is backpressure aimed exactly at the cause: the faster a goroutine allocates, the more collector work it is made to do, and the more collector work it does, the slower it allocates. The mechanism is self-limiting, and the limiting force lands on the code producing the pressure instead of being smeared evenly across the process. ### What it looks like from the outside Three things surprise people the first time they meet it. 1. **The time is attributed to your own code.** In a CPU profile the assist shows up as `runtime.gcAssistAlloc` and the marking routines beneath it, sitting directly under whichever function performed the allocation. The natural reading is "my decoder got slower", and the decoder did not change at all. 2. **The cost is bursty and uneven.** Assists happen only during a mark phase, and only when background marking is behind. Two identical requests can cost very different amounts depending on when they arrive, which moves the tail of a latency distribution far more than it moves the mean. 3. **The payer is frequently not the culprit.** Allocation rate is a property of the whole process. A cheap request path sharing a process with an expensive one will be charged assists it did nothing to cause. ### It is not a pause A mark assist is not a stop-the-world pause, and conflating the two is the most common mistake on this topic. A stop-the-world pause suspends every goroutine in the process; Go has two short ones per cycle and they exist for separate reasons. An assist blocks exactly one goroutine — the one that allocated — and only until it has scanned what it owes, after which its allocation completes and it carries on. Nothing else in the program stops. That distinction matters practically: assist-driven latency is fixed by allocating less, while pause-driven latency is a different problem with different remedies. ### Why not just run more background workers? Two reasons. First, extra background workers take CPU from the program with no relationship to who is causing the pressure — a quiet goroutine would be slowed down to pay for a noisy one. Second, an uncapped collector can consume an arbitrary share of the machine, which is precisely the failure mode the budget exists to prevent. Charging the allocator ties cost to cause and keeps the background budget honest. ### The one-sentence version Go budgets background marking so it cannot starve the program, and closes the resulting shortfall by billing it to whoever is allocating — so the busiest allocator pays for the marker's backlog.
- Does a goroutine that never allocates during a GC cycle ever get charged for a mark assist?No. The charge is applied inside the allocator, so a goroutine that allocates nothing during the mark phase performs no assist work. It can still be affected indirectly — background mark workers occupy Ps, and the two short stop-the-world phases stop everyone — but it is never billed scanning work of its own.
- Is a mark assist the same thing as a stop-the-world pause?No. An assist blocks only the goroutine doing it, and only until it has scanned the work it owes, after which its allocation completes and it carries on; everything else keeps running. A stop-the-world pause suspends every goroutine in the process. Go has two short ones per cycle and they are a separate mechanism with separate causes.
- Why does the runtime charge the allocator instead of simply running more background mark workers?Extra background workers would take CPU from the program with no relation to who is causing the pressure, and the background budget is deliberately capped so the collector cannot starve the application. Charging the allocator ties the cost to its cause and is self-limiting: heavy allocation slows itself down, which is what stops marking from falling permanently behind.
saying these in an interview costs you the question
- Says all Go garbage-collection work happens on dedicated collector threads
- Calls a mark assist a stop-the-world pause
- Thinks every goroutine is charged equally during a collection
- Believes allocation cost in Go is constant regardless of collector state
- Thinks assists only happen when the program triggers collection explicitly