How does Go's runtime decide how much marking work an allocating goroutine owes as a mark assist?
answer
- credit and debt, not a flat cost
- denominated in bytes allocated
- a ratio the collector recomputes mid-cycle
- assist a little extra, bank the rest
- no credit and no work means parking
basics
~20 sThe collector sets an assist ratio of scan work per byte allocated. Every allocation debits the goroutine's assist credit; when that credit goes negative the goroutine must scan off the debt before its allocation proceeds, or park until credit arrives.
solid answer
~50 sThe runtime keeps a per-goroutine assist credit, denominated in bytes. Each allocation debits it by the bytes allocated, scaled by an assist ratio the collector computes from how much scan work is left against how much allocation it expects before marking must finish. While the credit is positive the allocation takes the fast path; when it goes negative the goroutine enters `runtime.gcAssistAlloc` and drains marking work until the debt is repaid. Two things soften this. A goroutine that assists deliberately scans a chunk larger than it owes, banking credit for its next allocations, so the cost arrives in lumps rather than on every call. And background mark workers flush surplus scan work into a shared credit bank that an indebted goroutine can draw on instead of scanning at all. If there is neither credit to take nor work to do, the goroutine parks on the runtime's assist queue until credit is released or the cycle ends — that is the stall.
go deeper
Remember the shape rather than the arithmetic: allocating while a collection is running can make your goroutine do collector work, and how much depends on how many bytes you allocate.
Be ready to walk the accounting out loud — a per-goroutine credit in bytes, debited by each allocation times a ratio the collector recomputes, with scanning performed inline once that credit runs out.
Explain what the credit machinery buys operationally: assists arriving in occasional chunks instead of on every allocation, and a parking path that only bites when the collector genuinely cannot keep up. Connect each to the latency signature you would expect.
The levers you own are the ratio's two halves — scan work your heap creates and bytes allocated per unit of useful work. Framing optimisation targets that way makes teams argue about allocation budgets instead of about collector settings.
### What has to be true by the end of the cycle Concurrent marking has a deadline: the mark phase must finish before the heap grows past the size at which this cycle was meant to complete. The runtime therefore has two quantities it can compare at any instant — how much scan work remains, and how much allocation it expects before the deadline. Their ratio is the price of allocating: **scan work owed per byte allocated**. That number is the assist ratio, and the collector recomputes it as the cycle proceeds. When marking is comfortably ahead, the ratio is small or the assist path never triggers; when marking is behind, the ratio climbs and allocation gets expensive. ### Assist credit, denominated in bytes Every goroutine carries a small piece of accounting state: its assist credit, a signed byte count (the `gcAssistBytes` field on the runtime's goroutine structure). Think of it as a prepaid balance of allocation. Each allocation debits that balance. While the balance stays non-negative, the allocation is as cheap as it ever is — the fast path in the allocator just subtracts and returns. When the balance goes negative, the goroutine is in debt, and the runtime diverts it into `runtime.gcAssistAlloc` before the object is handed back. Note what is being counted. The charge is proportional to **bytes**, not to the number of allocations. A thousand tiny structs and one large buffer of the same total size owe roughly the same assist. This is what makes the mechanism a genuine brake on allocation *volume* rather than a per-call tax. ### Repaying the debt Inside the assist path the goroutine becomes, briefly, a mark worker. It pulls objects from the collector's work queues and scans them — following their pointer fields, marking what they reach — through `runtime.gcDrainN`, which drains a bounded amount of work rather than draining until the queues are empty. Once it has scanned enough to bring its credit back to non-negative, it leaves the assist path and its allocation completes. The important property is that this is *inline*, synchronous, and on the goroutine's own stack. It is not deferred to a helper and it is not a pause: the rest of the program is untouched. ### Over-assisting, and why Entering and leaving the assist machinery costs meaningfully more than scanning a handful of objects. If the runtime charged exactly the debt every time, a hot allocation loop would bounce in and out of the assist path continuously. So an assisting goroutine deliberately scans a fixed chunk of work larger than what it strictly owes, and the surplus is banked as positive credit. The next several allocations then run on the fast path, spending that credit, until the balance goes negative again. The visible effect is that assist cost arrives in occasional lumps rather than as a smooth per-allocation tax — which is another reason the symptom shows up in the tail of a latency distribution rather than in the average. ### Borrowing from the background workers Assists are the collector's fallback, not its primary strategy, and the runtime works hard to keep them rare. Background mark workers that complete more scan work than their own share is accounted for flush the surplus into a shared credit bank. A goroutine that finds itself in debt checks that bank first: if there is enough banked credit, it takes what it needs and its allocation proceeds without scanning anything at all. In a well-paced cycle on a service with moderate allocation, this is the common case — goroutines dip into debt, take credit the background workers have already earned, and never actually scan. Assists only bite when the background workers themselves cannot keep ahead. ### The stall There is one more state, and it is the "stall" half of the picture. A goroutine can owe work, have no credit of its own, find no banked credit, and find no marking work it can perform right now — for example, the work queues are momentarily drained but the cycle has not been declared complete. It cannot proceed (that would let allocation outrun marking) and it cannot pay (there is nothing to do). So it parks on the runtime's assist queue and is woken when credit is flushed by a background worker or when the cycle ends. A goroutine in that state is not burning CPU; it is blocked, and it looks like latency with no matching CPU time anywhere in the process — a genuinely confusing signature if you do not know the mechanism exists. ### Practical consequences - The knobs that matter to you are the ratio's two halves: how much scan work your heap creates (pointer density, live-object count) and how many bytes you allocate per unit of useful work. - Assist cost is not uniform across requests. It is concentrated on whichever goroutine happens to allocate when the collector is furthest behind. - Because credit is per goroutine and banked in chunks, the cost distribution is lumpy by design; averaging it away hides it.
- Why does an assisting goroutine scan more work than it strictly owes?Because entering and leaving the assist path costs more than the scanning does for a small debt. By over-assisting, the goroutine banks positive credit that covers its next several allocations, so a hot allocation loop pays in occasional larger chunks instead of on every single call. It trades a smooth tax for a lumpier but cheaper one.
- What does a background mark worker do with scan work it completes beyond its own share?It flushes the surplus as credit into a shared bank. A goroutine that owes an assist checks that bank first and, if there is enough credit there, takes it and lets the allocation proceed without scanning anything. That is how a well-paced cycle keeps assists rare even under steady allocation.
- What makes a goroutine park rather than assist?It owes work, has no credit of its own, and can find neither banked credit nor marking work it can perform right now — for instance because the work queues are momentarily drained while the cycle is not yet complete. It parks on the runtime's assist queue and is woken when credit is flushed or the cycle ends. That shows up as latency with no matching CPU time.
- Is the assist charge per allocation or per byte?Per byte. A thousand small structs and one large buffer of the same total size owe roughly the same assist. That is what makes the mechanism a brake on allocation volume rather than a per-call tax, and it is why reducing the number of allocations without reducing the bytes rarely helps much.
saying these in an interview costs you the question
- Says every allocation during a cycle scans a fixed amount
- Thinks the charge is per allocation rather than per byte
- Denies that an indebted goroutine can take credit banked by background workers
- Treats the assist ratio as a constant the user configures
- Confuses assist credit with the amount of free heap