skip to content

Why does Go's GC pacer start a cycle before the heap reaches the goal GOGC sets?

level: middleimportance: should knowfreq 44%

answer

  1. the goal is a finish line, not a start line
  2. the program allocates during the mark
  3. estimate the work, estimate the rate
  4. aim to finish exactly on the goal
  5. corrects itself from last cycle's error

basics

~20 s

Marking runs concurrently while the program keeps allocating, so Go's pacer triggers early enough that marking finishes just as the heap reaches the goal. Triggering late overshoots the goal; triggering early wastes CPU collecting a half-full heap.

solid answer

~50 s

The goal GOGC sets is where the heap should be when a cycle *ends*, not where it starts. Go marks concurrently with the running program, and the program keeps allocating throughout the mark, so the runtime must begin at a trigger point below the goal. The pacer picks that point by estimating two things from the previous cycle: how much scan work marking will take - live heap plus goroutine stacks and globals - and how fast the program is allocating, assuming collection gets roughly 25% of `GOMAXPROCS` while it marks. It is a feedback loop: each cycle it compares where the heap actually landed against the goal and corrects the trigger. Triggering late overshoots the goal, with mark assists pulling allocating goroutines in to help; triggering early pays for a cycle that reclaims less.

go deeper

for a junior

Know that Go marks the heap while your code keeps running, so a collection has to begin before the heap reaches its goal rather than at it. Naming that gap is enough at this level.

for a middle

Be able to state the two estimates behind the trigger - the scan work ahead and the program's allocation rate - plus the CPU share assumed for marking, and say what triggering too late or too early costs.

for a senior

Talk about it as a feedback controller: the pacer measures where the heap actually landed against the goal and corrects, so an allocation spike overshoots once and then settles. Judge pacing over several cycles, not one.

for a principal

Frame the trigger as a control problem balancing a CPU budget against memory headroom, and insist on knowing which of the two a service is genuinely short of before anyone proposes moving a knob.

## Two different heap sizes Talking about Go's pacer requires keeping two numbers apart. - The **goal** (sometimes called the heap target) is the heap size the runtime is aiming to *end* the next cycle at. `GOGC` sets it as a growth percentage over the live heap: at the default `GOGC=100`, roughly double. - The **trigger** is the heap size at which the cycle actually *begins*. It always sits below the goal. If the two were the same number, the collector would start marking at the moment the heap arrived at its ceiling and would then keep allocating past it for the whole duration of the mark. The goal would be an announcement of failure rather than a target. ## Why the gap has to exist Go's collector is concurrent: marking runs alongside the application goroutines rather than pausing them for the duration. That is the property the whole design is built to protect - only two very short stop-the-world phases per cycle. The price is that the mutator keeps allocating while the marker works, and every byte it allocates during the mark lands on the heap the cycle was supposed to be capping. So the runtime has to start early enough that, by the time marking completes, the extra allocation has taken the heap to - and ideally not past - the goal. Finishing the cycle exactly at the goal is the pacer's objective function. There is a second reason the gap cannot simply be made enormous "to be safe". Go's collector is non-generational and non-compacting: there is no cheap young-generation pass to fall back on, and no fallback full compaction either. Every cycle scans everything reachable. A cycle started far too early scans the same live set and reclaims less garbage for it, which is pure CPU waste. ## What the pacer estimates To place the trigger, the runtime needs to predict how long marking will take and how much will be allocated in that time. It works from measurements the previous cycle produced: 1. **Scan work.** How many bytes marking will have to traverse: the live heap plus the roots - every goroutine's stack and the global variables - which must be rescanned each cycle. 2. **Allocation rate.** How fast the program has been putting bytes on the heap. 3. **The CPU share it may use.** The collector is designed to take about 25% of `GOMAXPROCS` while marking. That fixes how quickly the estimated scan work can be got through. From those it computes how much heap growth will happen during the mark, and subtracts it from the goal. That difference is the trigger. ## It is a controller, not a formula None of those inputs is exact. The live set changes, the allocation rate changes, and the amount of work varies. So the pacer runs as a feedback loop: at the end of every cycle it compares where the heap actually finished against the goal it was aiming for, and uses that error to adjust the next trigger. A workload that suddenly starts allocating twice as fast will overshoot once, then be met with an earlier trigger on the following cycle. This is why pacing behaviour is best judged over several cycles rather than one. A single overshoot after a traffic change is the controller doing its job; a persistent one means an input assumption is wrong. ## What going wrong in each direction looks like **Too late (trigger too high).** The heap sails past the goal before marking is done - an overshoot. The runtime's backstop is mark assists: goroutines that allocate during a cycle are made to do a share of the marking work themselves, which slows allocation until the marker catches up. Latency shows up in allocating code paths rather than in a pause. **Too early (trigger too low).** The cycle begins over a heap that is not yet full of garbage. It costs a full scan of the live set to reclaim relatively little, GC CPU share rises, and application throughput falls without any memory being saved. ## Where this leaves the mental model `GOGC` decides *where the finish line is*; the pacer decides *when to set off* so the collector arrives there on time, given how fast the program is running and how much CPU the collector is allowed. Keeping those two responsibilities separate is what makes the rest of the collector's behaviour legible: a heap sitting comfortably below its goal may already be mid-cycle, and a heap slightly above it is usually a cycle that started a touch late rather than anything broken.

  • What happens if the heap overshoots the goal anyway?
    Nothing catastrophic: the cycle completes over a larger heap, mark assists slow allocating goroutines so marking can catch up, and the pacer feeds the error back into an earlier trigger for the next cycle. A single overshoot after a traffic change is normal; a persistent one means an estimate is systematically wrong.
  • Why not just stop the world at the goal and mark everything then?
    The pause would be proportional to the live heap, which is exactly the property Go's design refuses to accept. Marking concurrently keeps stop-the-world phases short and independent of heap size, and the price of that choice is that the program keeps allocating during the mark - which is what the trigger gap pays for.
  • What CPU share does the pacer assume it will get for marking?
    About 25% of GOMAXPROCS. That budget is what converts an estimate of scan work into an estimate of elapsed mark time, and therefore into how much heap growth to leave room for. If the collector could use more CPU it could start later; if it used less it would have to start earlier.

saying these in an interview costs you the question

  • Says the cycle starts exactly when the heap reaches the goal
  • Thinks the trigger is a fixed fraction of the goal
  • Assumes marking is instantaneous so timing does not matter
  • Claims the heap can never exceed the goal
  • Believes GOGC sets the trigger rather than the goal