A report service runs with GOGC=25 to hold memory down, and now burns a third of its CPU collecting during its hourly spike. How do you confirm that and choose a better value?
answer
- measure the spike, not the process lifetime
- cumulative fractions hide a spiky workload
- a quarter of headroom for an allocation-heavy job
- assist frames on the worker goroutines
- same spike, two settings, two curves
basics
~20 sMeasure collection CPU over the spike window alone, not since process start: diff the runtime's GC CPU-seconds metric across the spike. Then replay that spike at higher GOGC values and take the largest peak heap you can afford.
solid answer
~50 sFirst scope the measurement to the spike. `runtime.MemStats.GCCPUFraction` is cumulative since start-up, so an idle fifty minutes dilutes it into uselessness; instead sample `runtime/metrics`' `/cpu/classes/gc/total:cpu-seconds` at the start and end of the spike and diff it, and take a CPU profile over that window only — `runtime.gcBgMarkWorker` frames plus `runtime.gcAssistAlloc` inside the report goroutines confirm the diagnosis. `GOGC=25` means the heap may only grow a quarter past the live set, so a workload that builds and discards a large working set per report restarts collection almost continuously and charges the allocating goroutines mark assists. Then run the same spike as a load test at 25, 100 and 400, and plot GC CPU share against peak heap; take the highest value whose peak still fits the memory the service is given, with headroom. `debug.SetGCPercent` lets you flip settings on a live instance without a redeploy.
code
go · 5 linesfunc gcCPUSeconds() float64 {
s := []metrics.Sample{{Name: "/cpu/classes/gc/total:cpu-seconds"}}
metrics.Read(s)
return s[0].Value.Float64()
}go deeper
Recall that a very low GOGC makes the collector run far more often, which costs CPU; the memory graph looking better does not mean the service got cheaper.
Explain why a small growth allowance is especially bad for a workload dominated by short-lived temporaries, and name the runtime numbers you would read to check it.
An interviewer expects the measurement discipline: scope GC CPU to the spike window, confirm with a profile taken during it, then compare settings on the same replayed load and choose against peak heap with headroom.
Be ready to say what the fix costs and who pays: a higher setting spends memory the cluster budget owns, while allocating less spends engineering time — and to bring both numbers to the people who own each budget.
## The shape of the problem A scheduled report generator idles for most of the hour and then does everything at once. Each report builds a large working set of short-lived temporaries and throws it away. Someone set `GOGC=25` because the memory graph looked alarming; memory duly flattened, and the CPU cost was bought back silently. Now the spike takes a third of every core in collection and reports finish late. ### Why 25 is pathological for this workload `GOGC=25` means: allow the heap to grow only 25% past the live set before collecting again. For a service whose live set is small but whose **allocation rate during the spike is enormous**, that allowance is consumed almost instantly. Cycles start back-to-back, and when allocation outruns the marker the runtime charges the **allocating goroutine** mark-assist work — so the cost does not appear as background GC, it appears as report generation being slow. The knob squeezed the *garbage headroom*, which is exactly the part of the heap this workload needs most: the temporaries are the workload. ### Step 1 — measure GC CPU for the spike, not for the process The trap here is scope. `runtime.MemStats.GCCPUFraction` is documented as the fraction of available CPU used by the collector **since the program started**. On a service that idles fifty minutes an hour, that number is averaged into meaninglessness, and it will not move much even when the spike is entirely GC-bound. Sample a windowed counter instead: ```go func gcCPUSeconds() float64 { s := []metrics.Sample{{Name: "/cpu/classes/gc/total:cpu-seconds"}} metrics.Read(s) return s[0].Value.Float64() } ``` Read it immediately before the spike begins and immediately after it ends; the difference is GC CPU-seconds **for that window**, and dividing by (wall seconds x `GOMAXPROCS`) gives the share of available CPU the collector took during the spike. `/gc/cycles/total:gc-cycles` differenced the same way tells you how many cycles ran, which is the other half of the story: an implausible cycle count is the signature of a too-low `GOGC`. ### Step 2 — confirm with a profile taken during the spike A CPU profile collected over the spike window (not a default profile of a mostly idle process) should show: - `runtime.gcBgMarkWorker` — background marking, the collector's own CPU. - `runtime.gcAssistAlloc` on the stacks of the report goroutines — allocation being taxed to keep up. This is the frame that explains why *latency* moved, not just utilisation. If neither appears in quantity, the collector is not your problem and the tuning conversation ends there. ### Step 3 — run the experiment, do not argue about it Replay the same spike shape as a load test at a few settings — say 25, 100, 400 — and record, for each: GC CPU seconds over the window, wall time for the batch, and **peak heap**. Two curves on one chart, GC CPU share and peak heap against `GOGC`, and the decision makes itself: the CPU curve falls steeply and then flattens, the memory curve rises linearly, and you take the last value before the CPU curve goes flat, provided the peak heap still fits. `debug.SetGCPercent` is what makes this cheap: an operations endpoint that flips the value lets you run consecutive spikes at different settings on one instance, and it returns the previous value so the instance can be put back. ### Step 4 — choose against the peak, not the average The number that matters is peak heap **during the spike**, because that is when the live set and the allowance are both largest. A value chosen from the idle hour will be wrong by the multiple that the spike's live set exceeds the idle one. Leave headroom for the live set growing over the next few months. ### The durable fix, stated honestly The knob is the lever you have today; it does not make the service allocate less. If the same temporaries are rebuilt per report, reusing them is a code change that lowers both curves at once and makes the tuning question smaller. Say that in the review — but do not let it block shipping a value that stops the bleeding this week. ### What a strong answer sounds like Name the scoping error first (cumulative versus windowed), then the mechanism (a quarter of headroom for an allocation-dominated workload), then the experiment, then the decision rule against peak heap. A candidate who jumps straight to "set it to 400" has skipped every step that makes the number defensible.
- Why is runtime.MemStats.GCCPUFraction a poor measurement for this service?It reports the collector's share of available CPU since the process started, so fifty idle minutes average away a spike that is genuinely GC-bound. Any conclusion drawn from it understates the problem. Windowed counters differenced across the spike, or a profile taken only during it, measure the thing you actually care about.
- What tells you in a CPU profile that the collector's cost is landing in request latency rather than in the background?Mark assist frames on the working goroutines' own stacks: the allocating goroutine is being made to do collection work because allocation is outrunning the marker. Background marking alone shows up in dedicated worker frames and steals throughput without directly stretching any one request; assists stretch the request that allocates.
- How would you compare two GOGC settings on the same spike without redeploying?Expose an authenticated operations endpoint that calls debug.SetGCPercent with the new value and keeps the returned previous value. Run one spike at each setting on the same instance and record windowed GC CPU-seconds and peak heap for each. Same code, same data, same machine — the only variable left is the setting, which is what makes the comparison worth anything.
- If raising GOGC fixes the CPU but the peak heap is then too large to fit, what is left?Reduce what is allocated. Reusing per-report buffers and avoiding rebuilding the same temporaries lowers both curves at once, so a smaller allowance costs less CPU. That is a code change with a real price, so scope it against the alternative of paying for the memory, and pick with the measured numbers rather than by preference.
saying these in an interview costs you the question
- Quotes GCCPUFraction from a long-running idle process as proof
- Profiles the whole hour instead of the spike window
- Suggests a GOGC value with no measurement behind it
- Assumes a low GOGC only costs memory, never CPU or latency
- Sizes the setting from average heap rather than peak heap
- Blames stop-the-world pauses for a throughput loss during the spike