skip to content

How do you decide a Go service's GOGC value when the platform team wants 30% less memory per replica?

level: principalimportance: should knowfreq 33%

answer

  1. the ask is a budget, not a setting
  2. bring an exchange rate, not an opinion
  3. headroom shrinks, the live set does not
  4. the peak sets the bill, not the average
  5. leave the knob where the budget owner can reach it

basics

~20 s

Treat GOGC as a measured exchange rate between two billed resources. Replay peak load at several values, chart peak heap against collection CPU, pick the knee that still meets the latency objective, and keep the value operator-settable.

solid answer

~40 s

The request is not "lower `GOGC`", it is "spend less memory", and `GOGC` only converts memory into collection CPU at a rate this workload determines. So bring the exchange rate: replay peak load at three or four values and chart peak heap, GC CPU share and p99 for each. The curve usually has a knee, and the honest answer is "30% of the memory costs this much CPU and this much p99" - a trade the budget owners can accept or reject. Two constraints must be said plainly: `GOGC` squeezes only garbage headroom, so if the live set is the problem no value delivers; and it must be sized against the peak live set, not the average. Keep it in the environment so it can change per environment without a release.

go deeper

for a junior

Know that GOGC trades memory for collection CPU in both directions, so "use less memory" is never free — something else on the bill goes up.

for a middle

Be able to produce the comparison: run the same load at a few settings and record peak heap and GC CPU for each, rather than reasoning about which value sounds reasonable.

for a senior

An interviewer expects you to distinguish live data from garbage headroom before promising any saving, and to size the value against measured peak load with headroom for growth.

for a principal

Own the decision as a budget negotiation: bring the exchange rate for this specific workload, name what the knob cannot do, keep the value operator-settable per environment, and attach a review trigger so it does not rot.

## The question behind the question "Cut memory 30%" is a budget instruction. `GOGC` is one of the few levers that can answer it without a code change, and it answers it by **buying memory with CPU**. Both are billed. So the useful contribution from the service owner is not a value but an **exchange rate** for this workload, and then a decision made jointly with whoever holds the budgets. ### What you bring to the room One chart, from one experiment: the service's real peak load replayed at three or four settings, with three numbers recorded at each. | GOGC | peak heap | GC share of CPU | p99 latency | |---|---|---|---| | 50 | ... | ... | ... | | 100 | ... | ... | ... | | 200 | ... | ... | ... | | 400 | ... | ... | ... | The shape is predictable and it is the point: **memory rises linearly with the setting, collection CPU falls hyperbolically**, so there is a knee, and past it you are paying a lot of one resource for very little of the other. The sentence you want to be able to say is "30% less memory costs us this many CPU-seconds per replica and this many milliseconds at p99" — a trade someone can accept or reject, rather than an opinion someone can argue with. ### The two constraints that must be said out loud **1. The knob only squeezes headroom.** The allowance is a proportion of the live set; lowering it never shrinks the live data. If the 30% has to come out of a cache, a big in-memory index, or per-request state that is genuinely reachable, `GOGC` cannot deliver it at any value, and pretending otherwise buys a service that collects continuously and still does not fit. Diagnose which half of the heap is being asked to shrink before agreeing to anything. **2. Size against the peak, not the average.** Peak heap tracks the live set *at the busiest moment*. A value derived from a steady-state average will be wrong by exactly the ratio the peak exceeds it, and the failure lands at the worst possible time. Whatever number is agreed, it is agreed against a measured peak, with headroom for the live set growing. ### Who decides, and how you make that possible The service owner owns the latency objective and the evidence; the platform owner owns the memory budget and can overrule with the numbers on the table. That settlement only works if the value is actually theirs to change: **keep it in the deployment environment, not hard-coded in a runtime call**, so a per-environment change is a config edit rather than a release. Then the decision is reversible at the speed of a restart, which is the real reason to prefer the environment variable. ### What you commit to A defensible outcome is not just a number: - The value, per environment, with the measured peak heap and the headroom assumption written next to it. - A dashboard carrying peak heap and collection CPU share, so the trade stays visible instead of decaying into folklore. - A review trigger — when the live set grows past some threshold, the number is re-derived, because the multiplier applies to the new live set too. - A stated alternative: allocating less lowers both curves at once and makes the whole negotiation smaller, but it costs engineering time. Price it, offer it, and let the people funding the cluster choose. ### The failure mode to name The usual organisational failure is a fleet-wide `GOGC` mandated as a standard. Services differ enormously in allocation rate and live-set size, so one value is generous for some and pathological for others — and the ones it hurts pay in CPU and latency, which are harder to attribute back to the mandate than a memory graph is. If a standard is unavoidable, make it a default with an explicit, evidence-backed opt-out rather than a rule, and make sure the opt-out is a config change no one has to beg for. ### What separates a strong answer Refusing to answer with a number. The strong answer establishes what the knob can and cannot do, produces the exchange rate for this workload, names peak-versus-average as the trap, and hands the decision to the budget owners with the evidence they need to make it — while keeping the value in a place they can change tomorrow.

  • When can lowering GOGC not deliver the memory saving at all?
    When the memory is live data rather than garbage headroom — a large cache, an in-memory index, per-connection state. The setting is a proportion of the live set, so squeezing it approaches the live set as a floor while collection CPU rises without bound. If most of the heap is genuinely reachable, the answer is to hold less data or to size the replica honestly, not to turn the knob.
  • Why is a single fleet-wide GOGC standard usually a bad policy?
    Services differ by orders of magnitude in allocation rate and live-set size, so one value is slack for some and near-continuous collection for others. The services it hurts pay in CPU and latency, which are far harder to trace back to the mandate than a memory chart is. A default with an evidence-backed, easy opt-out gets the consistency without the tail of quiet victims.
  • What do you write down alongside the value you agree on?
    The measured peak heap it was derived from, the load it was measured under, the collection CPU share at that setting, and the condition that triggers a re-derivation — typically the live set growing past a threshold. Without that, the number becomes folklore within two quarters and nobody can say whether it is still right.

saying these in an interview costs you the question

  • Answers with a recommended number and no measurement
  • Promises a memory saving without checking whether the heap is live data or headroom
  • Derives the value from average heap instead of peak heap
  • Hard-codes the setting in application code so the budget owner cannot change it
  • Accepts a fleet-wide mandate without an opt-out path
  • Frames it as memory versus CPU only, ignoring the latency objective