How do you decide what GOMEMLIMIT a service gets, and whether being killed beats thrashing?
answer
- a budget split, not a single number
- which failure do you prefer
- restart cost decides the posture
- per deployment, never compiled in
- someone else pays for headroom
basics
~10 sTreat GOMEMLIMIT as splitting a fixed allowance between Go-managed and other memory, then choose a failure: a soft ceiling buys degradation instead of a kill, which pays only when a restart is expensive.
solid answer
~50 sThere are two decisions and only the first is arithmetic. Sizing is a budget split: measure the non-Go footprint, reserve it plus a margin out of the allowance, and derive the ceiling per deployment rather than compiling it in, since the same binary runs under different allowances. The second decision is the real one - a soft ceiling exchanges an out-of-memory kill for degradation, and that is an improvement only if a slow instance beats a restarted one. For a stateless request server a fast kill and replacement is usually cleaner. For an in-memory aggregation service a restart discards the window and the rebuild may cost more than the slowdown. Then make the degraded mode visible: alert on GC CPU share and time spent near the ceiling, or you have traded a loud failure for a silent one.
code
go · 8 lines// One binary, many allowances: never compile the ceiling in.
if v := os.Getenv("MEMORY_ALLOWANCE_BYTES"); v != "" {
if total, err := strconv.ParseInt(v, 10, 64); err == nil && total > 0 {
// Reserve 15% for the binary, cgo and mapped files.
debug.SetMemoryLimit(total / 100 * 85)
}
}
// No allowance known: leave the default in place rather than guess.go deeper
The point to remember is that the ceiling has to be smaller than the container's allowance because other memory shares that allowance, and that it should come from deployment configuration rather than being fixed in the code.
Be able to describe the sizing procedure concretely: measure the non-Go footprint, reserve it plus a margin for allocation bursts, and confirm the working set fits in what remains before committing to a value.
Argue the failure trade for a specific service rather than in general. Show that you know a soft ceiling buys degradation, and that degradation is only an improvement when a restart is expensive and slow instances do not amplify load.
Own the whole proposal: the measurement, the posture, the alerting that makes the chosen failure visible, and the capacity ask. Expect to be overruled by whoever pays for fleet-wide headroom or carries the pager, and prepare the evidence that answers them.
## Two decisions wearing one name Asking "what should the ceiling be" hides two questions. The first is a budget split and has a defensible numeric answer. The second is a policy choice about which failure the service should have, and it has no numeric answer at all — which is why it belongs to someone who can be overruled. ## Decision one: the split The container's allowance has to cover everything: the mapped binary, anything allocated through cgo, memory-mapped files, and the memory the Go runtime manages. Only the last of those is what the ceiling governs. So the procedure is: 1. **Measure the fixed cost.** Run with a deliberately small ceiling so the Go part is bounded, and take the difference between the reported process footprint and the runtime-managed total. That is what you must reserve. 2. **Reserve a margin.** Collection is concurrent, so allocation bursts can carry the total briefly above the target. A ceiling flush against the reserve turns a normal transient into a kill. 3. **Give the rest to the runtime**, and verify by soak test that the real working set fits inside it with room for garbage. The part teams get wrong most often is step zero: **the ceiling belongs to the deployment, not the binary.** The same image runs in a small instance and in a large batch instance, so a hardcoded value is wrong everywhere except one place. Derive it at startup from the allowance the platform gives the process, or take it from configuration; if it cannot be determined, fall back to the default behaviour rather than to a guess. ## Decision two: which failure do you want Setting a ceiling does not remove the failure — it changes its shape. Without one, a service that outgrows its allowance dies abruptly and is replaced. With one, it survives and gets slow: the collector takes an increasing share of the CPU as live data approaches the ceiling, and throughput falls while the process stays up. Neither is universally better, and the deciding factor is **what a restart costs**: - **A stateless request server** loses nothing on restart. Its callers are already built to retry another instance, so a kill is cheap and a degraded instance that keeps accepting requests it cannot serve in time is comparatively expensive. Here, less ceiling and faster failure is a defensible posture. - **An in-memory aggregation service** loses the window it has accumulated. Restarting may mean replaying a backlog, re-reading a large input, or producing a gap in output that downstream consumers notice. Degradation that lets it finish the window is worth real latency. Here, the ceiling earns its place. There is a third case worth naming: services whose degradation propagates. If slow responses cause callers to retry, the degraded instance can generate more load on the system than a dead one, and the calculation flips regardless of restart cost. ## The interaction people forget A thrashing process usually fails its health check anyway. If the platform's health check has a short timeout, the degraded instance gets terminated too — so the team pays the latency damage *and* takes the restart, which is the worst of both. Choosing degradation as a posture therefore obliges you to look at what health checks do under GC pressure, and to decide deliberately whether a slow instance should be considered unhealthy. That decision belongs with the ceiling, not separately from it. ## Making the chosen failure visible A kill is loud: restarts are counted, and someone notices. Degradation is quiet. If you adopt a ceiling, you must add the signal that replaces the one you gave up: - the share of CPU spent in garbage collection, alerting when it rises above a normal band; - how close runtime-managed memory runs to the ceiling, over time, rather than instantaneously; - the service's own latency, which is the thing you actually care about. Without those, adopting a ceiling means incidents stop appearing in the restart graph and start appearing in customer complaints. ## Who owns it, and who can overrule The service owner proposes the number and the posture, because they know the working set and what a restart costs. Two other parties have standing to overrule. The first is whoever owns the memory budget across the fleet. Headroom is not free — a reserve multiplied across every replica of every service is a real bill, and "give it more room" competes with other uses of the same capacity. The second is whoever owns the paging policy. A posture of "degrade rather than die" changes what wakes people up and what the on-call runbook says. If a team chooses degradation without adding the alerting to see it, the person carrying the pager has legitimate grounds to reject the change until the signal exists. The strongest version of this proposal, therefore, is not a number. It is: here is the measured fixed cost, here is the working set from the soak test, here is the failure we are choosing and why a restart is or is not cheap for this service, here is what we will alert on to see it, and here is the capacity we are asking for.
- A team proposes rolling a fleet-wide GOMEMLIMIT default across every service. What would you push back on?A single default cannot express the split, because the non-Go footprint differs per service, and it silently imposes one failure posture on services whose restart costs differ enormously. A better fleet-wide default is a mechanism — derive the ceiling from the allowance, reserve a measured fixed cost, emit the GC-pressure signals — leaving the number and the posture to each service owner.
- How does an orchestrator's health check interact with the posture you choose?A thrashing process often fails a short-timeout health check, so a team that chose degradation gets the restart anyway plus the latency damage first. If you intend a service to degrade rather than die, you must decide deliberately whether a slow instance counts as unhealthy and tune the check accordingly; otherwise the posture exists only on paper.
- What evidence would make you accept a request for a larger memory allowance rather than a lower ceiling?A soak test showing that cycle spacing collapses and throughput falls off a knee above the current working set, with the live set — not garbage — occupying the ceiling. That demonstrates the collector has nothing left to reclaim, so no ceiling value helps. Combined with a measured non-Go reserve, it turns a capacity request into a documented shortage rather than a preference.
- Why should the ceiling not be compiled into the binary?The same image runs under different allowances — a small always-on instance and a large batch run — so any fixed value is wrong in most of them. Deriving it at startup from the allowance the process is given keeps one binary correct everywhere, and makes the number reviewable in deployment configuration where the allowance itself is set.
saying these in an interview costs you the question
- Picks a round number with no measured non-Go reserve
- Hardcodes the ceiling in the binary for every deployment
- Treats a ceiling as strictly safer with no cost named
- Adopts degradation without alerting on GC pressure
- Ignores that a thrashing instance still fails health checks