A Go ingestion service's live heap climbs under steady input — how do you find the unbounded backlog?
answer
- rising RSS is not proof of anything
- watch the heap after each collection
- inuse_space, and diff two profiles
- look at where the producers are parked
- no blocked sender means no backpressure
basics
~20 sConfirm the growth is live data with GODEBUG=gctrace=1, then take a heap profile with inuse_space to see which structure holds the bytes, and check the goroutine profile: if no producer is parked in a channel send, the backlog is absorbing work that a bounded channel should have stalled.
solid answer
~50 sFirst separate real growth from collector lag: run with `GODEBUG=gctrace=1` and watch the heap size reported at the end of each cycle. If it rises monotonically under constant input, the live set is growing. Then take a heap profile — `go tool pprof -inuse_space` against `/debug/pprof/heap` — and look at what retains the bytes; an unbounded backlog shows up as one growing slice, map or oversized channel of unprocessed items, allocated on the intake path. The confirming signal is the goroutine profile: in a correctly bounded pipeline under overload, producer goroutines are parked in a channel send. If nothing is parked and the heap is still growing, some intermediate structure is absorbing the rate gap instead of transmitting it. The fix is to bound that structure so the send blocks; raising `GOMEMLIMIT` only converts an out-of-memory kill into a collector-bound service.
code
text · 3 linesGODEBUG=gctrace=1 ./ingest 2>gc.log # is the live heap really growing?
go tool pprof -inuse_space http://localhost:6060/debug/pprof/heap
curl -s 'http://localhost:6060/debug/pprof/goroutine?debug=2' | head -40go deeper
Know that Go gives you a heap profile and that GODEBUG=gctrace=1 prints a line per collection, and that memory held by items still waiting to be processed cannot be collected.
Walk the sequence: confirm live-heap growth from the collection trace, find the retaining allocation site with an inuse_space profile, and name the structure that is holding unprocessed items.
Demonstrate the discriminating step — reading the goroutine profile for producers parked in a channel send — and explain why its absence proves the pressure is being stored rather than transmitted. Say how you would verify the fix.
Own the aftermath: what the intake ceiling should be, what the service does when it is reached, and which signal the team pages on so this is caught as saturation rather than as an out-of-memory kill.
## Step 1 — is the live heap really growing? Rising RSS alone proves nothing: Go's collector is not compacting and does not immediately return memory to the OS, so a service with a bursty allocation profile can hold a large resident set with a small live set. The cheap discriminator is `GODEBUG=gctrace=1`, which makes the runtime print one line per collection to standard error, including the heap size at the end of the cycle. Under constant input, a healthy service shows those end-of-cycle sizes oscillating around a stable level. A backlog shows them stepping upwards, cycle after cycle, forever. `runtime/metrics` gives the same picture programmatically if you would rather not restart with a GODEBUG setting. While you are there, note the collection frequency. A growing live set makes the collector run more often over a bigger graph, so CPU spent in GC climbs alongside — which is why these incidents usually present first as latency, not as memory. ## Step 2 — what holds the bytes? Take a heap profile and read it in `inuse_space` mode, which reports memory currently held rather than cumulative allocation: ``` go tool pprof -inuse_space http://localhost:6060/debug/pprof/heap ``` You are looking for a single allocation site that dominates and grows between two profiles taken minutes apart. A backlog has a recognisable shape: the retaining site is on the *intake* path — the decode loop, the batching helper, the enqueue wrapper — and the retained objects are the event or job type itself. Taking two profiles and diffing them makes the growth unambiguous. ## Step 3 — backlog or leak? The two look identical on a graph of heap over time, and they are fixed differently, so this distinction is the whole diagnosis. A **leak** retains objects whose work is finished: completed requests kept in a map nobody deletes from, a cache with no eviction, a slice of results appended to forever. Its size tracks *total traffic served*. A **backlog** retains work that has not happened yet. Its size tracks *how far behind the consumers are* — arrival rate minus drain rate, integrated over time. Two more tells: it disappears the moment input pauses, because the consumers catch up; and its growth rate changes with consumer speed, not with traffic volume. ## Step 4 — the goroutine profile decides it This is the step that turns a guess into a conclusion. Fetch the goroutine profile (`/debug/pprof/goroutine?debug=2` gives full stacks) and look at where the producers are. In a correctly bounded pipeline under overload, the producing goroutines are parked in a channel send. That is backpressure working: the stall is being transmitted. If instead every producer is running happily, the heap is growing, and the consumers are the slow part, then something between them is absorbing the rate gap — an unbounded staging slice, a per-key map of pending items, a channel created with an enormous capacity, or a goroutine spawned per item so that the queue has become "the scheduler's run queue". The goroutine count itself is worth reading: a steadily climbing count under steady input means the unbounded thing is goroutines rather than a slice, which is the same failure wearing a different costume. ## Step 5 — fix it where the pressure should have been felt Bound the intake, so that the producer's send blocks and the stall reaches the source. Compute the ceiling you are choosing: capacity times item size, plus the workers' in-flight items. Then decide what happens when it is full — wait, wait with a deadline, or shed with a counter — because that is now a real code path rather than an implicit "grow". ## What not to do Raising `GOMEMLIMIT` is the reflex, and it is the wrong lever here. It is a *soft* limit: as the live heap approaches it the collector runs more and more aggressively to stay under it, and since a growing backlog cannot be collected, the service degrades into a collector-bound crawl instead of dying quickly. The runtime does cap GC's CPU share so the process does not livelock entirely, but you have converted a clear memory failure into a murky latency failure. Lowering `GOGC` has the same character. Both are legitimate tuning knobs for a service whose live set is stable and too large; neither bounds anything that is genuinely unbounded. Adding more workers is worth testing but is not a fix either: if the consumers are limited by a downstream dependency, more of them just move the queue somewhere else. ## Verifying the fix Reproduce with a producer deliberately faster than the consumer, and watch for the heap to plateau instead of climb, the goroutine profile to show parked senders, and either the enqueue block time or the shed counter to become non-zero. All three changing together is the proof that the pressure is now being transmitted rather than stored.
- How do you distinguish this from a classic leak in the heap profile?Look at what is retained and how it scales. A leak holds finished work and grows with total traffic served; a backlog holds unprocessed items on the intake path and grows with the gap between arrival and drain rate. Pause the input: a backlog drains and the heap falls, a leak does not move.
- Why is raising GOMEMLIMIT the wrong response here?It is a soft limit, so as the live heap approaches it the collector simply works harder to stay under it. A growing backlog is live and cannot be collected, so you trade an out-of-memory kill for a collector-bound service with degraded latency. It is a tuning knob for a stable-but-large live set, not a bound on something genuinely unbounded.
- The goroutine count is climbing steadily too. What does that suggest?That the unbounded thing is goroutines rather than a slice — typically a `go` per incoming item, so the run queue has become the backlog. It is the same failure: work arriving faster than it completes, with no construct forcing the producer to wait. The fix is the same shape, a fixed set of consumers behind a bounded intake.
saying these in an interview costs you the question
- Treats rising RSS as proof the live heap is growing
- Jumps to raising GOMEMLIMIT or lowering GOGC as the fix
- Never checks whether any producer is parked in a send
- Reads alloc_space instead of inuse_space for retained memory
- Calls it a leak and hunts for a missing delete
- Adds workers without checking the downstream bottleneck