A Go pipeline stage decoding JSON into map[string]any keeps the collector busy while its live heap stays flat. How do you use the heap and allocs profiles to find and cut the allocation volume?
answer
- the live view is flat by design
- pacing follows allocation, not retention
- count objects as well as bytes
- cumulative down to your own frame
- decode into a type, not a map
basics
~20 sOpen the profile at alloc_space and alloc_objects, not the live view: a flat live heap is what churn looks like. Then cut objects per event - a concrete struct instead of map[string]any, json.RawMessage for unread fields - and re-measure.
solid answer
~50 sA flat `inuse_space` is exactly what churn looks like, so that view will tell you nothing; open the same profile with `-sample_index=alloc_space`, then `-sample_index=alloc_objects`. Collection frequency tracks allocation volume rather than live size, which is why the collector never rests while memory looks fine. Read cumulative down from the top until you reach your own frames — the flat leader will be inside `encoding/json`, and that is not where you fix anything. Decoding into `map[string]any` is usually the finding: every event becomes a map plus one heap-boxed value per field, so `alloc_objects` explodes. The fixes are all about producing fewer objects for the same output: decode into a concrete struct, keep fields the stage does not inspect as `json.RawMessage`, reuse one `json.Decoder` over the stream instead of unmarshalling per event, and reuse scratch buffers. Prove it with a benchmark reporting allocations per operation and a fresh profile over the same input.
code
go · 13 linesfunc handle(r io.Reader, out chan<- event) error {
dec := json.NewDecoder(r)
for {
var m map[string]any
if err := dec.Decode(&m); err != nil {
if errors.Is(err, io.EOF) {
return nil
}
return err
}
out <- toEvent(m)
}
}go deeper
Take away the core fact: a program can be perfectly steady in memory and still allocate enormously, and the cumulative columns of a memory profile are where that shows up.
Explain why decoding into map[string]any costs a map plus a boxed value per field, and how decoding into a struct or leaving fields as raw bytes removes those objects.
Walk the whole loop out loud: pick the cumulative view, read down to your own frames, name the object-count finding, choose fixes that preserve output, and verify with both a benchmark and a fresh profile.
Decide how much of the stage is worth reshaping for the allocation you measured, what the change costs in readability and schema coupling, and what measurement the team must repeat so the churn does not come back next quarter.
## The symptom and why it confuses people A stage of a streaming pipeline reads events off a connection, decodes each one, and hands a value to the next stage. Memory usage is steady and unremarkable. Yet the process burns a large fraction of its CPU inside the garbage collector, and throughput per core is poor. The instinct is to hunt for a leak, and the instinct is wrong. Go's collector is paced by **allocation**: it starts a cycle when the heap has grown by a proportion of the live heap since the last cycle. A stage that allocates enormously but retains nothing therefore triggers cycle after cycle while `inuse_space` sits flat. Live size tells you how big each cycle's marking job is; allocation rate tells you how *often* you pay for one. So the first move is a view change, not a new tool. ## Step 1: read the cumulative view Take a heap profile as usual, then open it twice: ``` go tool pprof -sample_index=alloc_space ./stage heap.out go tool pprof -sample_index=alloc_objects ./stage heap.out ``` `alloc_space` names the stacks that produced the bytes; `alloc_objects` names the stacks that produced the *count*. The gap between them is the diagnosis. Twelve gigabytes across eight hundred million objects is a very different problem from twelve gigabytes across four thousand objects: the first is per-item allocation and boxing, the second is buffer sizing. Read **cumulative**, not flat. The flat leader in an allocation profile is almost always a shared standard-library routine that everyone calls. It is true and useless. Walk down the cumulative ordering until you hit the first function your team owns; that is the frame where a decision was made. ## Step 2: understand why map[string]any is expensive Decoding an object into `map[string]any` allocates, per event: * the map itself, plus its backing storage, which grows as keys are inserted; * a key string per field; * an interface value per field whose payload usually has to be heap-allocated, because a decoded number, string or nested map does not fit in the interface word; * a nested map (and its whole chain again) per nested object, and a `[]any` per array. A five-field event with one nested object can easily cost a dozen-plus objects. Multiply by the event rate and `alloc_objects` is the column that screams. Decoding the same event into a struct with typed fields allocates the strings it must keep and little else — the numeric and boolean fields land in the struct's own memory with no per-field boxing. ## Step 3: cut objects without changing output The constraint that makes this a real engineering problem is that the stage must keep producing exactly what it produced before. Ordered by payoff: 1. **Decode into a concrete type.** Define the struct for the fields the stage actually uses. This removes the map, the per-field boxing and most of the key strings in one change. 2. **Leave what you do not read undecoded.** A field typed `json.RawMessage` is captured as bytes and not turned into objects. For a pass-through stage that inspects two fields and forwards the rest, this is the single biggest win. 3. **Stream with one decoder.** `json.NewDecoder(r)` and a loop of `Decode` reuses internal state across events, instead of reading each event into a fresh `[]byte` and calling `json.Unmarshal` on it. 4. **Reuse the destination.** Decoding into the same struct variable each iteration, and resetting slices with `s = s[:0]` rather than allocating new ones, keeps steady-state allocation near zero for the loop scaffolding. 5. **Reuse large scratch buffers** through a per-worker buffer or a `sync.Pool` — but only after the profile says buffers are the cost, because pooling small objects usually adds complexity for nothing. What you should *not* do is start with pooling or with turning knobs. The profile has named the allocation sites; fix the ones it named. ## Step 4: prove it Two measurements, not one: * A benchmark over a fixed corpus with allocations-per-operation reported (`-benchmem`, or `testing.B.ReportAllocs`). This is the fast, repeatable signal while you iterate, and it is exact if you raise the profiling resolution for that run. * A fresh profile from the real stage under real load, opened at `alloc_space` and `alloc_objects`, over a comparable window. Confirm the absolute totals fell and that the stack you fixed dropped in the ordering rather than merely being overtaken by something else. Then check the outcome you actually cared about: collector CPU share and throughput. If allocation volume halved and collector CPU did not move, your model of the problem was wrong and you should re-derive it before changing more code. ## Traps worth naming * Declaring victory on a percentage. Percentages shift when other work changes; compare absolute bytes and object counts over a fixed input. * Reading `inuse_space` because it is the default view and concluding there is nothing to find. * Blaming the decoder. The decoder allocates what you asked it to build. * Fixing a stack that is large in bytes but rare, while the millions-of-objects stack sits second on the list.
- Why will inuse_space not help you here?Because the objects are already gone. The live view shows what is still reachable at the snapshot, and a stage that retains nothing has a small, stable live set no matter how much it churns. The evidence lives in the cumulative columns, which count everything ever allocated, collected or not.
- The top flat entry in the allocation profile is inside encoding/json. Is that where you make the change?No. That frame is where the bytes were requested, not where the decision was made. Switch to the cumulative ordering and walk down to the first function you own — typically the one that chose `map[string]any` or unmarshalled a field nobody reads. Fixing what you asked the decoder for is what moves the number.
- How do you prove the change worked rather than just looks better?Two measurements. A benchmark over a fixed corpus reporting allocations and bytes per operation gives an exact, repeatable before-and-after. Then re-profile the real stage under comparable load and compare absolute `alloc_space` and `alloc_objects` totals over the same window, and confirm collector CPU actually fell.
- When is a sync.Pool the right answer for this stage, and when is it not?It is right when the profile shows a small number of large, reusable buffers dominating `alloc_space` and their lifetime is confined to one pass through the stage. It is wrong for the millions-of-small-objects case: pooling tiny values adds bookkeeping and lifetime hazards for little gain, and entries can be dropped by the collector anyway.
saying these in an interview costs you the question
- Chases a leak because the collector is busy
- Reads only inuse_space and concludes nothing is wrong
- Blames encoding/json instead of what it was asked to build
- Adds pooling before measuring what allocates
- Compares percentages instead of absolute totals over a fixed input
- Assumes collection frequency depends only on live heap size