How would you prove that struct padding, not live data, is inflating a Go service's heap?
answer
- state the hypothesis in bytes first
- live bytes and live objects, not cumulative
- sizeof minus the sum of the fields
- multiply the waste by the live count
- fix the generator template, not its output
basics
~20 sMeasure rather than guess. Compare unsafe.Sizeof of the struct against the sum of its field sizes, multiply the difference by the live instance count from a heap profile, and check that the product accounts for the missing memory.
solid answer
~50 sStart from a heap profile in `inuse_space` mode to find which type dominates, and use `inuse_objects` to get the live count. Then compute the per-instance waste: `unsafe.Sizeof(record{})` minus the sum of the field sizes, which you can also get for a whole package from the `fieldalignment` analyzer in `golang.org/x/tools`, run through `go vet -vettool`. Padding times count either explains the gap or it does not — if the type is dominated by what its pointers reach, reordering fields will win you nothing. When the arithmetic does justify it, change the declaration order at the source: for structs emitted by a code generator, that means the template, not the generated file, or the next regeneration silently reverts it. Then re-measure the same heap profile and pin the new size with a `unsafe.Sizeof` assertion in a test.
code
go · 13 linestype recordSchemaOrder struct {
ok bool
id int64
score float64
kind uint16
} // unsafe.Sizeof == 32
type recordRegrouped struct {
id int64
score float64
kind uint16
ok bool
} // unsafe.Sizeof == 24go deeper
Be ready to say that unsafe.Sizeof reports the true in-memory size including padding, and that it can be larger than the field sizes added up.
Expect to turn the hypothesis into arithmetic: waste per instance times the live object count, compared against the memory you cannot account for.
Show the full loop — profile, quantify, change it where regeneration cannot revert it, re-measure the same metric, and say out loud when the saving is too small to bother with.
Own where this effort belongs in a memory budget. Holding fewer objects almost always beats holding smaller ones, and a codebase-wide reordering campaign usually costs more in churn than it returns.
## The situation this arises in A code generator turns a wire schema into Go structs, emitting fields in the order the schema declares them: a one-byte version, a four-byte flag word, an eight-byte sequence number, a two-byte kind. Nobody reads the generated file, the fields look tiny, and the service holds millions of these records in a few large slices. Resident memory is a third higher than the arithmetic anyone did on the back of an envelope. Padding is a plausible culprit — and it is exactly the kind of hypothesis that must be measured before it is acted on, because the fix touches generated code that everything else depends on. ## Step one: find the type that matters Take a heap profile and read it in `inuse_space` mode — the live bytes, not the cumulative allocation totals — and confirm which type or which allocation site holds the memory. Switch the same profile to `inuse_objects` for the live instance count. Two numbers come out: bytes held and objects alive. If the dominant entry is not the generated record type at all, stop; the padding theory is already dead. ## Step two: quantify the waste per instance The measurement is a subtraction: ``` padding per instance = unsafe.Sizeof(record{}) - (sum of the field sizes) ``` For the header above that is 24 minus 14, ten wasted bytes in a 24-byte value — 42% of it. Multiply by the live object count. If the product is a meaningful share of the gap you are chasing, you have your explanation; if it is 3% of it, you have eliminated a hypothesis, which is also progress. The `fieldalignment` analyzer from `golang.org/x/tools`, run over the package with `go vet -vettool=...`, does the same arithmetic for every struct in the package and prints the size it could be, which is the faster route when you do not yet know which type to look at. ## Step three: decide whether it is worth it Three things commonly make the saving smaller than it looks: - **The size class.** Individually heap-allocated objects are rounded up to one of the allocator's size classes, so a struct that shrinks from 24 bytes to 20 still occupies 24. The saving is real only when the shrink crosses a class boundary — or when the values live in one big slice, where elements are packed contiguously and the saving is exact. That distinction is worth checking before promising a number. - **Pointers dominate.** If each record holds fields that point at other objects, the memory those reach usually dwarfs the padding, and the leverage is in holding fewer objects rather than smaller ones. - **It may not be a latency win.** A smaller record improves cache utilisation and gives the collector less to walk, but if the workload is not memory-bound the benchmark will not move. Measure the footprint you set out to reduce — heap and RSS — and do not oversell throughput you did not observe. ## Step four: make the change where it survives Reorder the fields in **descending alignment**: the eight-byte sequence, then the four-byte flags, then the two single-byte fields. The header drops from 24 bytes to 16, a third of the memory for that slice. For generated code the edit belongs in the generator's template or its field-emission order, never in the output file — an edit to generated code is reverted by the next run, quietly, and the regression will surface weeks later as the same memory report. This is also the case where field reordering costs nothing in readability: nobody reads generated structs, so the usual objection that schema-ordered fields document the schema does not apply. ## Step five: re-measure and pin it Run the same `inuse_space` profile again on the same workload and confirm the drop matches the prediction; a benchmark with `-benchmem` is a good complement when the type is allocated in a hot path. Then stop the number from drifting: a test asserting `unsafe.Sizeof(record{}) == 16` fails the moment a field is inserted in the middle, and the failure text is the natural place to record why the fields are grouped as they are. If the type must build for both 64-bit and 32-bit targets, assert per target rather than hard-coding one machine's answer. ## What good judgment sounds like The strong version of this answer is not the field ordering — that part is mechanical. It is the discipline around it: a hypothesis stated in bytes, a profile that confirms or kills it, a change made where regeneration cannot undo it, a re-measurement against the original metric, and honesty about the cases where 40% of a struct is padding and it still does not matter because the service only ever holds a thousand of them.
- The struct is 33% smaller but the benchmark shows the same nanoseconds per operation. Was the work wasted?Not necessarily, but the claim has to change. You set out to reduce footprint, and the heap profile shows you did; a smaller element also means fewer cache lines touched per scan and less for the collector to walk. If the workload was never memory-bound, latency was not the metric to promise. Report the heap and RSS you moved, not a speedup you did not observe.
- Where does removing padding buy you nothing at all?Where the values are individually heap-allocated and the shrink does not cross an allocator size class, so each object still occupies the same rounded-up block. Also where the type is instantiated a few thousand times, where the record's real weight is in what its pointer fields reach, and where the value only ever lives briefly on the stack.
- How do you keep the ordering from being undone by the next person?Two guards. A test asserting `unsafe.Sizeof` of the type, whose failure message explains the grouping, catches a field inserted by hand. And for generated types, making the change in the generator itself is the only version that survives — a reordered output file is reverted by the next regeneration without anyone noticing.
- Would you apply this across the whole codebase once you have the analyzer wired up?No. Run it to find candidates, then act only on types that exist in very large numbers. Applied everywhere it produces a large diff over structs allocated once, trades schema-shaped or domain-shaped declarations for byte counts nobody will benefit from, and buries the two changes that actually mattered.
saying these in an interview costs you the question
- Reorders fields by eye without measuring anything first
- Reads cumulative allocation totals instead of live in-use bytes
- Edits generated code rather than the generator that produced it
- Ignores that most of a record's weight may be what its pointers reach
- Promises a latency improvement from a pure footprint change
- Forgets that individually allocated objects round up to a size class