skip to content

A Go agent binary for edge devices grew several megabytes in one release — how do you find why?

level: seniorimportance: should knowfreq 28%

answer

  1. control the variables first
  2. two builds, same toolchain, same flags
  3. list symbols by size and diff them
  4. many small new symbols, not one big one
  5. confirm by removing the suspect and remeasuring

basics

~20 s

Rebuild both revisions with the same toolchain and flags, compare file sizes, then list each binary's symbols with go tool nm sorted by size and diff the two lists. A jump in type metadata, method tables and reflection code points at name-based dispatch defeating the linker's pruning.

solid answer

~50 s

Make the comparison honest first: build the previous and current revisions on the same machine, same toolchain version, same flags, same target platform, and confirm with `go version -m` that nothing else moved. Then stop guessing and look: `go tool nm -size -sort size` on each binary, diff the symbol lists, and ask what appeared rather than what grew a little. In a case like this the answer is usually a class of symbols, not one big function — type descriptors, name strings and method tables that used to be pruned. From there, find what made them reachable: a package added in this release that looks types up by name at run time, or that marshals arbitrary values, pulls the reflection machinery into the reachable set, and once method lookup by name is reachable the linker retains exported methods it previously dropped. Confirm by building the branch with that one dependency replaced by a static registration table and measuring again.

code

text · 9 lines
text
go build -o agent.old ./cmd/agent   # previous revision
go build -o agent.new ./cmd/agent   # current revision
ls -l agent.old agent.new

go tool nm -size -sort size agent.old | head -50 > old.syms
go tool nm -size -sort size agent.new | head -50 > new.syms
diff old.syms new.syms

go version -m agent.new

go deeper

for a junior

Know that you can look inside a compiled Go binary and list its symbols with their sizes, and that comparing two builds means keeping the toolchain, flags and target identical.

for a middle

Explain what could make a binary grow without new code being written: metadata the linker used to remove is now reachable, typically through lookup by name at run time.

for a senior

Demonstrate the whole loop — control the variables, diff the symbols, recognise the class of what appeared, confirm with a third build, then choose between removing, confining and accepting the cost.

for a principal

Own the guardrail rather than the incident: an artifact-size budget checked on every build, and a stated position on what the team is allowed to spend to keep dynamic behaviour.

## Make the measurement trustworthy before you interpret it A size regression is one of the few production problems you can reproduce exactly, so do that first. Build both revisions: - on the same machine, with the same toolchain version — a toolchain upgrade alone moves size in both directions; - for the same target operating system and architecture; - with identical build flags and tags. `go version -m ./agent` prints the toolchain and the build settings recorded in each binary, which is the fastest way to catch an accidental difference in target or flags. If the two builds differ in anything but source, you are comparing noise. ## Look at symbols, not at the total `go tool nm -size -sort size ./agent | head -50` lists the largest symbols in a binary. Do it for both and diff the lists. Two shapes of answer come out: - **One or a few large new symbols.** Someone embedded an asset, generated a very large table, or vendored a big new component. This is the easy case and the total tells you almost everything. - **A broad new population of small symbols.** Hundreds or thousands of entries that did not exist before, clustered around type metadata, name strings and method tables. This is the interesting case and it means something changed what the linker is *allowed* to remove. The second shape is the signature of defeated dead-code elimination. The linker normally keeps a method only when it is called directly or when its type is boxed into an interface whose method name and signature match. Once a reachable code path looks methods up by a name computed at run time, the linker can no longer enumerate the possibilities and falls back to retention — and the retained methods drag their callees in behind them. ## Find the cause, not just the symptom With a candidate in hand, work backwards to what made it reachable: - Read the release's diff for anything that resolves a type or method **by string**: a registry keyed by type name, a plugin table, a decoder that maps a message kind to a handler by looking the name up. - Remember it can arrive indirectly. A package added for an unrelated reason may format arbitrary values or serialise unknown shapes, and reflection comes with it. - Check the module graph too: a dependency upgrade can introduce name-based dispatch inside code you did not write. ## Prove it The honest confirmation is a third build. Take the current revision, replace the suspected mechanism with a statically wired equivalent — an explicit map from a string key to a function value, populated by each type's own package at initialisation — and measure again. If the megabytes come back off, you have your cause and, incidentally, your fix. If they do not, your hypothesis was wrong and you have lost half an hour rather than a week. ## Then decide what to do about it Diagnosis and remedy are separate. Options, roughly in order of preference for a constrained target: - **Remove the need for name-based lookup** by generating or hand-writing the registration table. Static wiring is also compiler-checked, which the string version is not. - **Confine it.** If one subsystem genuinely needs dynamic dispatch, keep it behind a boundary and out of the artifact shipped to the constrained target, rather than letting it be reachable from every build. - **Accept it, with a number.** Sometimes the feature is worth the megabytes. Say so explicitly, record the new baseline, and make the artifact size a checked build output so the next regression is caught the day it lands rather than at the next device rollout. ## Keep the regression from recurring A size budget enforced in the build is worth more than any amount of after-the-fact archaeology: record the artifact size on every build, and fail when it crosses an agreed ceiling. The check costs nothing and turns a slow, deferred discovery into an immediate one, which matters most on targets where the artifact must fit a partition or be pushed over a slow link. ## What interviewers are checking That you control the variables before drawing conclusions; that you know a concrete tool for looking inside an artifact instead of reasoning from intuition; that you can recognise the difference between one big thing and a broad class of newly retained metadata; and that you confirm a hypothesis with a measurement rather than a plausible story.

  • The symbol diff shows thousands of small new entries rather than one big one. What does that pattern tell you?
    That something changed what the linker may discard, not what you wrote. A broad population of newly retained type metadata, name strings and method tables is the signature of dead-code elimination being disabled by run-time lookup by name.
  • How do you confirm the cause instead of just believing the story?
    Build a third artifact from the same revision with the suspected mechanism replaced by static registration, and measure. If the size returns to the old baseline, the hypothesis is confirmed and you already hold the fix.
  • What would you put in the build so this is caught next time?
    A recorded artifact size on every build and a hard ceiling that fails the pipeline when it is crossed. It converts a discovery made at device-rollout time into one made on the pull request that caused it.

saying these in an interview costs you the question

  • Compares binaries built with different toolchain versions or targets
  • Guesses at the cause from the diff without measuring
  • Assumes bigger binary means more of their own code
  • Never inspects the artifact's symbols at all
  • Reaches for stripping before understanding what grew