How do you use `go tool pprof -diff_base` on two saved CPU profiles to find what got slower between builds?
answer
- subtract one profile from another
- rows can come out negative
- percentages relative to which profile
- same corpus, same duration, same machine
- diff a build against itself to find the noise floor
basics
~20 sRun go tool pprof -diff_base=old.pprof new.pprof. pprof subtracts the base sample by sample, so rows are deltas: positive means the new profile spends more there, negative means less. It is only meaningful if both profiles cover comparable work.
solid answer
~50 sWith two profile artifacts saved per build, `go tool pprof -diff_base=build-4821.pprof build-4899.pprof` loads the newer one and subtracts the older, so every `top`, `list` and flame-graph value becomes a delta: positive rows grew, negative rows shrank, and percentages are computed relative to the base profile. `-base` does the same subtraction with percentages relative to the resulting diff. The subtraction is arithmetic, so the answer is only as good as the comparability of the two captures — same input corpus, same duration, same machine and CPU count, ideally the same profiling window. If the durations differ, scale before you read, and treat small deltas as sampling noise rather than findings. Start at `top -cum` and package granularity, because a compiler change to inlining can shuffle flat time between functions without any real regression, then narrow to the functions whose delta survives a second capture.
code
text · 6 lines$ go tool pprof -diff_base=build-4821.pprof build-4899.pprof
(pprof) top
flat flat% cum cum%
0.28s 14.00% 0.52s 26.00% main.(*generator).renderField
0.06s 3.00% 0.06s 3.00% runtime.mallocgc
-0.14s -7.00% -0.14s -7.00% bytes.Equalgo deeper
Know that pprof can subtract one profile from another and that the resulting numbers are differences, so negative values mean the newer run spent less there.
Explain the invocation, what negative rows mean, and which profile the percentages are computed against with each of the two subtraction flags.
Show the judgment: comparable captures, an established noise floor, cumulative reading first because inlining shifts attribution, and recognising that a flat CPU diff points at off-CPU time.
Own the practice around it — what metadata every saved profile carries, which workload is the canonical one to profile, and when a measured regression is worth a build's delay versus being tracked.
## The mechanics pprof can subtract one profile from another. Given two files it loads the positional one as the profile under study and the flag's argument as the base: ``` go tool pprof -diff_base=build-4821.pprof build-4899.pprof ``` Every sample stack present in both contributes its difference; stacks only in the new profile contribute their full weight positive; stacks only in the base contribute negative. From then on every view — `top`, `list`, `peek`, the flame graph — displays deltas rather than absolutes, and rows can be negative. There are two subtraction flags and they differ in what the percentages are relative to. `-base` subtracts and computes percentages against the resulting difference. `-diff_base` subtracts and computes percentages against the *base* profile, so a row reading 14% means it grew by 14% of the old total — usually the more interpretable framing when you are asking how much worse a build got. ## Why this is a senior question rather than a flag question The subtraction is trivial; the interpretation is where people go wrong. **Comparability.** A profile is a sample of one run of one workload. Subtracting two profiles is only meaningful if the runs were alike: the same inputs (for a build-step code generator, the same schema files, not last week's tree), a comparable duration, the same machine class and the same number of CPUs. A profile collected over 30 seconds against one over 10 will show every row growing, and none of that is a regression. If durations differ, scale the base to match before reading it — pprof exposes a normalisation option for exactly this — but the safer discipline is capturing both under the same harness so no scaling is needed. **Noise.** Samples are a statistical estimate. On a short capture, a row moving by a couple of percent is inside the noise, and chasing it wastes the day. The cheap defence is a third capture: run the base build twice and diff it against itself. Whatever shows up in *that* diff is your noise floor, and only deltas comfortably above it are real. **Attribution shifting.** Between builds the compiler's inlining decisions can change, and a function that used to be inlined into its caller suddenly appears in its own right — a big positive row next to a big negative row with no change in total cost. Reading the diff at cumulative rather than flat weight, or grouping at package granularity first, makes these cancel out and keeps you looking at genuine movement. **What a CPU diff cannot see.** It only accounts for time on a CPU. If the new build got slower because it waits longer — on a lock, on a channel, on I/O — the CPU profiles can be almost identical while wall-clock latency doubles. A flat CPU diff in the face of a real slowdown is itself a finding: it tells you the cost moved off-CPU and that a CPU profile is the wrong instrument for the next step. ## A workable procedure 1. **Keep the artifacts identifiable.** A directory of profiles, one per build, is only useful if each carries the commit, the toolchain version, the input corpus and the capture duration. A profile with no provenance cannot be diffed against anything with confidence, and this is the single most common reason the file tree of saved profiles goes unused. 2. **Establish the noise floor** by diffing two captures of the same build. 3. **Diff the suspect pair** with `-diff_base` and read `top -cum` first: which subtree grew. 4. **Narrow** with `peek` to see whether the growth arrives from a new caller, and with `list` to see whether a specific line grew. 5. **Confirm** by recapturing both, or by bisecting builds if the delta straddles many commits. 6. **Verify the fix the same way**, diffing after against before — which is the same command with the arguments in the other order. ## The honest limits A diff localises movement; it does not explain it. A row that grew may have grown because it is genuinely slower, because it is called more often, because work moved into it from somewhere else, or because the input got bigger. Only the first of those is a code regression, and distinguishing them is the work `peek` and the surrounding context do. Treat the diff as a pointer, and always confirm the story against the code change between the two builds before you act on it.
- The CPU diff between a fast build and a slow one is essentially flat, yet the slow build takes twice as long. What does that tell you?That the extra time is not spent on a CPU. The added latency is waiting — on a lock, a channel, a syscall or a remote call — and a CPU profile is blind to it by construction. Take the flat diff as a positive result: it rules out a compute regression and redirects you to instrumentation that records waiting rather than execution.
- How do you tell a real regression from sampling noise in a pprof diff?Establish a noise floor first: capture the same build twice and diff those two against each other. Anything that appears there is noise. Then only trust deltas comfortably larger than it, prefer longer captures over short ones, and reproduce the finding with a second pair of captures before acting on it.
- Why can a diff show a large positive row next to a large negative one with no real change in cost?Usually because attribution moved. A different inlining decision between builds makes a callee appear as its own frame instead of being folded into its caller, so weight shifts between two rows while the subtree total is unchanged. Reading the diff cumulatively, or at package granularity, makes such pairs cancel and keeps genuine movement visible.
- What has to be stored alongside each saved profile for a diff months later to be trustworthy?The commit or build id, the toolchain version, the machine class and CPU count, the input corpus, and the capture duration. Without them you cannot tell whether a delta is code, environment or workload, and the sources needed for annotated line views cannot be checked out at the right revision either.
saying these in an interview costs you the question
- Diffs profiles of different durations without scaling
- Treats every positive row in a diff as a regression
- Never establishes a noise floor before trusting a delta
- Concludes nothing regressed because the CPU diff is flat
- Saves profile artifacts with no build or workload metadata