skip to content

A `perf record -g` capture of a production Linux service produces a report where nearly every call stack is only one or two frames deep. Why does frame-pointer unwinding fail, and what do `--call-graph dwarf` and `--call-graph lbr` do differently?

level: seniorimportance: should knowfreq 44%

answer

  1. -g is not free and not automatic
  2. optimised builds drop the frame pointer
  3. copy the stack now, unwind later
  4. hardware branch records are shallow
  5. check stack depth before trusting a report

basics

~20 s

Plain -g means frame-pointer unwinding, and optimised x86-64 builds usually compile with -fomit-frame-pointer, freeing that register for general use — so there is no chain to walk. --call-graph dwarf copies part of the user stack per sample and unwinds it offline using DWARF debug data; --call-graph lbr uses the CPU's branch records.

solid answer

~40 s

`-g` is shorthand for `--call-graph fp`, which walks the frame-pointer chain. On x86-64 that register is a general-purpose register, and optimised builds have historically compiled with `-fomit-frame-pointer`, so the chain simply is not there and the unwind stops after a frame or two. Three ways out. Rebuild the hot code with `-fno-omit-frame-pointer` — cheapest at sample time and what recent Fedora and Ubuntu releases now do for distro packages. Or use `--call-graph dwarf`, where perf copies a chunk of the user stack with every sample (8 KB by default, tunable as `--call-graph dwarf,<size>`) and unwinds offline using DWARF unwind information; that needs debug data present and produces far larger `perf.data` files. Or `--call-graph lbr` on Intel, which reads the CPU's last-branch records: nearly free, but limited hardware depth and user-space only.

code

bash · 8 lines
bash
# frame pointers present: cheap, complete, includes kernel frames
perf record -g -F 99 -p 4242 -- sleep 30

# no frame pointers: copy 16 KB of stack per sample, unwind offline
perf record --call-graph dwarf,16384 -F 49 -p 4242 -- sleep 30

# self overhead rather than cumulative
perf report --stdio --no-children

go deeper

for a junior

Know that -g asks perf for call stacks and that a report showing only one or two frames means the stacks failed to unwind, not that the program has no callers.

for a middle

Explain the three mechanisms: frame pointers walked cheaply at sample time, DWARF unwinding from a copied stack chunk after the fact, and hardware branch records — and why optimised builds break the first one.

for a senior

Show the operational tradeoff you would actually make on a production host: reduced sample rate and a short window for dwarf mode, awareness of perf.data growth, mixed-build truncation at library boundaries, and validating stack depth before trusting the report.

for a principal

Own the build policy: whether the fleet ships release binaries with -fno-omit-frame-pointer so that any incident can be profiled, what that costs in performance, and how debug symbols are made retrievable at analysis time without shipping them everywhere.

## What a call graph costs A sample without a stack tells you which function was executing. A sample *with* a stack tells you who called it, which is usually the actionable half — `memcpy` being hot means nothing until you know which code path is calling it. Getting that stack means unwinding from the sampled instruction pointer back through the caller chain, and there are exactly three mechanisms available, each with a different bargain. ## Frame pointers: `--call-graph fp` (what `-g` means) A frame pointer is a register holding the base of the current stack frame, with each frame storing the caller's value. Unwinding is then trivial: follow the linked list. It is cheap enough to do in the kernel during the sample itself. The catch is that on x86-64 the frame pointer register is a perfectly good general-purpose register, and compilers have long defaulted to `-fomit-frame-pointer` at optimisation levels above `-O0`. When the register is repurposed there is no chain, and the unwind terminates almost immediately. That is exactly the symptom in the question: one or two frames, then nothing. It gets worse in mixed builds. Your application may have frame pointers while the system libraries it calls do not, so stacks truncate the moment they pass into `libc` or a third-party shared object — a partial, misleading picture that is arguably more dangerous than none. **The fix at the source:** rebuild with `-fno-omit-frame-pointer`. The runtime cost is small — one register and a couple of instructions per call — and both Fedora (from Fedora 38) and Ubuntu (from 24.04 LTS) now build their distribution packages this way precisely so that production profiling works. On your own service, adding the flag to release builds is usually the single highest-leverage change for profilability. ## DWARF: `--call-graph dwarf` When you cannot rebuild, DWARF unwinding works on binaries that have no frame pointers at all. The mechanism is blunt: with each sample perf copies a chunk of the user stack — 8 KB by default — into the sample record along with the register state. Nothing is unwound at sample time. Later, `perf report` or `perf script` replays that stack memory against the DWARF unwind information for the binaries involved and reconstructs the frames. The consequences follow from the mechanism: - **`perf.data` gets very large.** 8 KB per sample at hundreds of samples a second per CPU adds up fast. Drop `-F` and shorten the capture. - **Deep stacks get truncated.** If the live stack is deeper than the copied window, the tail is missing. `--call-graph dwarf,16384` or larger buys depth at the cost of even more data. - **Debug information must be available** at analysis time — the `-dbgsym` package on Debian/Ubuntu, `dnf debuginfo-install` on Fedora/RHEL — and it must match the exact build. - **Overhead is real** at capture time too, because copying stack memory on every sample is not free. ```bash perf record --call-graph dwarf,16384 -F 49 -p 4242 -- sleep 30 perf report --stdio --no-children ``` ## LBR: `--call-graph lbr` Intel CPUs maintain last-branch records, a small hardware ring of recent branch source/target pairs. perf can reconstruct a call chain from them essentially for free — no stack copying, no frame pointers, no debug data. The limits are hardware limits. The ring is shallow (a couple of dozen entries depending on the microarchitecture), so deep stacks are cut off; it is user-space only, so kernel frames are absent; and it is Intel-specific, which matters on a mixed or Arm fleet. For a shallow, hot user-space call path it is the cheapest good answer available. ## Reading the report once stacks work One more thing separates people who have used this from people who have read about it: `perf report` defaults to `--children`, which shows *cumulative* overhead — time in a function plus everything it called. That is what you want for finding an expensive subtree. `--no-children` shows *self* overhead, the time in the function's own instructions, which is what you want for finding the actual hot loop. Misreading one for the other sends people optimising a wrapper that costs nothing itself. ## Choosing, in practice 1. Can you rebuild the hot code? `-fno-omit-frame-pointer` and plain `-g`. Best signal, lowest capture cost, works for kernel frames too. 2. Production binary you cannot touch, and debug data available? `--call-graph dwarf` with a reduced sample rate and a short window. 3. Intel host, shallow user-space stack, minimal overhead required? `--call-graph lbr`. And always sanity-check the first report: if the deepest stacks are two frames, stop and fix the unwinding before you draw any conclusion from the data. A truncated profile does not fail loudly — it quietly attributes everything to the wrong place.

  • Why does `--call-graph dwarf` make `perf.data` so much larger than frame-pointer mode?
    Because it defers the unwind. Each sample carries a copy of a chunk of the user stack — 8 KB by default — plus register state, so perf can reconstruct frames later against DWARF unwind data. Frame-pointer mode walks the chain in the kernel and stores only the resulting addresses. At a few hundred samples per second per CPU the difference is orders of magnitude, so you lower `-F` and shorten the capture window.
  • Your stacks unwind correctly but the frames read as `[unknown]` and hex addresses. Is that the same problem?
    No — that is symbolization, not unwinding. perf recovered the return addresses but cannot map them to names because the binaries are stripped and no matching debug information is installed. Install the `-dbgsym` or debuginfo package for the exact build, or point perf at a filesystem that has them. Fixing unwinding and fixing symbols are separate steps, and confusing them wastes a lot of time.
  • What is the difference between `perf report --children` and `--no-children`?
    `--children`, the default, shows cumulative overhead: time in a function plus everything it called, which surfaces expensive subtrees and puts `main` near the top. `--no-children` shows self overhead — samples that landed in that function's own instructions — which is what identifies the actual hot loop. Reading one as the other leads people to optimise a wrapper that costs nothing by itself.

saying these in an interview costs you the question

  • Assumes -g always produces complete call stacks
  • Thinks DWARF unwinding needs no extra data at analysis time
  • Ignores the perf.data size explosion from dwarf mode
  • Believes LBR gives unlimited stack depth
  • Reads cumulative children overhead as time in the function itself

context