skip to content

perf and CPU profiling

perf samples where CPU time actually goes, from a one-line perf stat to a flame graph of a hot service. Expect it whenever the interview scenario is "the box is pinned at 100% CPU — now what?".

on this pageshow

questions

5

Walk through turning a `perf record` capture of a hot Linux service into a flame graph, and explain how to read the result — what the width of a box means, and what the horizontal axis does not mean.

level: seniorimportance: must knowfreq 58%

answer

  1. capture, script, fold, render
  2. semicolon-joined stack plus a count
  3. wide equals expensive
  4. the top edge is where the CPU was
  5. left-to-right is not a timeline

basics

~20 s

Record with call graphs, dump the samples as text with perf script, fold each stack to one line with stackcollapse-perf.pl, then render with flamegraph.pl. Box width is the share of samples containing that frame; the x axis is only alphabetical ordering, never time.

solid answer

~50 s

The pipeline is four steps: `perf record -F 99 -a -g -- sleep 30` to capture, `perf script` to emit one text record per sample, `stackcollapse-perf.pl` to fold each stack into a single semicolon-joined line with a count, and `flamegraph.pl` to render an SVG. Reading it: the y axis is stack depth, with callers below and callees above. A box's **width** is the fraction of samples in which that frame appeared on the stack — so wide means expensive, and the flat plateaus along the top are the frames that were actually on-CPU. The x axis carries no meaning at all: frames are merged and sorted alphabetically so identical stacks collapse together, so left-to-right is not a timeline and adjacency implies nothing. And because sampling only fires while a thread runs, blocked time is invisible — a narrow flame graph on a slow service means it is not CPU-bound.

code

bash · 7 lines
bash
# 30 seconds of system-wide sampling at 99 Hz with call graphs
perf record -F 99 -a -g -- sleep 30

# decode, fold, render
perf script > out.perf
./stackcollapse-perf.pl out.perf > out.folded
./flamegraph.pl out.folded > flame.svg

go deeper

for a junior

Know that a flame graph comes from a perf capture with call graphs, that box width means how often a function appeared in the samples, and that wider means more expensive.

for a middle

Be able to name the pipeline — perf record with -g, perf script, stack folding, then rendering — and explain that the folded format is one semicolon-joined stack per line with a sample count.

for a senior

Demonstrate interpretation under pressure: the top edge is where the CPU actually was, the x axis is alphabetical merging rather than time, a thin graph on a slow service is a real finding, and a truncated profile must be fixed before any conclusion is drawn.

for a principal

Own how profiling fits your incident process: what capture is safe to take on a live production host, whether differential flame graphs are part of the release-regression workflow, and what has to be true about symbols and build flags for any of this to work at 3am.

## Why a flame graph and not a report table `perf report` gives you a sortable, navigable tree, and for a single dominant hotspot that is enough. A flame graph earns its place when the cost is *spread*: a dozen call paths each taking a few percent, or an expensive leaf reached from several unrelated callers. The visualisation makes the shape of the profile legible in seconds, which matters when you are looking at it during an incident. ## The pipeline ```bash perf record -F 99 -a -g -- sleep 30 perf script > out.perf ./stackcollapse-perf.pl out.perf > out.folded ./flamegraph.pl out.folded > flame.svg ``` Step by step: 1. **`perf record -F 99 -a -g -- sleep 30`** — sample at 99 Hz across all CPUs, with call graphs, for thirty seconds. 99 rather than 100 is deliberate: an off-round frequency avoids locking in step with periodic activity that ticks at a round rate. `-a` is system-wide; `-p <pid>` narrows to one process. Thirty seconds at 99 Hz gives a few thousand samples per CPU, comfortably enough for stable proportions. 2. **`perf script`** — decodes the binary `perf.data` into a text record per sample: the process, timestamp, event, and the resolved call chain, one frame per line. 3. **`stackcollapse-perf.pl`** — folds each multi-line stack into one line: `func_a;func_b;func_c 42`, a semicolon-joined path plus how many samples ended there. This is the format that decouples the renderer from the profiler, which is why the same renderer serves many different profilers. 4. **`flamegraph.pl`** — emits an interactive SVG with click-to-zoom and search. Recent perf versions also ship `perf script report flamegraph`, which produces an HTML flame graph directly and skips the external scripts — convenient when you cannot install anything on the box. ## Reading the picture - **Y axis is stack depth.** The bottom is the root of the stack (often the thread entry point), each box above it is a callee of the box below. - **Width is sample share.** A box's width is the proportion of the collected samples in which that frame was somewhere on the stack. Wide equals expensive. That is the whole message of the graphic. - **The X axis means nothing.** This is the point interviewers probe, because everyone's instinct on seeing a wide horizontal chart is "time runs left to right". It does not. Sibling frames are sorted alphabetically so that identical stacks merge into one wide box instead of scattering into hundreds of slivers. Two adjacent boxes are neighbours in the alphabet, not in time. - **The top edge is where the CPU actually was.** A frame at the top of a tower is a leaf — the sample was taken while executing *that* function's own instructions. Long flat plateaus along the top are the real hot code. A wide box with a lot stacked on top of it is expensive because of what it calls, not because of itself. - **Colour is usually meaningless** in the default palette — warm hues chosen for the flame look, randomised per frame so adjacent boxes are distinguishable. Some variants encode the DSO or user/kernel; never read significance into colour without knowing which palette produced it. ## What the graph cannot tell you **Off-CPU time is absent.** CPU sampling fires only while a thread is running. Time waiting on disk, on a lock, on a database, on a socket produces zero samples. So a service that is slow but whose flame graph is thin and boring has just given you a real answer: the latency is not CPU. That negative result should redirect the investigation, not prompt a longer capture. **Truncated stacks lie quietly.** If the capture was taken without working unwinding, the graph will be a forest of two-frame stumps and the attribution will be wrong without ever looking broken. Check stack depth before you interpret anything. **Sampling is statistical.** A frame at 1% of a few thousand samples is noise. Do not chase slivers. **Recursion merges.** A recursive function appears as a tall tower of identical boxes, which is expected rather than a bug. ## Variations worth knowing - **Icicle graph** — the same data drawn top-down (root at top), which suits reading leaf-first. - **Differential flame graph** — two captures rendered as one, colouring frames by whether they got wider or narrower. This is the strongest tool for "it got slower after the deploy", because it isolates the delta rather than making you eyeball two pictures. - **Per-process filtering** — on a system-wide capture, `perf script` output includes the command name, so you can fold a single service out of a whole-machine profile. ## The interview answer Name the four commands, then spend most of your breath on interpretation: width is samples, top edge is on-CPU, x axis is alphabetical, and blocked time never appears. That last pair is what separates someone who has actually read flame graphs during an incident from someone who has seen one in a slide deck.

  • Your service is slow but its flame graph is thin and unremarkable. What have you learned?
    That the latency is not CPU-bound. Sampling fires only while a thread is on-CPU, so waiting on disk, a lock, a database round trip or a socket contributes no samples at all. An unremarkable CPU profile is a genuine negative result: stop widening the capture and go after I/O latency, contention or a slow dependency instead.
  • Why does the flame graph pipeline go through a folded text format rather than rendering perf.data directly?
    Because folding decouples the renderer from the profiler. A folded line is just a semicolon-joined stack and a count, so anything that can emit stacks can feed the same renderer, and the intermediate file is small, diffable and easy to filter with ordinary text tools. It also makes differential rendering trivial, since you are comparing two sets of counted strings.
  • You have captures from before and after a deploy. How do you compare them?
    Render a differential flame graph rather than eyeballing two pictures. It draws one graph and colours frames by whether they grew or shrank between the two folded files, which isolates the regression directly. Make sure both captures used the same duration, sample rate and unwinding method, otherwise the widths are not comparable and the colours are meaningless.

A flame graph is a cross-section of the call stack, not a timeline: think of shining a strobe on the stack thousands of times and stacking the photographs, so the shapes that recur most often become the widest.

saying these in an interview costs you the question

  • Reads the horizontal axis as a timeline
  • Thinks colour encodes how hot a frame is
  • Assumes a thin flame graph means nothing is wrong
  • Interprets a truncated two-frame profile at face value
  • Optimises a wide box that has callees stacked above it

context

open as a page

A Linux box is pinned at 100% CPU and you have already identified the guilty process. What does `perf top` show you that a process-level view cannot, and how does it collect that information?

level: juniorimportance: should knowfreq 52%

basics

~20 s

perf top is a live sampling profiler. It interrupts the CPU many times a second, records which instruction was executing, resolves it to a function, and ranks functions by how many samples landed in them — so it names the hot code, not just the hot process.

open as a page

As an ordinary non-root user on a Linux host, `perf record` fails with a permission error about performance monitoring operations. Which sysctl governs that, what do its values mean, and what additionally blocks perf inside a container?

level: middleimportance: should knowfreq 38%

basics

~10 s

The sysctl is kernel.perf_event_paranoid, exposed at /proc/sys/kernel/perf_event_paranoid. Higher values restrict unprivileged use: 2 permits user-space measurement only, 1 also allows kernel profiling, -1 removes restrictions. Granting CAP_PERFMON is the privileged alternative.

open as a page

A `perf record -g` capture of a production Linux service produces a report where nearly every call stack is only one or two frames deep. Why does frame-pointer unwinding fail, and what do `--call-graph dwarf` and `--call-graph lbr` do differently?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Plain -g means frame-pointer unwinding, and optimised x86-64 builds usually compile with -fomit-frame-pointer, freeing that register for general use — so there is no chain to walk. --call-graph dwarf copies part of the user stack per sample and unwinds it offline using DWARF debug data; --call-graph lbr uses the CPU's branch records.

open as a page

You run `perf stat` against a CPU-bound Linux process and see roughly 0.3 instructions per cycle together with a high cache-miss count. What is that telling you about the workload, and what would an IPC near 3 mean instead?

level: middleimportance: nice to knowfreq 35%

basics

~20 s

Low instructions-per-cycle with heavy cache misses means the CPU is stalling on memory rather than doing work — the fix is data layout and access patterns. A high IPC means the core is genuinely retiring instructions, so the code is doing too much work, and you optimise the algorithm.

open as a page