skip to content

Walk through turning a `perf record` capture of a hot Linux service into a flame graph, and explain how to read the result — what the width of a box means, and what the horizontal axis does not mean.

level: seniorimportance: must knowfreq 58%

answer

  1. capture, script, fold, render
  2. semicolon-joined stack plus a count
  3. wide equals expensive
  4. the top edge is where the CPU was
  5. left-to-right is not a timeline

basics

~20 s

Record with call graphs, dump the samples as text with perf script, fold each stack to one line with stackcollapse-perf.pl, then render with flamegraph.pl. Box width is the share of samples containing that frame; the x axis is only alphabetical ordering, never time.

solid answer

~50 s

The pipeline is four steps: `perf record -F 99 -a -g -- sleep 30` to capture, `perf script` to emit one text record per sample, `stackcollapse-perf.pl` to fold each stack into a single semicolon-joined line with a count, and `flamegraph.pl` to render an SVG. Reading it: the y axis is stack depth, with callers below and callees above. A box's **width** is the fraction of samples in which that frame appeared on the stack — so wide means expensive, and the flat plateaus along the top are the frames that were actually on-CPU. The x axis carries no meaning at all: frames are merged and sorted alphabetically so identical stacks collapse together, so left-to-right is not a timeline and adjacency implies nothing. And because sampling only fires while a thread runs, blocked time is invisible — a narrow flame graph on a slow service means it is not CPU-bound.

code

bash · 7 lines
bash
# 30 seconds of system-wide sampling at 99 Hz with call graphs
perf record -F 99 -a -g -- sleep 30

# decode, fold, render
perf script > out.perf
./stackcollapse-perf.pl out.perf > out.folded
./flamegraph.pl out.folded > flame.svg

go deeper

for a junior

Know that a flame graph comes from a perf capture with call graphs, that box width means how often a function appeared in the samples, and that wider means more expensive.

for a middle

Be able to name the pipeline — perf record with -g, perf script, stack folding, then rendering — and explain that the folded format is one semicolon-joined stack per line with a sample count.

for a senior

Demonstrate interpretation under pressure: the top edge is where the CPU actually was, the x axis is alphabetical merging rather than time, a thin graph on a slow service is a real finding, and a truncated profile must be fixed before any conclusion is drawn.

for a principal

Own how profiling fits your incident process: what capture is safe to take on a live production host, whether differential flame graphs are part of the release-regression workflow, and what has to be true about symbols and build flags for any of this to work at 3am.

## Why a flame graph and not a report table `perf report` gives you a sortable, navigable tree, and for a single dominant hotspot that is enough. A flame graph earns its place when the cost is *spread*: a dozen call paths each taking a few percent, or an expensive leaf reached from several unrelated callers. The visualisation makes the shape of the profile legible in seconds, which matters when you are looking at it during an incident. ## The pipeline ```bash perf record -F 99 -a -g -- sleep 30 perf script > out.perf ./stackcollapse-perf.pl out.perf > out.folded ./flamegraph.pl out.folded > flame.svg ``` Step by step: 1. **`perf record -F 99 -a -g -- sleep 30`** — sample at 99 Hz across all CPUs, with call graphs, for thirty seconds. 99 rather than 100 is deliberate: an off-round frequency avoids locking in step with periodic activity that ticks at a round rate. `-a` is system-wide; `-p <pid>` narrows to one process. Thirty seconds at 99 Hz gives a few thousand samples per CPU, comfortably enough for stable proportions. 2. **`perf script`** — decodes the binary `perf.data` into a text record per sample: the process, timestamp, event, and the resolved call chain, one frame per line. 3. **`stackcollapse-perf.pl`** — folds each multi-line stack into one line: `func_a;func_b;func_c 42`, a semicolon-joined path plus how many samples ended there. This is the format that decouples the renderer from the profiler, which is why the same renderer serves many different profilers. 4. **`flamegraph.pl`** — emits an interactive SVG with click-to-zoom and search. Recent perf versions also ship `perf script report flamegraph`, which produces an HTML flame graph directly and skips the external scripts — convenient when you cannot install anything on the box. ## Reading the picture - **Y axis is stack depth.** The bottom is the root of the stack (often the thread entry point), each box above it is a callee of the box below. - **Width is sample share.** A box's width is the proportion of the collected samples in which that frame was somewhere on the stack. Wide equals expensive. That is the whole message of the graphic. - **The X axis means nothing.** This is the point interviewers probe, because everyone's instinct on seeing a wide horizontal chart is "time runs left to right". It does not. Sibling frames are sorted alphabetically so that identical stacks merge into one wide box instead of scattering into hundreds of slivers. Two adjacent boxes are neighbours in the alphabet, not in time. - **The top edge is where the CPU actually was.** A frame at the top of a tower is a leaf — the sample was taken while executing *that* function's own instructions. Long flat plateaus along the top are the real hot code. A wide box with a lot stacked on top of it is expensive because of what it calls, not because of itself. - **Colour is usually meaningless** in the default palette — warm hues chosen for the flame look, randomised per frame so adjacent boxes are distinguishable. Some variants encode the DSO or user/kernel; never read significance into colour without knowing which palette produced it. ## What the graph cannot tell you **Off-CPU time is absent.** CPU sampling fires only while a thread is running. Time waiting on disk, on a lock, on a database, on a socket produces zero samples. So a service that is slow but whose flame graph is thin and boring has just given you a real answer: the latency is not CPU. That negative result should redirect the investigation, not prompt a longer capture. **Truncated stacks lie quietly.** If the capture was taken without working unwinding, the graph will be a forest of two-frame stumps and the attribution will be wrong without ever looking broken. Check stack depth before you interpret anything. **Sampling is statistical.** A frame at 1% of a few thousand samples is noise. Do not chase slivers. **Recursion merges.** A recursive function appears as a tall tower of identical boxes, which is expected rather than a bug. ## Variations worth knowing - **Icicle graph** — the same data drawn top-down (root at top), which suits reading leaf-first. - **Differential flame graph** — two captures rendered as one, colouring frames by whether they got wider or narrower. This is the strongest tool for "it got slower after the deploy", because it isolates the delta rather than making you eyeball two pictures. - **Per-process filtering** — on a system-wide capture, `perf script` output includes the command name, so you can fold a single service out of a whole-machine profile. ## The interview answer Name the four commands, then spend most of your breath on interpretation: width is samples, top edge is on-CPU, x axis is alphabetical, and blocked time never appears. That last pair is what separates someone who has actually read flame graphs during an incident from someone who has seen one in a slide deck.

  • Your service is slow but its flame graph is thin and unremarkable. What have you learned?
    That the latency is not CPU-bound. Sampling fires only while a thread is on-CPU, so waiting on disk, a lock, a database round trip or a socket contributes no samples at all. An unremarkable CPU profile is a genuine negative result: stop widening the capture and go after I/O latency, contention or a slow dependency instead.
  • Why does the flame graph pipeline go through a folded text format rather than rendering perf.data directly?
    Because folding decouples the renderer from the profiler. A folded line is just a semicolon-joined stack and a count, so anything that can emit stacks can feed the same renderer, and the intermediate file is small, diffable and easy to filter with ordinary text tools. It also makes differential rendering trivial, since you are comparing two sets of counted strings.
  • You have captures from before and after a deploy. How do you compare them?
    Render a differential flame graph rather than eyeballing two pictures. It draws one graph and colours frames by whether they grew or shrank between the two folded files, which isolates the regression directly. Make sure both captures used the same duration, sample rate and unwinding method, otherwise the widths are not comparable and the colours are meaningless.

A flame graph is a cross-section of the call stack, not a timeline: think of shining a strobe on the stack thousands of times and stacking the photographs, so the shapes that recur most often become the widest.

saying these in an interview costs you the question

  • Reads the horizontal axis as a timeline
  • Thinks colour encodes how hot a frame is
  • Assumes a thin flame graph means nothing is wrong
  • Interprets a truncated two-frame profile at face value
  • Optimises a wide box that has callees stacked above it

context