skip to content

In a flame graph from a sampling profiler, what do a frame's width and the vertical stacking each mean?

level: middleimportance: must knowfreq 52%

answer

  1. Width is a share, not a duration
  2. Boxes merged across many stack samples
  3. Vertical means caller below, callee above
  4. Left to right is only a merge order
  5. Exposed top edge is self cost

basics

~20 s

Width is the share of collected stack samples containing that frame — its share of the sampled resource, not elapsed time and not a call count. Vertical stacking is caller-to-callee ancestry, merged across every sample.

solid answer

~50 s

A flame graph renders the merged tree of stack samples a sampling profiler collected over a window. **Width** is proportional to the number of samples in which that frame appeared at that position, so it reads as a share of the sampled resource — CPU time, allocated bytes, blocked time. **Vertical stacking** is ancestry: a box sits directly on top of the caller it was reached from, and can never be wider than its parent. **Left-to-right position means nothing.** Children are ordered by a stable convention, usually alphabetically, purely so the merge is deterministic and two graphs are comparable. The horizontal axis is not time and adjacency is not execution order — the standard misreading, and an easy one for anyone who spends their day in a distributed-trace waterfall, where the axis genuinely is wall-clock time and each bar is one real occurrence.

code

text · 4 lines
text
main;handleUpload;parseCsv;splitLine 4127
main;handleUpload;parseCsv;parseFloat 2318
main;handleUpload;writeBatch;compress 9042
main;scheduler;sweepExpired 613

go deeper

for a junior

Recall the two axes: width is the share of collected samples that included that frame, and a box sits on top of whatever called it. Remember that the horizontal direction is not time.

for a middle

Explain how merged stack samples become the picture, distinguish a box's total width from its exposed top edge, and say why left-to-right position carries no meaning at all.

for a senior

Demonstrate a reading workflow on a real graph: ignore the base, find where width narrows, spot plateaus, search by frame name, and know when a fleet-wide graph is hiding a single bad instance.

for a principal

Own how profiles are compared across time and builds rather than read one at a time, and what conventions — window selection, differencing, per-instance filtering — keep engineers from drawing confident conclusions from a misread picture.

## From stack samples to the picture A sampling profiler collects many **stack samples** over a window: each one is the sequence of frames that was on the stack at the instant of an interrupt. Identical sequences are merged and counted, producing a tree in which every distinct stack ends in a leaf carrying a sample count. A flame graph is a rendering of exactly that tree — root frames along the base, each callee drawn as a box resting on its caller, and every box sized in proportion to the samples in the subtree beneath it. Nothing is lost or added in the drawing. Everything you can read off the picture is a property of the merged counts. ## What each axis means - **Width** — the share of collected samples in which that frame was on the stack at that position. A frame 22% of the total width was present in roughly 22% of samples, and since samples proxy the sampled resource, that reads as 22% of the CPU, of the allocated bytes, or of the blocked time. - **Vertical** — ancestry. Caller below, callee above. A child can never be wider than its parent, because it was only ever reached through it. - **Left to right** — nothing at all. Children are laid out in a stable order, typically alphabetical, so the merge is deterministic and two graphs of the same workload line up. Adjacency does not imply sequence, concurrency or anything else. - **Depth** — stack depth only. A tall graph is a deeply nested one, not a slow one. One more distinction does real work: the **exposed top edge** of a box — the part with nothing drawn above it — is where samples landed *in that frame itself*. Total width is cumulative cost including everything it called; exposed top edge is self cost. ## Why people read it as a timeline | | Flame graph | Flame chart | Trace waterfall | |---|---|---|---| | Horizontal axis | share of samples; order is a merge convention | time | wall-clock time | | One box is | a frame merged across many samples | one recorded invocation | one span | | Scope | a window of a process or a fleet | one recorded execution | one request | | Answers | which code consumed the resource | what ran when in this recording | where this request spent its time | The two neighbours in that table are why the misreading is so persistent. A **flame chart** looks almost identical but genuinely puts time on the horizontal axis, with one box per invocation. A **trace waterfall** also runs left to right in time. An engineer who lives in those views will instinctively read a flame graph's left edge as 'earlier' and its width as 'took this long', and both readings are wrong: the same frame's samples may have come from moments scattered across the whole window and from many different threads or hosts. ## Reading one in practice 1. **Ignore the base.** The root and the framework frames beneath your code are wide by construction — everything passed through them. A wide box at the bottom tells you almost nothing. 2. **Scan upward for where a wide box narrows** into many thin ones. That split point is where the cost stops being concentrated. 3. **Look for plateaus.** A wide box with a flat, exposed top edge is a frame burning the resource itself rather than delegating, and is usually the thing worth optimising. 4. **Search by frame name**, not by position. A function often appears under several different parents, and the tool's search will sum its total share across all of them. 5. **Compare two graphs by differencing them**, not by eyeballing shapes side by side. Left-to-right position is a merge artefact, so visual comparison misleads. ## Pitfalls worth naming - **Orientation.** Some tools draw the icicle form, root at the top growing downward. Same data, flipped; the axis rules are unchanged. - **Missing or inlined frames.** If the profiler cannot walk the stack fully, samples get attributed to whatever frame it could see, and the graph confidently blames the wrong code. A suspiciously flat, shallow graph is the tell. - **Merging across a fleet.** A frame at 8% of a fleet-wide graph can be 56% of one instance. If the problem is one bad node, aggregate width hides it — filter to the instance first. - **Window length.** A graph over a long window averages a short spike away. Narrow the window to the incident before reading it. ## What interviewers listen for - Says 'share of samples' rather than 'duration' when explaining width. - Volunteers that the horizontal axis is not time before being prompted. - Distinguishes cumulative width from the exposed top edge. - Knows that a wide frame at the base is a routing fact, not a finding.

  • A frame near the base of a flame graph is nearly the full width. What does that tell you?
    Very little on its own. Frames near the base are wide because almost every sampled stack passed through them, which is a routing fact rather than a finding. Read upward instead: look for where that width narrows into many thin children, and for a box whose exposed top edge is wide, since that is the frame consuming the resource itself.
  • Two flame graphs of the same service look different left to right. Does that mean the workload changed?
    Not by itself. Horizontal position is a merge convention, usually alphabetical ordering of children, so a single new or absent frame shifts everything beside it without any change in cost. Compare by frame name and share, ideally with a differential view that subtracts one graph from the other, rather than by matching shapes visually.

It is a census of the program, not a diary: it says what proportion of the sampled moments each code path occupied, never when any of them happened.

saying these in an interview costs you the question

  • Reads the horizontal axis as elapsed time
  • Thinks left-to-right order shows execution order
  • Says a wide frame at the base is the bottleneck
  • Treats frame width as a count of calls
  • Assumes the graph describes a single request
  • Ignores the difference between cumulative width and self cost