skip to content

Reading an Execution Trace

runtime/trace writes a binary trace that go tool trace opens as timeline, goroutine and syscall views, and a flight recorder keeps the last seconds so you can snapshot a rare spike.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

questions

4

What does the timeline view in go tool trace show, and which other views does it serve?

level: middleimportance: should knowfreq 42%

answer

  1. one row per P, time along the bottom
  2. colour where something ran, gaps where nothing did
  3. click a span, follow the arrow to the waker
  4. the index also lists derived profiles
  5. MMU asks how starved the application was

basics

~20 s

The timeline in go tool trace plots wall-clock time horizontally, one row per P (Go's scheduler slot), showing which goroutine ran where, plus GC and heap rows. The same page serves goroutine analysis, blocking profiles, scheduler latency, and minimum mutator utilization.

solid answer

~50 s

`go tool trace file.trace` serves a landing page of views. The centrepiece is the timeline — "View trace by proc" and "View trace by thread" in current Go — where the x-axis is wall-clock time and each row is a P or an OS thread; coloured spans show which goroutine occupied it, with separate rows for GC work and a heap-size graph. Clicking a span gives that goroutine's identity and the stack that blocked or unblocked it, and arrows connect an unblock to the goroutine that caused it. Around the timeline the tool derives aggregate views from the same events: goroutine analysis (per-goroutine breakdowns of where time went), network, synchronization and syscall blocking profiles, a scheduler latency profile, the user-defined task and region views, and a minimum mutator utilization curve. The blocking views are served in pprof format, so you can open them with `go tool pprof` instead of the browser.

go deeper

for a junior

Know that go tool trace opens a local web UI whose main view is a time-ordered timeline of what ran, and that the same page links several derived profiles computed from the same events.

for a middle

Explain what a row and a span mean in the by-proc timeline, what the GC and heap rows add, and what the arrows between spans represent.

for a senior

Demonstrate a reading strategy: shape first at full zoom-out, then drill to named goroutines and their block reasons, then confirm against an aggregate view before you write the conclusion down.

for a principal

Decide how much of this skill a team needs: who is expected to read a trace during an incident, what is captured automatically versus on demand, and when the answer is simply to add request timing instead.

## The landing page Running `go tool trace gateway.trace` parses the file and serves a small index over localhost. Everything on that index is derived from the same event stream; the difference between the entries is whether you want the *timeline* or an *aggregate*. ## The timeline This is the view people mean when they say "looking at the trace". Current Go serves it as **View trace by proc** and **View trace by thread**. - The horizontal axis is wall-clock time across the captured window; you zoom and pan with keyboard shortcuts (`w`/`s` to zoom, `a`/`d` to pan), and the tool shows the duration of any selection. - Each row is a **P** (a scheduler slot able to run Go code) in the by-proc view, or an OS thread in the by-thread view. A coloured span on a row means a goroutine occupied that slot for that interval. - Extra rows carry runtime-wide state: garbage-collector activity across the window, the number of goroutines in existence, and the heap size as it grows and drops. - Clicking a span shows which goroutine it was, its start function, and the stack at the event; selecting a range summarises what ran in it. - Arrows link causally related events — the goroutine that closed a channel or released a lock is connected to the goroutine its action woke up. This is the feature you cannot get anywhere else: it turns "this goroutine was asleep for 200 ms" into "and here is who eventually woke it". The honest reading skill is knowing what shape to look for. Continuous full-width colour on every row means the machine is busy doing work. Wide horizontal gaps mean the process had nothing running — the interesting question is then what everyone was waiting for. Colour concentrated on one row while the others are empty means the work is not parallel at all. ## The aggregate views From the same events the tool computes: - **Goroutine analysis** — goroutines grouped by their start function, with a per-goroutine breakdown of execution time and the various waiting categories. Good for "which class of goroutine spent its life doing what". - **Network blocking profile** — where goroutines blocked waiting on network readiness. - **Synchronization blocking profile** — where they blocked on channels and locks. - **Syscall blocking profile** — where they blocked inside syscalls. - **Scheduler latency profile** — time spent runnable but not yet running. - **User-defined tasks / regions** — the spans a program annotated deliberately. - **Minimum mutator utilization (MMU)** — a curve showing, for windows of increasing length, the worst fraction of CPU that was available to application goroutines rather than the collector. A curve that stays near 1 means the collector never starved the program for long; a curve that dips toward 0 at short windows means there were brief intervals where application work got almost nothing. The four blocking-style views are served as pprof-format profiles, so the same data can be opened in `go tool pprof` and explored as a graph or flame graph — useful when the browser timeline is too large to navigate comfortably. ## Choosing a view A workable habit: - Start on the timeline, zoomed out, and look at the *shape* of the window. Ninety percent of the value of a trace is visible in that first glance: busy, idle, or serialised. - Zoom into the anomaly and click spans until you can name the goroutines involved and what they were blocked on. - Then use an aggregate view to check whether what you found is representative or a one-off, since the timeline shows individual events and the profiles show distributions. ## Practical friction Large traces are slow to load and heavy in the browser; a smaller capture is more usable than a bigger one. The viewer runs entirely locally against a file, so you can copy a trace off a production host and read it on your laptop — which is the normal workflow, since you do not want to be zooming a UI served from a machine under load. And the whole tool is read-only: it never changes the traced program, so there is nothing to undo after a session.

  • What do the arrows between spans in the go tool trace timeline mean?
    They show causality: a goroutine that unblocked another one — by sending on a channel, closing it, releasing a lock, or completing an operation the other was waiting on — is connected to the goroutine it woke. Following an arrow backwards from a long wait is how you find the party that was actually holding things up, rather than guessing from the blocked side alone.
  • What does a minimum mutator utilization curve that dips near zero at short windows tell you?
    It says there were brief intervals in which application goroutines got almost no CPU, because collector work was occupying it. Whether that matters depends on your latency budget: a service with a 10 ms tail target cares a lot about a dip at the 1 ms window, while a batch job that only cares about throughput can ignore it entirely.
  • Why is a smaller trace often more useful than a longer one when reading it in the browser?
    The viewer has to parse and render every event, so a very large trace is slow to load and awkward to pan and zoom — sometimes slower to open than it was to record. A tightly scoped few seconds that definitely contains the behaviour you are chasing is easier to navigate and no less conclusive than a minute of mostly uninteresting time.

The timeline is a train-station departure board replayed second by second: each platform is a P, each coloured block is a train that occupied it, and the arrows show which arrival released which departure.

saying these in an interview costs you the question

  • Thinks each timeline row is one goroutine rather than a P
  • Reads the timeline as a flame graph of hot functions
  • Ignores the gaps and only looks at the coloured spans
  • Believes go tool trace needs a server-side agent
  • Cannot say what minimum mutator utilization measures
open as a page

What does Go's execution tracer actually record, and why keep captures to a few seconds?

level: middleimportance: should knowfreq 38%

basics

~20 s

Go's execution tracer records every traced runtime event with a timestamp — goroutine lifecycle, block reasons, syscalls, scheduler and GC phases — rather than sampling. That exhaustiveness makes files grow with activity, so captures stay short.

open as a page

A Go service stalls for four seconds twice a day; how do you get an execution trace of it?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Run runtime/trace.FlightRecorder continuously: it keeps a recent window of trace data in memory, and an in-process watchdog that notices the stall calls WriteTo to dump it. On-demand capture starts too late for an event this short.

open as a page

How do you capture a Go execution trace from a running program or a test, and how do you open it?

level: juniorimportance: nice to knowfreq 26%

basics

~10 s

Call runtime/trace.Start with a file and defer trace.Stop, or run go test -trace=out, or fetch /debug/pprof/trace?seconds=5 from a service. Open the resulting file with go tool trace, which serves a browser UI.

open as a page