skip to content

How do you capture a Go execution trace from a running program or a test, and how do you open it?

level: juniorimportance: nice to knowfreq 26%

answer

  1. three ways in, one way out
  2. an io.Writer now, a browser UI later
  3. the test flag sits beside -cpuprofile
  4. the HTTP path needs a duration, not a snapshot
  5. flush before you close the file

basics

~10 s

Call runtime/trace.Start with a file and defer trace.Stop, or run go test -trace=out, or fetch /debug/pprof/trace?seconds=5 from a service. Open the resulting file with go tool trace, which serves a browser UI.

solid answer

~40 s

There are three ways to produce a Go execution trace. In code, `runtime/trace.Start(w)` writes the trace to any `io.Writer` until `trace.Stop()` runs, which is what you use around a specific piece of work. In tests and benchmarks, `go test -trace=bench.trace` records the whole run. In a service that imports `net/http/pprof`, `GET /debug/pprof/trace?seconds=5` returns a five-second window as a downloadable file. All three produce the same binary format, and you read it with `go tool trace file.trace`, which parses it and serves an interactive web UI on localhost. The two things beginners get wrong: `trace.Stop()` must actually run (flush before the file is closed, and before the process exits), and this is a trace, not a profile — it is a timeline of runtime events, not a ranking of hot functions.

code

go · 12 lines
go
f, err := os.Create("gateway.trace")
if err != nil {
	log.Fatal(err)
}
defer f.Close()

if err := trace.Start(f); err != nil {
	log.Fatal(err)
}
defer trace.Stop()

serveBroadcastRound() // the work you want to see

go deeper

for a junior

Be ready to name the three capture paths — runtime/trace.Start and Stop in code, go test -trace, and /debug/pprof/trace?seconds=N — and to say that go tool trace opens the file in a browser UI.

for a middle

Explain why Stop must run before the file is closed, why the HTTP endpoint takes a duration when a heap snapshot does not, and why you scope the capture to the window you care about.

for a senior

Show judgment about when a trace is the right tool at all, how long a window to take on a live service, and how you keep the pprof endpoint off a public listener while still being able to capture on demand.

for a principal

Own the standing posture: which environments carry a trace endpoint, who is allowed to trigger a capture, where trace files land, and how you keep that capability from becoming an unmonitored debug backdoor.

A Go **execution trace** is a recording of what the Go runtime did over a window of wall-clock time: when each goroutine was created, started running, blocked, was unblocked and finished; when a P (a scheduler slot for running Go code) picked work up or went idle; when goroutines entered and left syscalls; and when the garbage collector moved between its phases. It is a different artefact from a profile. A CPU profile answers "where did the time go, aggregated over the run"; an execution trace answers "what was happening at 10:42:03.418, and what was each goroutine waiting for". ## The three ways to capture one **1. From inside the program — `runtime/trace`.** `trace.Start(w io.Writer) error` turns the tracer on and streams events to `w`; `trace.Stop()` turns it off and flushes. Wrap exactly the work you want to see: ```go f, err := os.Create("gateway.trace") if err != nil { log.Fatal(err) } if err := trace.Start(f); err != nil { log.Fatal(err) } // ... run the workload ... trace.Stop() f.Close() ``` Order matters: `trace.Stop()` before `f.Close()`, because Stop is what flushes the buffered events. If you use `defer`, remember defers run LIFO, so `defer f.Close()` written first and `defer trace.Stop()` written second gives the right order. A process that exits with the tracer still running leaves a file that ends mid-stream, and `go tool trace` will usually refuse it or show only the part it could parse. **2. From a test or benchmark — `go test -trace`.** `go test -trace=t.trace ./...` records the entire test binary run. This is the cheapest way to get a trace of a reproducible workload, and it pairs naturally with `-bench`, so you can look at a benchmark's concurrency behaviour rather than only its ns/op number. **3. From a live service — the pprof HTTP endpoint.** A program that imports `net/http/pprof` exposes `/debug/pprof/trace`, which starts the tracer, waits, stops it, and returns the bytes. The `seconds` query parameter chooses the window: `curl -o stall.trace 'http://localhost:6060/debug/pprof/trace?seconds=5'`. Without it the handler defaults to a one-second trace, which is almost always too short to be interesting. Unlike `/debug/pprof/heap`, which is an instantaneous snapshot, a trace is inherently a *window*, which is why it takes a duration. Only one execution trace can be running at a time in a process; asking for a second one while the first is active returns an error rather than interleaving them. ## Opening it `go tool trace stall.trace` parses the file, converts it for the viewer, prints a `http://127.0.0.1:PORT` URL and serves an interactive UI: a timeline of the recorded window plus a set of derived views. The file itself is a compact binary format — do not expect to `less` it, and do not commit it to a repository; a few seconds from a busy process can be tens of megabytes. ## What it is not It is not a sampled profile, so it will not hand you a flame graph of hot functions. It is not free — the tracer instruments the runtime, so it costs a few percent of CPU while running and produces data fast, which is why every capture path above is scoped to a bounded window rather than left on forever. And it is per-process: a trace tells you what one Go binary did, not what a distributed request did across services. ## Practical checklist - Decide the window first. Five to ten seconds of a *representative* period beats a minute of idling. - Make sure the interesting work actually happens between Start and Stop; a trace of a program's startup tells you about initialisation, not about steady state. - Keep the endpoint off public listeners if you expose it over HTTP — anyone who can reach it can force a trace and read your program's internal timing. - Save the file with the incident time in its name. Traces are only useful when you know what window they cover.

  • What happens if the process exits without runtime/trace.Stop having run?
    Events still buffered in the runtime are never flushed and the file ends mid-stream. `go tool trace` will typically fail to parse it, or show only a truncated window. The fix is to make Stop unconditional — `defer trace.Stop()` — and to close the file after Stop, not before, since Stop is what writes the tail.
  • Why does /debug/pprof/trace take a seconds parameter when /debug/pprof/heap does not?
    A heap profile is a snapshot of state at one instant, so there is nothing to wait for. An execution trace is a recording of events over time, so the handler must turn the tracer on, wait, turn it off, and return what accumulated. Without the parameter the handler defaults to one second, which is usually too short to contain anything interesting.
  • Can you run the execution tracer permanently in production?
    Not as a normal posture. The tracer instruments the runtime, so it costs a few percent of CPU, and the data volume is large enough that a busy service produces tens of megabytes per second. The usual approaches are short, deliberate windows, or a flight recorder that keeps only a recent slice of trace data in memory and dumps it when something interesting happens.

A profile is the summary bill at the end of the month; an execution trace is the itemised timeline of every transaction in a five-second window.

saying these in an interview costs you the question

  • Thinks go tool trace ranks hot functions like a CPU profile
  • Calls runtime/trace.Start and never calls trace.Stop
  • Confuses /debug/pprof/trace with /debug/pprof/profile
  • Closes the trace file before calling trace.Stop
  • Expects the trace file to be readable text
  • Believes tracing needs a special build tag or rebuilt binary