What does Go's execution tracer actually record, and why keep captures to a few seconds?
answer
- not a sample — a record of each event
- every transition gets a timestamp
- goroutine created, blocked, unblocked, ended
- volume scales with activity, not seconds
- a couple of percent CPU, megabytes fast
basics
~20 sGo's execution tracer records every traced runtime event with a timestamp — goroutine lifecycle, block reasons, syscalls, scheduler and GC phases — rather than sampling. That exhaustiveness makes files grow with activity, so captures stay short.
solid answer
~50 sThe execution tracer is exhaustive, not statistical. Whenever the runtime does something interesting to a goroutine — creates it, schedules it onto a P, blocks it on a channel, mutex, syscall or network read, unblocks it, ends it — an event with a nanosecond timestamp (and often a stack) goes into a per-P buffer, along with GC phase transitions and heap-size changes. That is why a trace can answer "what was this goroutine waiting on at that instant", which no sampled profile can. The price is volume: a busy process can emit tens of megabytes per second, and the tracer costs a small but real slice of CPU — roughly a couple of percent in recent Go, much worse before the tracer was overhauled. So you take a five-to-ten-second window around something you care about, not an hour of steady state.
go deeper
Remember the headline: the execution tracer records runtime events with timestamps instead of sampling, which is why you take a short window rather than leaving it running.
Be able to list what is recorded — goroutine lifecycle with block reasons, scheduler and syscall transitions, GC phases — and explain why exhaustive recording makes file size track activity rather than duration.
Show that you can budget the capture: which instance, how many seconds, what latency cost you are accepting, and why a trace answers latency questions while a profile answers throughput questions.
Frame the tradeoff for a team: what diagnostic coverage is worth a couple of percent of CPU and a large artefact, and where you would rather invest in request-level timing that spans process boundaries.
## Exhaustive, not sampled The distinction that matters most is the sampling model. A CPU profile is **statistical**: a timer fires periodically, the runtime records the stack of whatever is running, and the tool aggregates thousands of those samples into an estimate of where time went. The execution tracer is the opposite — it is **event-driven and exhaustive**. Every time the runtime performs a traced transition, it appends a record with a nanosecond-resolution timestamp to a buffer belonging to the current P. Roughly what gets recorded: - **Goroutine lifecycle** — created (with the creation stack), started running on a P, blocked (and *why*: channel send/receive, select, mutex, network poller, sleep, GC assist), unblocked (and by whom), ended. - **Scheduler activity** — Ps starting and stopping, goroutines being handed between them, the runtime going idle. - **Syscalls** — entry and exit, including whether the goroutine's P was handed off while it was blocked in the kernel. - **Garbage collection** — phase transitions, sweep and mark activity, heap-size changes over the window. - **User annotations** — the log, task and region events a program emits deliberately. Because every event carries a timestamp, the trace reconstructs a *timeline*: at any instant in the window you can say which goroutines were running, which were runnable but waiting for a P, and which were blocked on what. That is the unique value. A profile tells you the top of a distribution; a trace tells you a story with an ordering. ## What it does *not* record It is not a call tracer. It does not log every function call, every allocation site, or every line executed. Stacks are attached to specific events (goroutine creation, blocking, unblocking, syscalls), not sampled continuously — so a trace is a poor way to find a hot loop, and an excellent way to find out why nothing was running at all. ## Why the window is short Two reasons, and both are worth stating explicitly. **Volume.** Event count scales with runtime activity, not with wall-clock time. A quiet process produces almost nothing; a gateway with tens of thousands of connections, each waking a goroutine on every frame, produces an enormous number of block/unblock pairs. Tens of megabytes per second is a realistic order of magnitude, and parsing scales with size — a very large trace can take longer to load in the viewer than it took to record. **Overhead.** The tracer adds work on hot runtime paths and writes buffers out continuously. In recent Go this is small — commonly cited as a couple of percent of CPU — but it is not zero, and it is highest exactly where the trace is largest: a process doing millions of scheduling transitions per second. Older Go releases were much more expensive, which is why "never trace in production" was received wisdom for years and is now too strong a rule. The practical consequence: pick a window of a handful of seconds that you believe contains the behaviour you want to explain, and take it deliberately. A 60-second trace of a mostly idle service is a large file that says nothing; five seconds spanning a latency spike is gold. ## Reading consequences of the event model Because it is exhaustive, a trace is *falsifiable* in a way a profile is not. If a goroutine does not appear as running in a window, it did not run — that is a fact, not an estimate. If you see a blocking event and then a long gap before the unblock, the wait is real, not a sampling artefact. This is why traces are the right tool for latency questions ("where did those four seconds go?") and profiles are the right tool for throughput questions ("what is burning the CPU?"). Because the tracer is per-process, a trace also cannot tell you about time spent in another service, in the kernel beyond syscall boundaries, or waiting for a peer on the network beyond "this goroutine was blocked in the poller". Pair it with request-level timing when the question crosses a process boundary. ## A note on cost budgeting When someone asks whether they can capture a trace on a production instance, the honest answer is: yes, for a few seconds, on one instance, at a moment you have chosen — and you should expect a small latency bump while it runs, plus the I/O of writing the file. That framing is much more useful than a blanket yes or no.
- Why can an execution trace explain a latency spike when a CPU profile cannot?A latency spike is usually time spent *not* running — blocked on a lock, waiting for a P, parked in the poller, or stalled behind the collector. A CPU profile only samples goroutines that are on-CPU, so idle waiting is invisible to it by construction. The trace records the block and unblock events with timestamps, so the wait itself is measurable.
- Does the size of a Go execution trace depend mostly on the duration you record?No — it depends on how many runtime events occur. Duration only sets the window. A mostly idle process traced for a minute may produce a tiny file, while a gateway performing millions of goroutine wake-ups per second can produce tens of megabytes in a second. Estimate size from activity, not from seconds.
- Does the tracer capture a stack for every recorded event?No. Stacks are attached to specific events where they are useful — goroutine creation, blocking, unblocking, syscall boundaries — not to every record, because attaching a stack is comparatively expensive. That is one reason the tracer is affordable, and also why a trace is a poor substitute for a CPU profile when you want hot call paths.
saying these in an interview costs you the question
- Says the execution tracer samples at a fixed frequency
- Claims tracing is free and can be left on permanently
- Expects the trace to list every function call
- Thinks trace size is set by duration alone
- Reaches for a trace to find a hot loop