skip to content

What the Tracer Shows

The execution tracer records goroutine transitions, syscalls, GC phases and P states; go tool trace turns that into timelines showing why a goroutine waited rather than where it burned CPU.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

explore

questions

12

In `go tool trace`'s goroutine analysis, what does a goroutine's scheduler wait time mean?

level: juniorimportance: should knowfreq 38%

answer

  1. ready, but nobody ran it
  2. queueing delay, not blocking
  3. measured from wake-up to first instruction
  4. grows when runnable goroutines outnumber Ps
  5. an idle CPU graph proves nothing

basics

~20 s

Scheduler wait is time the goroutine was runnable but not running: it had work to do and sat queued, waiting for one of the GOMAXPROCS logical processors to pick it up. It is queueing delay, not blocking.

solid answer

~40 s

The goroutine analysis page of `go tool trace` splits each goroutine's lifetime into non-overlapping buckets: execution, network wait, sync block, blocking syscall, scheduler wait, GC sweeping and GC pause. Scheduler wait is the gap between the instant the goroutine became runnable — created by `go`, woken by a channel operation, a mutex unlock, a timer, or I/O readiness — and the instant a P actually started running it. It is pure queueing delay: the goroutine had nothing left to wait for and still did not run. It grows when runnable goroutines outnumber the available Ps, when GC background mark workers and assists are occupying Ps, or when the OS is not promptly scheduling the process's threads. Because it is not blocking, you do not fix it by making that goroutine's own code faster.

go deeper

for a junior

Be ready to say in one sentence that scheduler wait means runnable but not yet running, and to name at least two other buckets the goroutine analysis page shows, such as network wait and sync block.

for a middle

Explain the mechanics: what event makes a goroutine runnable, why only GOMAXPROCS goroutines can run Go code at once, and why garbage collection cycles push this number up without any change to your code.

for a senior

Show that you use the bucket split as a routing decision on a live incident, and that you do not accept an idle host CPU graph as evidence against processor contention.

for a principal

Frame it as a capacity question: whether the answer is bounding goroutine concurrency, cutting allocation, or giving the process more processors, and what each of those costs the team to change.

## What the goroutine analysis view is A **goroutine** is Go's unit of concurrency: a function started with `go`, multiplexed by the Go runtime onto operating-system threads. `go tool trace` opens a runtime execution trace and, among its pages, serves a **goroutine analysis** table. That table groups goroutines by the function they started in and breaks each one's total lifetime into buckets that do not overlap and add up to the total: - **Execution** — actually running Go code on a processor. - **Network wait** — parked by the runtime's network poller waiting for a socket to become readable or writable. - **Sync block** — blocked on a channel operation, a `sync.Mutex`, a `sync.WaitGroup` and similar. - **Blocking syscall** — inside a system call that blocked (a file read, a DNS lookup, a `cgo` call). - **Scheduler wait** — runnable, but not yet running. - **GC sweeping** and **GC pause** — sweeping work charged to this goroutine, and time it was frozen for a stop-the-world phase. ## Scheduler wait precisely Go's runtime keeps a bounded number of logical processors, called **P**s; `GOMAXPROCS` sets how many exist, and a goroutine can only run Go code while it is attached to one. A goroutine becomes **runnable** at a definite instant: when `go f()` creates it, when the value it was waiting for is sent on a channel, when the mutex it wanted is released, when a timer fires, or when the poller sees its socket is ready. From that instant it sits on a run queue until a P takes it and starts executing. **That interval is the scheduler wait time.** The important property is that nothing external is holding the goroutine back. Its input has arrived. Every millisecond in this bucket is latency added by queueing alone, and it appears in your end-to-end numbers exactly as if the work had been slow. ## What drives it up - **More runnable goroutines than Ps.** If you have eight Ps and two hundred goroutines that all became runnable at once, one hundred and ninety-two of them are accruing scheduler wait by definition. Wake storms are the classic source: closing a channel that many goroutines are receiving from, or a batch arriving all at once. - **Ps consumed by garbage collection.** During a collection cycle the runtime dedicates background mark workers to roughly a quarter of the Ps, and allocating goroutines can be charged mark assist work on top. Your goroutines are then competing for fewer effective Ps, and their scheduler wait rises even though your own code did not change. - **The operating system not running the threads.** If the host is oversubscribed or the process is CPU-throttled, a thread holding a P may itself not be on a core. From inside the trace this looks like goroutines waiting to be scheduled. - **`GOMAXPROCS` smaller than the CPU you thought you had.** The host CPU graph can look half idle while every P Go owns is saturated. ## Reading it, and what not to conclude Sort the goroutine analysis table by total time and look at which bucket dominates for the goroutines on your critical path. Dominant **scheduler wait** points at contention for processors; dominant **sync block** points at your own coordination (a channel that nobody is receiving from, a hot mutex); dominant **network wait** points downstream, at whatever you are calling; dominant **blocking syscall** points at the kernel — files, DNS, `cgo`. The four suggest completely different fixes, which is the whole reason the tracer separates them. `go tool trace` also serves a **scheduler latency profile** in the same pprof shape as a CPU profile, attributing time-to-be-scheduled to stacks, so you can see which paths queue rather than just that queueing exists. Two misreadings are common. The first is treating scheduler wait as blocking — it is the opposite of blocking; the goroutine is ready. The second is assuming an idle host CPU graph proves scheduler wait must be zero: Go can only use as many cores as it has Ps, and short wake bursts vanish into a one-minute CPU average while still costing your p99 milliseconds. ## The fix shape Because it is queueing, you reduce it by reducing competition: bound how many goroutines are runnable at once rather than spawning one per item, spread wake-ups instead of releasing them in a thundering herd, cut allocation so collection cycles take fewer Ps away, and make sure the process actually has the processors you believe it has.

  • How does scheduler wait differ from the sync block bucket in the same table?
    Sync block is time with nothing to do — parked on a channel operation or a mutex, waiting for another goroutine. Scheduler wait begins only after that wake-up, when the goroutine is ready and merely queued for a processor. A slow handoff lands in sync block; a slow pickup after the handoff lands in scheduler wait. Both add to end-to-end latency and they have different fixes.
  • The host's CPU graph shows idle cores. Can scheduler wait still be significant?
    Yes, and it often is. Go only runs goroutines on the Ps it has, so if GOMAXPROCS is below the core count the spare cores are irrelevant. Wake bursts are also short: two hundred goroutines queueing for eight Ps for five milliseconds barely moves a per-minute CPU average but is plainly visible in the trace.
  • Which part of go tool trace tells you where the queueing comes from, not just how much there is?
    The scheduler latency profile. It is served in the same pprof form as a CPU profile and attributes time-to-be-scheduled to stacks, so you can see which wake-up paths wait longest instead of only reading a per-goroutine total. The goroutine analysis table gives the magnitude; that profile gives the attribution.

It is the time a boarding pass holder spends at the gate after the flight is called: nothing is missing, there is simply no jet bridge free yet.

saying these in an interview costs you the question

  • Says scheduler wait means the goroutine was blocked on a channel
  • Reads it as time waiting for the OS scheduler rather than a P
  • Assumes an idle host CPU graph proves scheduler wait is zero
  • Confuses it with time parked inside a blocking syscall
  • Tries to fix it by optimising the goroutine's own function body
open as a page

In `go tool trace`, what does a goroutine's GC mark assist time mean, and which goroutines pay it?

level: middleimportance: should knowfreq 30%

basics

~20 s

Mark assist is garbage-collection marking work done by an application goroutine itself, charged in proportion to how much it allocated during a collection cycle. Allocation-heavy goroutines pay it, and it lands directly in their latency.

open as a page

What does the timeline view in go tool trace show, and which other views does it serve?

level: middleimportance: should knowfreq 42%

basics

~20 s

The timeline in go tool trace plots wall-clock time horizontally, one row per P (Go's scheduler slot), showing which goroutine ran where, plus GC and heap rows. The same page serves goroutine analysis, blocking profiles, scheduler latency, and minimum mutator utilization.

open as a page

What does Go's execution tracer actually record, and why keep captures to a few seconds?

level: middleimportance: should knowfreq 38%

basics

~20 s

Go's execution tracer records every traced runtime event with a timestamp — goroutine lifecycle, block reasons, syscalls, scheduler and GC phases — rather than sampling. That exhaustiveness makes files grow with activity, so captures stay short.

open as a page

In Go's runtime/trace, how does a task differ from a region, and how is each one ended?

level: middleimportance: should knowfreq 34%

basics

~20 s

A runtime/trace task is a logical operation that may span goroutines: NewTask returns a context carrying it, and Task.End closes it. A region is one interval inside a single goroutine, ended by the goroutine that started it.

open as a page

Your Go worker pool's p99 job latency climbs while host CPU sits half idle — how do you read a `go tool trace` to tell scheduler wait from blocking syscalls?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Read the goroutine analysis for the worker goroutines and see which bucket dominates. Scheduler wait means they were queued for a processor; blocking syscall means they were parked in the kernel. Idle CPU fits both.

open as a page

A Go service stalls for four seconds twice a day; how do you get an execution trace of it?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Run runtime/trace.FlightRecorder continuously: it keeps a recent window of trace data in memory, and an in-process watchdog that notices the stall calls WriteTo to dump it. On-demand capture starts too late for an event this short.

open as a page

In runtime/trace, how do you annotate a stage that fans work out to several goroutines?

level: seniorimportance: should knowfreq 26%

basics

~20 s

Pass the context returned by trace.NewTask into every goroutine so their events land under one task, and let each goroutine start and end its own region. Never share a trace.Region across goroutines, and end the task only after the workers finish.

open as a page

In runtime/trace, how would you find which stage of a hung ingestion workflow never finished?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Capture a trace while the process is still stuck and look for the task with no end event. Its last recorded events are the start of a region that never ended and the log messages before it, which name the stage and the item.

open as a page

How do you capture a Go execution trace from a running program or a test, and how do you open it?

level: juniorimportance: nice to knowfreq 26%

basics

~10 s

Call runtime/trace.Start with a file and defer trace.Stop, or run go test -trace=out, or fetch /debug/pprof/trace?seconds=5 from a service. Open the resulting file with go tool trace, which serves a browser UI.

open as a page

When should you guard runtime/trace.Logf with trace.IsEnabled, and what does the guard actually save?

level: middleimportance: nice to knowfreq 20%

basics

~20 s

Guard when building the arguments costs something. trace.Logf already skips its own formatting while tracing is off, but the caller still evaluates and boxes the arguments; trace.IsEnabled skips that work. Treat the answer as advisory only.

open as a page

Your `go tool trace` shows a stop-the-world pause of several milliseconds instead of the usual short one. What explains it?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

A stop-the-world pause is usually dominated by the time it takes to stop every logical processor, not by the work done once stopped. Suspect threads the host is not scheduling, or a caller such as runtime.ReadMemStats.

open as a page