skip to content

In `go tool trace`'s goroutine analysis, what does a goroutine's scheduler wait time mean?

level: juniorimportance: should knowfreq 38%

answer

  1. ready, but nobody ran it
  2. queueing delay, not blocking
  3. measured from wake-up to first instruction
  4. grows when runnable goroutines outnumber Ps
  5. an idle CPU graph proves nothing

basics

~20 s

Scheduler wait is time the goroutine was runnable but not running: it had work to do and sat queued, waiting for one of the GOMAXPROCS logical processors to pick it up. It is queueing delay, not blocking.

solid answer

~40 s

The goroutine analysis page of `go tool trace` splits each goroutine's lifetime into non-overlapping buckets: execution, network wait, sync block, blocking syscall, scheduler wait, GC sweeping and GC pause. Scheduler wait is the gap between the instant the goroutine became runnable — created by `go`, woken by a channel operation, a mutex unlock, a timer, or I/O readiness — and the instant a P actually started running it. It is pure queueing delay: the goroutine had nothing left to wait for and still did not run. It grows when runnable goroutines outnumber the available Ps, when GC background mark workers and assists are occupying Ps, or when the OS is not promptly scheduling the process's threads. Because it is not blocking, you do not fix it by making that goroutine's own code faster.

go deeper

for a junior

Be ready to say in one sentence that scheduler wait means runnable but not yet running, and to name at least two other buckets the goroutine analysis page shows, such as network wait and sync block.

for a middle

Explain the mechanics: what event makes a goroutine runnable, why only GOMAXPROCS goroutines can run Go code at once, and why garbage collection cycles push this number up without any change to your code.

for a senior

Show that you use the bucket split as a routing decision on a live incident, and that you do not accept an idle host CPU graph as evidence against processor contention.

for a principal

Frame it as a capacity question: whether the answer is bounding goroutine concurrency, cutting allocation, or giving the process more processors, and what each of those costs the team to change.

## What the goroutine analysis view is A **goroutine** is Go's unit of concurrency: a function started with `go`, multiplexed by the Go runtime onto operating-system threads. `go tool trace` opens a runtime execution trace and, among its pages, serves a **goroutine analysis** table. That table groups goroutines by the function they started in and breaks each one's total lifetime into buckets that do not overlap and add up to the total: - **Execution** — actually running Go code on a processor. - **Network wait** — parked by the runtime's network poller waiting for a socket to become readable or writable. - **Sync block** — blocked on a channel operation, a `sync.Mutex`, a `sync.WaitGroup` and similar. - **Blocking syscall** — inside a system call that blocked (a file read, a DNS lookup, a `cgo` call). - **Scheduler wait** — runnable, but not yet running. - **GC sweeping** and **GC pause** — sweeping work charged to this goroutine, and time it was frozen for a stop-the-world phase. ## Scheduler wait precisely Go's runtime keeps a bounded number of logical processors, called **P**s; `GOMAXPROCS` sets how many exist, and a goroutine can only run Go code while it is attached to one. A goroutine becomes **runnable** at a definite instant: when `go f()` creates it, when the value it was waiting for is sent on a channel, when the mutex it wanted is released, when a timer fires, or when the poller sees its socket is ready. From that instant it sits on a run queue until a P takes it and starts executing. **That interval is the scheduler wait time.** The important property is that nothing external is holding the goroutine back. Its input has arrived. Every millisecond in this bucket is latency added by queueing alone, and it appears in your end-to-end numbers exactly as if the work had been slow. ## What drives it up - **More runnable goroutines than Ps.** If you have eight Ps and two hundred goroutines that all became runnable at once, one hundred and ninety-two of them are accruing scheduler wait by definition. Wake storms are the classic source: closing a channel that many goroutines are receiving from, or a batch arriving all at once. - **Ps consumed by garbage collection.** During a collection cycle the runtime dedicates background mark workers to roughly a quarter of the Ps, and allocating goroutines can be charged mark assist work on top. Your goroutines are then competing for fewer effective Ps, and their scheduler wait rises even though your own code did not change. - **The operating system not running the threads.** If the host is oversubscribed or the process is CPU-throttled, a thread holding a P may itself not be on a core. From inside the trace this looks like goroutines waiting to be scheduled. - **`GOMAXPROCS` smaller than the CPU you thought you had.** The host CPU graph can look half idle while every P Go owns is saturated. ## Reading it, and what not to conclude Sort the goroutine analysis table by total time and look at which bucket dominates for the goroutines on your critical path. Dominant **scheduler wait** points at contention for processors; dominant **sync block** points at your own coordination (a channel that nobody is receiving from, a hot mutex); dominant **network wait** points downstream, at whatever you are calling; dominant **blocking syscall** points at the kernel — files, DNS, `cgo`. The four suggest completely different fixes, which is the whole reason the tracer separates them. `go tool trace` also serves a **scheduler latency profile** in the same pprof shape as a CPU profile, attributing time-to-be-scheduled to stacks, so you can see which paths queue rather than just that queueing exists. Two misreadings are common. The first is treating scheduler wait as blocking — it is the opposite of blocking; the goroutine is ready. The second is assuming an idle host CPU graph proves scheduler wait must be zero: Go can only use as many cores as it has Ps, and short wake bursts vanish into a one-minute CPU average while still costing your p99 milliseconds. ## The fix shape Because it is queueing, you reduce it by reducing competition: bound how many goroutines are runnable at once rather than spawning one per item, spread wake-ups instead of releasing them in a thundering herd, cut allocation so collection cycles take fewer Ps away, and make sure the process actually has the processors you believe it has.

  • How does scheduler wait differ from the sync block bucket in the same table?
    Sync block is time with nothing to do — parked on a channel operation or a mutex, waiting for another goroutine. Scheduler wait begins only after that wake-up, when the goroutine is ready and merely queued for a processor. A slow handoff lands in sync block; a slow pickup after the handoff lands in scheduler wait. Both add to end-to-end latency and they have different fixes.
  • The host's CPU graph shows idle cores. Can scheduler wait still be significant?
    Yes, and it often is. Go only runs goroutines on the Ps it has, so if GOMAXPROCS is below the core count the spare cores are irrelevant. Wake bursts are also short: two hundred goroutines queueing for eight Ps for five milliseconds barely moves a per-minute CPU average but is plainly visible in the trace.
  • Which part of go tool trace tells you where the queueing comes from, not just how much there is?
    The scheduler latency profile. It is served in the same pprof form as a CPU profile and attributes time-to-be-scheduled to stacks, so you can see which wake-up paths wait longest instead of only reading a per-goroutine total. The goroutine analysis table gives the magnitude; that profile gives the attribution.

It is the time a boarding pass holder spends at the gate after the flight is called: nothing is missing, there is simply no jet bridge free yet.

saying these in an interview costs you the question

  • Says scheduler wait means the goroutine was blocked on a channel
  • Reads it as time waiting for the OS scheduler rather than a P
  • Assumes an idle host CPU graph proves scheduler wait is zero
  • Confuses it with time parked inside a blocking syscall
  • Tries to fix it by optimising the goroutine's own function body