skip to content

Your Go worker pool's p99 job latency climbs while host CPU sits half idle — how do you read a `go tool trace` to tell scheduler wait from blocking syscalls?

level: seniorimportance: should knowfreq 44%

answer

  1. the buckets are mutually exclusive
  2. idle CPU fits several opposite stories
  3. runnable and queued, or parked in the kernel?
  4. check whether the processors were busy at all
  5. unbuffered results channel parks every worker

basics

~20 s

Read the goroutine analysis for the worker goroutines and see which bucket dominates. Scheduler wait means they were queued for a processor; blocking syscall means they were parked in the kernel. Idle CPU fits both.

solid answer

~60 s

Take a trace over a bad window and go straight to the goroutine analysis, grouped by the worker function. The buckets are mutually exclusive, so whichever dominates routes the whole investigation. Dominant **scheduler wait** means the workers were ready and could not get a processor — check the timeline's runnable-goroutine count and the per-processor rows to see whether Ps were busy with your code, with garbage-collection mark workers, or not running at all, and check whether `GOMAXPROCS` matches the CPU you actually have, since Go cannot use cores it has no P for. Dominant **blocking syscall** means the workers were in the kernel — file, DNS or `cgo` calls — and idle CPU is exactly what you would expect; the fix is more concurrency or fewer syscalls, not more cores. Dominant **sync block** points at your own coordination, typically workers parked sending on an unbuffered results channel that a slow consumer is draining. Use the scheduler latency profile to attribute queueing to stacks once you know queueing is the problem.

code

go · 13 lines
go
jobs := make(chan Job, 1024)
results := make(chan Result) // unbuffered: senders park here

var wg sync.WaitGroup
for w := 0; w < 8; w++ {
	wg.Add(1)
	go func() {
		defer wg.Done()
		for j := range jobs {
			results <- process(j)
		}
	}()
}

go deeper

for a junior

Know that a trace can tell you whether a goroutine was running, queued, blocked in a syscall or waiting on a channel, and that those are different problems.

for a middle

Explain why goroutines parked in syscalls leave the CPU idle, and why runnable goroutines can also leave it idle when the process has fewer logical processors than the host has cores.

for a senior

Walk an ordered investigation under pressure: dominant bucket first, then processor occupancy, then one change and a re-trace, and say why the candidate fixes would harm each other if applied blind.

for a principal

Own the standing decision about what the team measures by default, so the next incident starts from evidence rather than from a CPU graph that cannot distinguish these cases.

## Why the symptom is ambiguous "p99 is climbing and the CPU graph is half idle" is one of the most common shapes an on-call engineer sees, and on its own it discriminates nothing. Idle CPU is consistent with at least four completely different stories: - The workers are **parked in the kernel** and consuming no CPU by definition, so the pool is too small for the amount of waiting it has to do. - The workers are **runnable and queued** behind too few logical processors, while the host reports idle cores Go will never use. - The workers are **blocked on each other** — a channel with nobody receiving, or a contended mutex. - The workers are **fine** and the queue in front of them is backing up because the producer or the consumer is the bottleneck. An execution trace separates these because it records state transitions per goroutine with timestamps, so it can say which of those four the time actually went to. ## Step one: the goroutine analysis, grouped by the worker function The goroutine analysis page splits each goroutine's total lifetime into non-overlapping buckets: execution, network wait, sync block, blocking syscall, scheduler wait, GC sweeping and GC pause. Group by the worker function and read the dominant bucket: | Dominant bucket | What it means | What it does not mean | |---|---|---| | Scheduler wait | Runnable, waiting for a processor | Not blocked; the input had arrived | | Blocking syscall | Inside the kernel: file, DNS, `cgo` | Not a CPU shortage | | Network wait | Parked on socket readiness | Usually a downstream problem | | Sync block | Channel or mutex coordination | Your program's own structure | | Execution | Genuinely running your code | This is a CPU-cost problem | ## Step two: follow the dominant bucket **If scheduler wait dominates.** The workers were ready and nothing ran them. Look at the timeline's counter rows for runnable goroutines and for threads, and at the per-processor rows. Three distinct pictures follow. If the processors were busy running your own code, you are simply out of parallelism and the host graph is misleading — Go can only run as many goroutines as it has Ps, so idle cores on the host are irrelevant unless the process is configured to use them. If the processors were busy with background mark workers, garbage collection is eating your parallelism and the real lever is allocation rate. If the processors were doing nothing at all during the wait, the threads themselves were not being run by the operating system, and the problem is outside the process. **If blocking syscall dominates.** The workers were parked in the kernel. Idle CPU is the expected observation, not an anomaly, and adding cores changes nothing. The levers are concurrency (more workers, since each one is mostly waiting), doing fewer syscalls per job, or batching. Watch for the second-order effect: goroutines blocked in syscalls hand their processor off, so the thread count can climb well above `GOMAXPROCS` in exactly this scenario. **If sync block dominates.** This is usually the pool's own plumbing rather than the outside world. The classic case in a jobs-and-results pool is an unbuffered results channel: every worker finishes its job, parks on the send, and no worker returns to the jobs channel, so the queue in front of the pool grows while the workers look busy waiting. The trace shows workers alternating short execution spans with long sync block spans in lockstep, which is the fingerprint of backpressure from the consumer rather than of slow work. **If execution dominates.** You have an ordinary CPU-cost problem in the job itself, and a CPU profile is the right next tool, not the trace. ## Step three: attribute, then change one thing Once you know the shape, `go tool trace` will serve a scheduler latency profile in pprof form, attributing time-to-be-scheduled to stacks, so you can find which wake-up paths queue rather than just knowing that queueing exists. Then make one change and re-trace the same window shape. The reason to be strict about that is that the four stories above have fixes that actively harm each other: adding workers helps a syscall-bound pool and makes a processor-starved one worse; adding cores helps the starved one and does nothing for the syscall-bound one. ## The trap The trap in this scenario is treating the CPU graph as evidence. It is an average over a coarse window, over cores Go may not be allowed to use, and it counts a goroutine queueing for a processor exactly the same as a goroutine sleeping on a socket: as idle. The buckets are the evidence; the CPU graph is the symptom that made you look.

  • Blocking syscall dominates for the workers. What does that say about pool size?
    That the pool is too small for the amount of waiting it does. A goroutine parked in a syscall hands off its processor and consumes no CPU, so idle cores are the expected observation rather than a contradiction. Raising worker count adds useful concurrency here — the opposite of the right move if the workers were queueing for processors instead.
  • Sync block dominates and the workers all park at the same point. What do you suspect?
    Backpressure from the consumer through an unbuffered or undersized results channel. Every worker finishes a job, parks on the send, and none returns to take the next job, so the jobs queue grows while the workers appear busy. The fix is on the consumer side or in the channel's capacity, not in the workers.
  • Scheduler wait dominates, and the per-processor rows show the processors were busy. What next?
    Determine what they were busy with. If it was your job code, you are out of parallelism and the host's idle cores are irrelevant unless the process can actually use them. If a large share was background mark work, garbage collection is consuming your parallelism, and allocation rate is the real lever rather than anything about the pool.
  • Why is the host CPU graph such weak evidence in this scenario?
    It averages over a window long enough to hide millisecond queueing bursts, it counts cores the process may not be configured to use, and it records a goroutine queueing for a processor and a goroutine sleeping on a socket identically, as idle. It is the symptom that makes you take a trace, not a diagnosis.

saying these in an interview costs you the question

  • Adds worker goroutines before knowing which bucket dominates
  • Treats idle host CPU as proof there is spare parallelism
  • Assumes a growing queue means the jobs themselves are slow
  • Ignores the results channel as a source of worker parking
  • Changes several knobs at once and re-measures nothing