skip to content

Your Go worker pool's p99 is uneven across workers while a CPU sits idle — what in the run queues explains it?

level: seniorimportance: should knowfreq 34%

answer

  1. the runtime never promised worker fairness
  2. wakes happen on the sender's processor
  3. nothing spills to the shared queue here
  4. an idle P has to go and find it
  5. microsecond jobs pay the search

basics

~20 s

Work concentrates on the P that created it. A worker woken by a send on the jobs channel becomes runnable on the sender's P, so idle Ps get work only by stealing, which costs a search and takes the runnext slot last.

solid answer

~50 s

It is queue locality rather than a bug. When the producer sends on the jobs channel it hands the value to a parked worker and readies that worker on its **own** P, in that P's `runnext` slot, so short jobs tend to run on the producer's P and the rest pile into that P's local run queue. Other Ps see nothing in the global run queue, because goroutines only spill there when a local queue overflows its 256 entries, so they must steal — which costs a spinning thread, up to four randomised passes, a grab of half a queue, and takes the victim's `runnext` only on the last pass. For jobs lasting microseconds that latency is a visible share of the work, and it lands on whichever workers sat at the back, which is what the per-worker histogram is showing. Before tuning, split wait time from service time; then batch jobs per send, and judge the pool by end-to-end throughput rather than per-worker fairness.

code

go · 13 lines
go
jobs := make(chan Job, 1024)
results := make(chan Result, 1024)

var wg sync.WaitGroup
for w := 0; w < 8; w++ {
	wg.Add(1)
	go func() {
		defer wg.Done()
		for j := range jobs {
			results <- do(j) // a few microseconds of work
		}
	}()
}

go deeper

for a junior

Know that goroutines are not spread evenly by default: the runtime prefers to run newly woken work on the processor that woke it.

for a middle

Explain the path a job takes — a send wakes a parked worker into the sender's P's fast slot — and why an otherwise idle P has to steal to get any of it.

for a senior

Diagnose before tuning: separate wait from service time per worker, and judge whether hand-off and stealing latency is a real share of a job that lasts microseconds.

for a principal

Own what the benchmark reports and why: aggregate throughput and latency with a stated pool-sizing rule, rather than tuning against a fairness number the runtime never guaranteed.

## The symptom A pool of worker goroutines reads jobs from one channel and writes to another. Each job takes single-digit microseconds and they are near-identical in cost. A per-worker latency histogram shows most workers tight around the mean while one or two carry a long tail, and at the same time the machine has a core doing nothing. Nothing in the application is unfair: every worker runs the same loop over the same channel. ## The mechanism Follow one job through the runtime. **The producer wakes a worker on its own P.** When the producer sends on the jobs channel and a worker is parked in `<-jobs`, the send hands the value straight to that waiting worker and marks it runnable. "Marks it runnable" means putting it in the **current P's `runnext` slot** — the producer's P, not an idle one. The woken worker therefore tends to run on the P that produced the job, immediately after the producer yields or blocks. **Its P's local run queue grows.** Each subsequent send readies another worker on the same P. The previous occupant of `runnext` is bumped to the tail of that P's 256-entry local run queue. Work is now concentrated: one P holds a queue of ready workers, others hold nothing. **The global run queue does not help.** Goroutines only spill into the global queue when a local queue overflows its 256 entries, and a pool of a few workers never gets close. So the other Ps have no shared pile to draw from. **The idle P must steal, and stealing takes time.** An idle P's thread checks the global queue, polls the network poller, marks itself spinning and makes up to four randomised passes over the other Ps, grabbing about half of a victim's local queue when it finds one. The victim's `runnext` is only taken on the final pass, after a deliberate pause. If no thread is currently spinning, one has to be woken first. That whole sequence is cheap in absolute terms — microseconds — but a job that lasts microseconds pays it as a visible share of its own latency, and it falls on whichever worker happened to be at the back of the concentrated queue. That is the tail your histogram is showing. ## Confirming it before acting Split the measurement in two. Record, per worker, both the **service time** (how long the job itself took) and the **wait** (enqueue to start). If service times match across workers and only the waits differ, the jobs really are uniform and the delay is placement and stealing latency. If service times differ, the jobs are not uniform and the run queues are innocent — you are looking at a data-skew problem wearing a scheduler costume. The other diagnosis worth ruling out is that the tail worker is not runnable at all: a worker blocked in a syscall or contending on a mutex looks slow for entirely different reasons, and the fix would be different too. ## What actually changes the number - **Batch the jobs.** Sending a slice of jobs per channel operation instead of one removes hand-offs altogether. For microsecond work this is usually the single biggest win, because the scheduling cost per job goes to near zero. - **Do trivial jobs inline.** If a job costs less than the hand-off, a pool is a net loss. Measure the same workload with no pool at all before defending the pool's existence. - **Size the pool against the number of Ps, not against intuition.** Beyond that number the extra workers only queue behind each other; below it you leave Ps for the runtime to fill by stealing. ## What does not change it - Adding more workers. It adds goroutines to the same concentrated queues. - Buffering the jobs channel. A send with a parked receiver still hands off directly and still readies that worker on the sender's P; a send that merely fills the buffer readies nobody. Buffering decouples producer from consumer, which is worth having for other reasons, but it does not spread wakes across Ps. - Trying to pin workers. `runtime.LockOSThread` binds a goroutine to a thread, not to a P, and it does not give you placement control — it takes flexibility away. ## The benchmark to report Since the point of the exercise is comparing two pool sizes, decide up front what the benchmark reports: **end-to-end throughput and total p99 for the whole batch**, with allocations from `-benchmem`, not per-worker fairness. Per-worker balance is not something the runtime ever promised, and optimising for it will push you toward changes that make the aggregate worse. Run the comparison long enough that the stealing transient is amortised, and state the pool-sizing rule you derived so the next person does not re-tune it by feel.

  • How would you confirm the imbalance is scheduler placement rather than uneven job cost?
    Record per-worker histograms of two separate quantities: the job's own service time and the wait from enqueue to start. If service times match across workers and only the waits differ, the jobs are uniform and the delay is placement plus stealing latency. If service times differ, the work is skewed and the run queues are innocent.
  • Does buffering the jobs channel change which P runs a woken worker?
    No. A send that finds a parked receiver hands off directly and readies that worker on the sender's P, buffered or not; a send that merely fills the buffer readies nobody at all. Buffering decouples the producer from a slow consumer, which is worth having, but only stealing spreads woken workers across Ps.
  • You are benchmarking two pool sizes. What should the benchmark report?
    End-to-end throughput and total p99 for the whole batch, with allocations from `-benchmem` — not per-worker fairness, which the runtime never promised. Past the number of Ps, extra workers only queue behind each other. For microsecond jobs the real lever is usually batching several jobs per channel operation, which removes the hand-offs entirely.

saying these in an interview costs you the question

  • Blames the OS for descheduling one thread
  • Adds more workers to fix the imbalance
  • Assumes goroutines are spread round-robin across Ps
  • Expects the global run queue to distribute every job
  • Treats per-worker fairness as the benchmark result