skip to content

In a Go service, throughput plateaus while CPU sits near 30% — how do block and mutex profiles find the lock?

level: seniorimportance: must knowfreq 45%

answer

  1. parked goroutines do not burn CPU
  2. which profile names the culprit?
  3. both are cumulative, so diff them
  4. channel waits appear in only one
  5. count against delay tells the shape

basics

~20 s

Low CPU with flat throughput means goroutines are parked, not computing. Enable both profilers coarsely on one instance and diff two captures: the mutex profile's top delay names the critical section, and the block profile separates lock waits from channel waits.

solid answer

~50 s

Low CPU with flat throughput says goroutines are parked, not computing — Go's runtime parks a goroutine waiting on a contended mutex rather than spinning on it, so a serialised service looks idle. I would enable `runtime.SetMutexProfileFraction` and `runtime.SetBlockProfileRate` at coarse rates on one instance, take two captures a known interval apart, and diff them, because both profiles are cumulative. The mutex profile is the one that names the culprit: it attributes waiters' delay to the holder's unlock, so a single middleware counter guarded by one `sync.Mutex` collapses into one huge entry instead of being smeared over every handler queued behind it. The block profile then separates lock waits from channel waits — a limiter token or a queue receive appears only there — and its contentions count distinguishes many short waits, meaning every request funnels through one lock, from a few long ones, meaning someone holds the lock across slow work.

code

go · 10 lines
go
type stats struct {
	mu     sync.Mutex
	counts map[string]int
}

func (s *stats) record(route string) {
	s.mu.Lock()
	defer s.mu.Unlock()
	s.counts[route]++
}

go deeper

for a junior

Know which profile to reach for: the mutex profile when you suspect a lock, the block profile when you want to see what goroutines are waiting on. Remember both must be enabled before they record anything.

for a middle

Explain why the mutex profile points at the lock holder while the block profile shows the waiters, and why goroutines parked on a contended mutex leave CPU utilisation low rather than high.

for a senior

Walk the whole procedure: enable coarsely on one instance, diff two captures under load, use the mutex profile to name the critical section, use the block profile to rule channel waits in or out, and read contentions against delay before choosing a fix.

for a principal

Own the follow-through: whether the fix is sharding, atomics or removing the shared state, what the next ceiling will be, and what you change so the next occurrence is diagnosed from data already being collected.

## Reading the symptom first "Throughput will not climb, CPU is well under the quota" is a specific signature. It rules out being compute-bound and it rules out the scheduler being saturated. Work is arriving and something is making the goroutines that would do it stop. In Go this looks idle because of how the runtime handles a contended lock: a goroutine that cannot take a `sync.Mutex` spins only very briefly and is then parked, its P handed to something else. A serialised program therefore does not burn CPU the way a spinlock-based one would. Low CPU is evidence *for* contention, not against it. ## The shape that produces it The common cause is a piece of shared state that every request touches. A middleware layer wrapping every handler, holding one in-process limiter and a small struct of counters behind a single mutex, is enough: ```go type stats struct { mu sync.Mutex counts map[string]int } func (s *stats) record(route string) { s.mu.Lock() defer s.mu.Unlock() s.counts[route]++ } ``` The critical section is a map increment — microseconds at worst. It does not matter. If every request passes through it, the whole service is a queue for one lock, and its ceiling is one over the lock's service time regardless of how many cores you add. ## The measurement 1. **Enable both profilers on one instance**, at coarse rates: a mutex fraction in the tens or hundreds, a block rate in the tens of microseconds. Coarse sampling is weighted towards long waits, so it preserves the ranking you care about. 2. **Take a delta.** Both profiles accumulate for the life of the process, so a single capture from a week-old process is dominated by history. Capture, wait a fixed interval under load, capture again, and work with the difference. 3. **Start with the mutex profile.** It fires only on genuine contention and records the **holder's** stack at the release, so the top entry by delay points straight at `record` — the critical section whose holding time cost everyone else. Had you started with the block profile, the same contention would be spread across every handler that happened to be waiting. 4. **Cross-check with the block profile.** This is where you separate the two kinds of waiting the symptom could be: waiting for a lock, or waiting on a channel. A limiter implemented as a token channel, or a handoff to a bounded queue, shows up only in the block profile. If the block profile's request-path stacks are dominated by a channel receive rather than by `Lock`, your ceiling is the limiter's capacity, not a critical section — a completely different fix. 5. **Read count against delay.** Enormous contentions count with modest per-event delay is the "everything goes through one lock" shape. A small count with huge delay is the "someone holds the lock across an RPC or a disk write" shape. ## The block profile's built-in noise Sort a block profile by total delay and the top entries are usually background goroutines parked in a `select`, waiting for work they are supposed to be waiting for. That is not contention and it is why the profile has a reputation for being useless. The discipline is to look only at stacks that lie on the request path and to use the mutex profile as the arbiter for lock questions. An entry in the mutex profile always represents time a goroutine genuinely lost to another goroutine's lock. ## What the two profiles will not tell you - Time in syscalls and network I/O appears in neither. If both profiles are quiet and throughput is still capped, look outside the runtime's synchronisation: a connection pool with a small maximum, a downstream service, or a per-instance concurrency limit. - Contention on locks internal to the runtime or to a library that does not use `sync.Mutex` will not necessarily show as a mutex sample. - Neither profile tells you the fix. They tell you which critical section is expensive; whether the answer is to shard the state, to use an atomic counter, to accumulate per-request and flush periodically, or to remove the shared state entirely, is a design call. ## Closing the loop After the change, repeat the same delta capture under the same load. The mutex profile's top entry should shrink and something else should become the ceiling — usually the next-largest shared thing. Contention work is iterative, and the profiles are how you tell you moved the bottleneck rather than merely moved the code. ## What a strong answer sounds like Name the symptom's meaning (parked, not spinning), enable both profilers deliberately at coarse rates on one instance, diff two captures, use the mutex profile to name the critical section and the block profile to rule channels in or out, and read count and delay together before proposing a fix.

  • The block profile's largest delay entry is a background goroutine parked in select. Is that your bottleneck?
    Almost certainly not. A goroutine that waits for work all day accumulates enormous delay by design. Idle waiting dominates raw block-profile totals, so restrict attention to stacks on the request path and let the mutex profile arbitrate lock questions, since it only fires on real contention.
  • The mutex profile is nearly empty but throughput is still capped. Where do you look next?
    Outside the runtime's synchronisation primitives. Neither profile covers syscalls or network waits, so check a connection pool's maximum, a semaphore or token channel limiting concurrency, a slow downstream dependency, or a configured per-instance request limit. The block profile's channel stacks are the first place to look.
  • Why does a delta between two captures matter more here than for a CPU profile?
    Because the block and mutex profiles accumulate for the process lifetime rather than covering a window. A capture from a week-old instance mixes today's contention with a startup burst. Two captures a known interval apart under known load give you contention per unit of time, which is what you can compare after a fix.
  • How would you tell a long-held lock from a lock everyone touches?
    Compare the contentions count with the delay. A very large count with modest delay per event means the critical section is short but universal, so the fix is to shard or remove the shared state. A small count with large delay means the lock is held across slow work, so the fix is to shorten the critical section.

saying these in an interview costs you the question

  • Concludes low CPU rules out lock contention
  • Reads a single lifetime capture instead of a delta
  • Sorts the block profile by delay and blames the top idle wait
  • Looks for the bottleneck under the waiting callers, not the holder
  • Enables both profilers at full rate across the whole fleet mid-incident
  • Assumes a quiet mutex profile means there is no bottleneck at all