skip to content

Block and Mutex Profiles

The off-by-default block and mutex profiles say where goroutines wait instead of compute. Interviewers reach for them the moment the scenario is idle CPU and a service that is still slow.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

questions

5

How does Go's mutex profile differ from its block profile in what each records?

level: middleimportance: must knowfreq 50%

answer

  1. one profile per side of the wait
  2. waiter's stack versus holder's stack
  3. only one of them sees channels
  4. the mutex entry is at the unlock
  5. both report contentions and delay

basics

~20 s

Go's block profile records the waiting goroutine's stack for any synchronisation wait — channels, select, mutex acquisition, WaitGroup. The mutex profile records contention on sync.Mutex and sync.RWMutex from the holder's side, attributing the waiters' delay to the unlock that released the lock.

solid answer

~50 s

They sit on opposite sides of the same wait. The block profile answers "what were my goroutines waiting on": it samples any blocking synchronisation event — a channel send or receive, a `select` with nothing ready, `sync.WaitGroup.Wait`, `sync.Cond.Wait`, and also waiting for a mutex — and records the **waiter's** stack. The mutex profile answers "which critical section made them wait": it only covers `sync.Mutex` and `sync.RWMutex`, it only fires when there was real contention, and the stack it records is the **holder's**, captured where the lock is released. So a goroutine parked on a channel appears only in the block profile, and a hot critical section shows up in the mutex profile under the function that owns the lock rather than under the callers queued behind it. Both report a contentions count and a delay in nanoseconds.

code

go · 10 lines
go
// waiting for a limiter token: block profile only,
// recorded on this goroutine's stack
<-tokens

// waiting for the lock: block profile, on this goroutine's stack
mu.Lock()
counts[route]++
// if anyone was queued, the mutex profile records the delay HERE,
// on the holder's stack, not on the waiters'
mu.Unlock()

go deeper

for a junior

Recall that the block profile covers any synchronisation wait, including channels, while the mutex profile covers only contention on sync.Mutex and sync.RWMutex. Know that both must be switched on first.

for a middle

Explain the inversion clearly: the block profile records the waiter's stack, the mutex profile records the holder's stack at the release, and both report a contentions count alongside a delay in nanoseconds.

for a senior

Demonstrate that you would not read the top delay line straight off. Idle background waits dominate the block profile, both profiles are cumulative, and the mutex profile is the one that names the offending critical section.

for a principal

Frame which of the two your teams should be able to get by default, given that one is narrow and cheap and the other is broad, noisy and easy to misread, and how that shapes what a service exposes in production.

## Two profiles, one phenomenon, opposite ends When a goroutine cannot make progress, somebody is waiting and — sometimes — somebody else is the reason. Go gives you one profile for each side of that sentence. ### The block profile: the waiter's view Enabled with `runtime.SetBlockProfileRate`, the block profile samples goroutines that **park on a synchronisation operation** and records the stack of the goroutine that parked, together with how long it stayed parked. What counts as a blocking event: - a send on a channel with no ready receiver or no buffer space; - a receive from a channel with nothing in it; - a `select` with no ready case and no `default`; - acquiring a `sync.Mutex` or `sync.RWMutex` that is already held; - `sync.WaitGroup.Wait` and `sync.Cond.Wait`. What does **not** count: time spent inside a blocking syscall or waiting on network I/O, and time spent running on a CPU. The block profile is about the runtime's own synchronisation primitives, not about the operating system. Because it records the waiter, its stacks read like "this function was stuck here for this long". That is exactly what you want when the question is *what is my request handler waiting for*. ### The mutex profile: the holder's view Enabled with `runtime.SetMutexProfileFraction`, the mutex profile is much narrower. It covers `sync.Mutex` and `sync.RWMutex` only, and it records an event only when a lock release actually had someone queued behind it — that is, when there was genuine contention. An uncontended `Lock`/`Unlock` pair produces nothing. The crucial detail, and the one interviewers probe: **the stack recorded is the stack of the goroutine that held the lock, captured at the point it releases**, and the delay attributed to it is the time other goroutines spent waiting. The profile therefore reads as "this critical section cost the rest of the program this much waiting". That inversion is deliberate. A hundred callers queued on one mutex produce a hundred different waiter stacks in the block profile, which fragments the signal; in the mutex profile they collapse into the single unlock site that is the actual bottleneck. ### Where they overlap and where they do not | | block profile | mutex profile | |---|---|---| | channel send/receive, `select` | yes | no | | `WaitGroup.Wait`, `Cond.Wait` | yes | no | | contended `sync.Mutex` / `sync.RWMutex` | yes (waiter's stack) | yes (holder's unlock stack) | | uncontended lock | no | no | | syscall or network wait | no | no | Mutex contention is the one thing both see, from opposite ends. Everything else is block-profile territory. ### Reading them Both profiles carry the same two sample values per stack: **contentions** (a count of sampled events) and **delay** (nanoseconds of waiting). Ranking by delay tells you where the time went; ranking by contentions distinguishes *a few very long waits* from *an enormous number of tiny ones*. Those two shapes call for different fixes: a long wait usually means someone holds a lock across expensive work, while a huge count of short waits usually means the critical section is fine but every request goes through it. Both accumulate for the life of the process, so a lifetime capture from a service that has been up for a week is dominated by history. Take two captures a known interval apart and work with the difference. ### The trap: idle waiting looks enormous A background goroutine that sits in `select` waiting for work all day contributes an enormous delay to the block profile and means nothing at all — it is *supposed* to be waiting. That is the single biggest reason people mistrust the block profile. The discipline is to look at stacks that lie on the request path, and to weigh contention counts alongside delay, rather than sorting by total delay and reading the top line. The mutex profile does not have this problem: it fires only when a release unblocks a waiter, so every entry in it represents a goroutine that genuinely lost time to another goroutine's lock. ### A one-line summary worth saying out loud Block profile: *who was waiting, and on what*. Mutex profile: *whose critical section made them wait*. Use the first to discover that requests are stalling, and the second to name the lock.

  • Why does the mutex profile attribute the delay to the unlock site rather than to the waiters?
    Because the unlock site names the critical section that caused the waiting, and it aggregates. Many different callers queued on one lock produce many waiter stacks but a single holder stack, so the bottleneck shows up as one large entry instead of being smeared across the call sites that happened to be unlucky.
  • A goroutine is blocked reading from a network connection. Which of the two profiles shows it?
    Neither. Both profiles cover the runtime's synchronisation primitives — channels, select, mutexes, WaitGroup, Cond — not syscalls or network I/O. Time parked on a socket read shows up in a goroutine dump of parked stacks, not in the block or mutex profile.
  • What do the contentions and delay values tell you when read together?
    Delay is total nanoseconds of waiting, contentions is the number of sampled events. High delay with a low count means someone holds the lock across slow work; high count with modest delay means the critical section is short but every request funnels through it. The two shapes need different fixes.

saying these in an interview costs you the question

  • Says the mutex profile records the waiting goroutine's stack
  • Expects channel waits to appear in the mutex profile
  • Thinks the block profile counts time in syscalls or network reads
  • Reads the largest block-profile delay as contention without checking the stack
  • Believes an uncontended Lock and Unlock pair produces a mutex sample
open as a page

In a Go service, throughput plateaus while CPU sits near 30% — how do block and mutex profiles find the lock?

level: seniorimportance: must knowfreq 45%

basics

~20 s

Low CPU with flat throughput means goroutines are parked, not computing. Enable both profilers coarsely on one instance and diff two captures: the mutex profile's top delay names the critical section, and the block profile separates lock waits from channel waits.

open as a page

Why is Go's block profile empty until you call runtime.SetBlockProfileRate?

level: juniorimportance: should knowfreq 30%

basics

~20 s

Go's block and mutex profilers are switched off by default because recording costs throughput. Until you call runtime.SetBlockProfileRate for blocking events, or runtime.SetMutexProfileFraction for lock contention, the runtime samples nothing and the profile comes back with no samples.

open as a page

Would you leave Go's block and mutex profilers enabled permanently in production, and at what rates?

level: principalimportance: should knowfreq 28%

basics

~20 s

Usually yes for the mutex profiler at a coarse fraction, and off or very coarse for the block profiler. Cost scales with how often goroutines block, so measure it per service and make both rates adjustable at runtime.

open as a page

What do the arguments to runtime.SetBlockProfileRate and SetMutexProfileFraction mean?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

They are not the same unit. runtime.SetBlockProfileRate takes nanoseconds: it aims for one sampled event per that many nanoseconds spent blocked, and 1 samples everything. runtime.SetMutexProfileFraction takes a fraction: on average one contention event in N is reported.

open as a page