skip to content

In a Go goroutine dump, how do you find which goroutine holds the sync.Mutex the others are waiting on?

level: seniorimportance: should knowfreq 35%

answer

  1. Go's mutex stores no owner
  2. the dump prints frame arguments in hex
  3. the receiver word is the lock's address
  4. the holder already returned from Lock
  5. group waiters by that address, then follow

basics

~20 s

Go records no mutex owner, so you infer it. Group blocked goroutines by the hex receiver printed for their sync.(*Mutex).Lock frame; the holder is the goroutine already past a Lock on that same address, parked on something else.

solid answer

~50 s

A `sync.Mutex` is two words with no owner field, so nothing in the dump ever says "held by goroutine 41" the way other runtimes' thread dumps do. You reconstruct it from two things Go does print: each goroutine's wait reason and age in its header line, and the first few argument words of each frame in hex. For a frame like `sync.(*Mutex).lockSlow(0xc0000b4010)`, that hex value is the address of the mutex itself, so grouping every blocked goroutine by it tells you exactly which lock instance is contended and how many are queued. The holder is then the goroutine whose stack shows it already returned from `Lock` on that address — it is inside the critical section, parked on something else entirely, typically a channel receive, a network read or a second mutex. Chase that second wait and you have the cycle.

code

text · 13 lines
text
goroutine 41 [chan receive, 61 minutes]:
main.(*reconciler).apply(0xc0000b4000)
	/src/reconcile.go:88 +0x9c
main.(*reconciler).sync(0xc0000b4000)
	/src/reconcile.go:61 +0x145

goroutine 57 [sync.Mutex.Lock, 61 minutes]:
sync.(*Mutex).lockSlow(0xc0000b4010)
	/usr/local/go/src/sync/mutex.go:171 +0x15d
sync.(*Mutex).Lock(...)
	/usr/local/go/src/sync/mutex.go:90
main.(*reconciler).status(0xc0000b4000)
	/src/reconcile.go:120 +0x45

go deeper

for a junior

Know that Go's dump shows each goroutine's wait reason and, past a minute, how long it has waited, and that a goroutine blocked in sync.(*Mutex).Lock is queued behind someone else's critical section.

for a middle

Explain that a sync.Mutex has no owner field, so the holder must be inferred, and that the hex value printed with a frame is the receiver — the mutex's own address for a Lock frame.

for a senior

Walk the whole inference: group waiters by lock address, find the goroutine already inside the critical section, follow its wait, and confirm with a second dump that nothing advanced.

for a principal

Decide when inference stops paying and instrumentation starts: a recurring cycle justifies a temporary owner-recording wrapper behind a build flag, and the team should agree in advance what evidence a hang must produce.

## Why the dump refuses to help you directly A `sync.Mutex` is deliberately tiny: a state word and a semaphore word. It stores no owner, no recursion count and no acquisition site, which is part of why it is fast and why it is not reentrant. The consequence for diagnosis is blunt: a Go goroutine dump contains no line naming the holder of a lock. Engineers arriving from runtimes whose thread dumps print an explicit owner look for that line, do not find it, and conclude the dump is useless. It is not — the information is there, just as an inference. ## What the dump does give you Two things carry the signal. **The header line.** Each goroutine is introduced with its number, its wait reason and, once it has been blocked for more than a minute, its wait age: something of the form `goroutine 57 [sync.Mutex.Lock, 61 minutes]`. Current releases print a specific wait reason for mutex acquisition; older ones printed the more generic semaphore wait. The age is the part people underuse — it is a free timeline. **The frame arguments.** After each function name the runtime prints the first few words of that frame's arguments in hex. For a method the first word is the receiver pointer. So `sync.(*Mutex).lockSlow(0xc0000b4010)` tells you the address of the mutex being acquired, and `main.(*reconciler).sync(0xc0000b4000)` tells you which instance of your own type is running. Values the runtime is unsure about — registers that may have been reused at that program counter — are printed with a trailing `?`; treat those as hints rather than facts. ## The procedure 1. **Collect every goroutine whose stack contains a mutex acquisition** — frames under `sync.(*Mutex).Lock`, or `sync.(*RWMutex).Lock` and `RLock` if the lock is a read-write one. 2. **Read the receiver hex for that frame and group by it.** All goroutines sharing an address are queued on the same mutex instance. If you see two clusters, you have two contended locks and probably a cycle rather than simple contention. 3. **Find the holder.** It is the goroutine that has *already returned* from `Lock` on that address: its stack shows a function you know takes that mutex, with frames *above* it that are doing something else. It is parked on some other wait — a channel receive, a network read, a second `Lock`. If the mutex is a field of a struct, its address is the struct's address plus the field offset, so `main.(*reconciler).sync(0xc0000b4000)` sitting on a mutex at `0xc0000b4010` is very likely the same object; when the mutex is the first field, the two addresses are identical. 4. **Follow the second wait.** If the holder is blocked acquiring another mutex whose address is the one a third goroutine holds, and that one is blocked on the first, you have a cycle and you can name every participant. If instead the holder is blocked on a channel receive or a remote call, you have a chain, not a cycle — and the fix is usually a timeout or a context at the far end rather than lock reordering. 5. **Use the ages.** Everyone in a cycle shares roughly the same age. A cluster of waiters that are all younger than one goroutine points at that older goroutine as the head of the chain: it blocked first, and everyone else piled up behind it. ## Confirming that it is stuck and not slow Take a second dump a minute later. Same goroutine numbers, same frames, ages advanced by exactly the elapsed time means nothing moved. If the goroutine numbers change and the frames differ, you are looking at slow progress or contention, not a deadlock, and the investigation goes elsewhere. ## Traps worth knowing - **Inlining merges frames.** A small function that took the lock may not appear as its own frame; the acquisition looks like it happened in its caller. - **`sync.RWMutex` inverts the shape.** One pending writer blocks all later readers, so the visible symptom is a long queue of `RLock` waiters while the actual culprit is a single writer waiting behind an existing reader that will not finish. - **A deferred `Unlock` proves nothing about liveness.** `defer mu.Unlock()` guarantees release when the function returns; if the function never returns, the deferred call never runs and the lock is held forever. - **The dump is a snapshot with skew.** It is taken while the world is not perfectly still, so a single frame that looks odd is not necessarily meaningful; the pattern across goroutines is. ## When inference is not enough If a lock cycle keeps recurring and the stacks are ambiguous, stop guessing and make ownership explicit for one release: wrap the mutex in a small debug type that records the acquiring goroutine's identity and acquisition site behind a build flag, then remove it once the cycle is understood. That is a heavier tool than reading a dump, but it turns an inference into a fact — and, unlike the dump, it survives inlining.

  • Two dumps taken a minute apart look identical. What does that add?
    It rules out slow progress. Same goroutine numbers, same frames and wait ages advanced by exactly the elapsed time mean nothing moved at all. A merely slow system shows different goroutine numbers and different frames between snapshots, which sends the investigation towards contention or a slow dependency instead of a lock cycle.
  • Why does a hex argument in a Go stack sometimes end with a question mark?
    The runtime prints `?` when it cannot be sure the printed word is still the argument's value at that program counter — with register-based calls a register may have been reused. Treat such a value as a hint: it is often right, but do not build a conclusion about which mutex is contended on a single question-marked address.
  • How does the picture change when the contended lock is a sync.RWMutex?
    Waiters appear under `RLock` as well as `Lock`, and a single pending writer blocks every reader that arrives after it. So a dump showing a long queue of readers is usually the symptom, not the cause — look for the one writer waiting behind an existing reader whose critical section never ends.
  • The holder is parked on a channel receive rather than another lock. Is that still a deadlock?
    It is a chain rather than a cycle, and it hangs just as hard. Everything queued behind the mutex waits for a value that may never arrive. The fix usually belongs at the far end — a context deadline or a timeout on whatever produces that value — rather than in the locking order.

saying these in an interview costs you the question

  • Expects the dump to name the lock owner
  • Ignores the hex receiver and cannot tell two mutexes apart
  • Assumes the longest-waiting goroutine holds the lock
  • Reads one dump and cannot distinguish hung from slow
  • Believes a deferred Unlock guarantees the lock is released