skip to content

Two goroutines lock two account mutexes in opposite order and wedge — what do you capture from the live process?

level: seniorimportance: nice to knowfreq 42%

answer

  1. the process is alive, so evidence exists
  2. stacks before restart, never after
  3. full detail, not aggregated counts
  4. the wait reason and its age
  5. mirrored arguments in the same frame

basics

~10 s

Capture full goroutine stacks before restarting: /debug/pprof/goroutine?debug=2 if net/http/pprof is registered, or GOTRACEBACK=all plus SIGQUIT. Two goroutines parked in sync.(*Mutex).Lock inside mirrored transfer calls are the cycle.

solid answer

~50 s

A lock-order cycle leaves the process alive — the listener still accepts, health checks still pass — so nothing crashes and you get exactly one chance at evidence before someone restarts it. Take a full goroutine dump first: `/debug/pprof/goroutine?debug=2` if `net/http/pprof` is registered, otherwise `runtime/pprof.Lookup("goroutine").WriteTo(w, 2)` from a debug endpoint; failing both, `GOTRACEBACK=all` and a `SIGQUIT`, which prints every stack but kills the process, so do it last. Read the dump for goroutines parked in `sync.(*Mutex).Lock`. Each header carries the wait reason and how long it has waited — `[semacquire, 9 minutes]` — and the frames above it show which transfer call the goroutine is inside, and therefore which account it already holds while waiting for the other. Two of those with mirrored arguments is the proof. The mutex profile will not help: contention is recorded when a contended lock is released, and these never are.

code

go · 14 lines
go
type Account struct {
	mu      sync.Mutex
	id      int64
	balance int64
}

func Transfer(from, to *Account, amount int64) {
	from.mu.Lock()
	defer from.mu.Unlock()
	to.mu.Lock() // a mirrored call is already holding this one
	defer to.mu.Unlock()
	from.balance -= amount
	to.balance += amount
}

go deeper

for a junior

Know that a Go process can be wedged and still look healthy, and that the first move is a goroutine dump, not a restart. Recognising semacquire in a stack as waiting on a lock is enough at this level.

for a middle

Be able to explain how to get full goroutine stacks — the pprof goroutine endpoint at debug=2, or the runtime/pprof lookup — and why debug=1 counts are not enough to identify which locks are involved.

for a senior

Demonstrate the whole capture sequence under time pressure, including the second dump that separates stuck from slow, and read the mirrored frames as proof. Say clearly why the race detector, the mutex profile and a CPU profile all come back empty here.

for a principal

Own the standing capability rather than the incident: whether every service exposes a debug endpoint, whether readiness checks exercise the paths that can wedge, and whether the fix ships with a concurrent test so the class of bug fails in CI instead of on call.

## Why this is a capture problem, not a crash Two goroutines each holding one account's mutex and waiting for the other's are permanently parked, but the process is otherwise healthy. The accept loop runs, `/healthz` returns 200, unrelated endpoints work, and the only symptom is that transfers stop completing and the in-flight count climbs. Nothing writes a stack trace on its own. The engineer looking at it has one decision to make first — capture, then restart — because a restart destroys the only evidence that distinguishes a lock cycle from a slow dependency, a stuck downstream call, or a leak. ## What to capture, in order **1. The goroutine profile at full detail.** If the binary registers `net/http/pprof` on an internal listener, `GET /debug/pprof/goroutine?debug=2` returns every goroutine's full stack as text, without disturbing the process. `debug=1` gives aggregated counts, which tells you *how many* are stuck but not *where*; for a cycle you want `debug=2`. Without the HTTP handler, the same text comes from `runtime/pprof.Lookup("goroutine").WriteTo(w, 2)` on any debug path the binary already exposes. **2. A second dump, thirty seconds later.** Two dumps distinguish stuck from slow. A goroutine present in both, in the same frame, with a growing wait age, is stuck. This step is what stops a wrong diagnosis of "the database is slow". **3. Anything cheap about the shape of the pile-up** — the in-flight request gauge, the goroutine count over time, the queue depth. A lock cycle produces a linear climb in goroutines blocked in the same frame, which is a distinctive curve. **4. Last, `GOTRACEBACK=all` plus `SIGQUIT`.** This prints all goroutine stacks and then kills the process. It is the fallback when the binary exposes no debug endpoint, and it is deliberately last because it ends the incident on its own terms. ## Reading the dump A goroutine blocked acquiring a `sync.Mutex` appears with the wait reason `semacquire` in its header, and once it has been blocked for minutes the runtime prints the age too: ``` goroutine 18 [semacquire, 9 minutes]: sync.(*Mutex).Lock(...) main.Transfer(0xc000112000, 0xc000112040, 0x1) ``` The header tells you it is waiting on a lock and roughly since when. The frames tell you *which* lock, indirectly but conclusively: the goroutine is inside `Transfer`, which by inspection takes the source account's lock first and the destination's second, so a goroutine parked at the second `Lock` already owns the first. The pointer arguments printed for the frame identify the two accounts, and the mirrored pair — one goroutine with `(A, B)` and another with `(B, A)`, both parked in the same frame — is the cycle written out. Two goroutines with the same wait age, stuck at the same line, in mirrored calls, is as close to a proof as a dump gets. ## What will not tell you - **The mutex profile.** It records contention events when a contended lock is *released*, attributing the delay to the holder. A lock that is never released contributes nothing, so the profile of a deadlocked process is empty exactly where you want data. It also requires `runtime.SetMutexProfileFraction` to have been enabled beforehand. - **The race detector.** It reports unsynchronised concurrent accesses that actually happened. A lock cycle is the opposite problem — too much synchronisation, correctly performed — and a `-race` build will sit there wedged, saying nothing. - **A CPU profile.** Parked goroutines burn no CPU. The profile will show an idle process, which is true and useless. - **`go vet`.** There is no lock-order analysis in the standard toolchain. Attaching a debugger to the live process is the other way in, and it can show the mutex state words directly and which goroutine owns each — worth it when the dump's frames are ambiguous, though it usually confirms what the dump already said. ## Turning the capture into a fix The dump gives you the two call sites; the fix is a design change at those sites. Locking both accounts in an order derived from a stable key such as the account id makes the mirrored pair impossible. Alternatively, remove the second lock from the picture entirely: one mutex over the ledger, or a single owner goroutine that applies every transfer in sequence, trades some throughput for a structure in which no goroutine ever holds two locks. `sync.Mutex` offers no lock timeout, and while `TryLock` exists, backing off and retrying is easy to get subtly wrong and does not by itself bound the retry storm. ## Making it reproducible Once you know the two call sites, a test that runs both directions concurrently in a tight loop reproduces it in seconds on a multi-core machine. It will hang, and `go test` kills a timed-out binary with a dump of every goroutine — which is the same evidence, in CI, before the change ships. That test is the artefact worth leaving behind, more than the incident writeup.

  • Why does the mutex profile show nothing for this?
    Because it records a contention event when a contended lock is released, blaming the holder for the delay. Locks caught in a cycle are never released, so they never produce an event. It also has to be switched on in advance with `runtime.SetMutexProfileFraction`. It is a tool for finding slow critical sections, not stuck ones.
  • The instance kept passing health checks throughout. What should the health check have done?
    Distinguish liveness from readiness and make readiness exercise the path that can wedge — a cheap transaction against the ledger, or a threshold on in-flight transfers and goroutines blocked in the same frame. A check that only proves the accept loop is alive will keep a wedged instance in rotation for as long as it takes a human to notice.
  • What do you change so it cannot recur?
    Acquire the two account locks in an order fixed by a stable key such as the account id, so mirrored callers converge on the same sequence; or restructure so no goroutine ever holds two account locks — one ledger lock, or one owner goroutine applying transfers in sequence. Then add a test that runs both directions concurrently, so a regression hangs in CI instead of in production.

saying these in an interview costs you the question

  • Restarts the process before capturing any goroutine stacks
  • Expects the Go runtime to abort a two-mutex lock cycle
  • Runs a -race build expecting the detector to report the deadlock
  • Looks in the mutex profile for a lock that is never released
  • Believes sync.Mutex.Lock can be given a timeout
  • Takes a CPU profile of a process whose goroutines are all parked