skip to content

Why can a goroutine that takes sync.RWMutex.RLock twice deadlock in Go?

level: middleimportance: should knowfreq 46%

answer

  1. it takes a third goroutine to bite
  2. writers must not starve
  3. a waiting writer shuts the door on new readers
  4. the inner acquisition queues behind that writer

basics

~20 s

Once a goroutine blocks in Lock, sync.RWMutex admits no new readers, so the waiting writer cannot be starved. A nested RLock then queues behind that writer while the writer waits for the outer read lock to be released, and all three block forever.

solid answer

~50 s

`sync.RWMutex` deliberately gives a waiting writer priority over arriving readers: once a goroutine is blocked in `Lock`, further `RLock` calls block too, otherwise a steady stream of readers could keep a writer out indefinitely. That rule is what makes recursive read locking unsafe. Picture three goroutines: A takes `RLock`; B calls `Lock` and blocks because A is inside; A now calls `RLock` again — from a helper method that also locks — and blocks behind B, while B is waiting for A. Nothing can move. It is timing-dependent, so it survives testing and appears under production write traffic. The fix is structural: never call a lock-taking method from inside a locked section. Put the body in an unexported helper that assumes the caller holds the lock, and let the exported methods take the lock exactly once.

code

go · 13 lines
go
func (r *Router) Count() int {
	r.mu.RLock()
	defer r.mu.RUnlock()
	return len(r.routes)
}

func (r *Router) Report() string {
	r.mu.RLock()
	defer r.mu.RUnlock()
	// Count takes RLock again. A writer blocked in Lock between
	// the two acquisitions shuts out new readers, and all three park.
	return fmt.Sprintf("%d routes", r.Count())
}

go deeper

for a junior

Know that a Go read lock is not reentrant: taking it again in the same goroutine is not automatically safe. Keep each acquisition and its matching release inside one small function.

for a middle

Be ready to walk the three-goroutine timeline aloud: outer read lock, a writer blocked behind it, the nested acquisition queued behind the writer. Explain why the writer-first rule exists in the first place.

for a senior

Show how you would find it in a live process from a goroutine dump, and how you refactor to unexported helpers that assume the lock is held rather than sprinkling more locking.

for a principal

The durable fix is a convention, not a patch. Decide where locking is allowed to live in a package and make it an explicit review rule, so this class of bug cannot be reintroduced by the next contributor.

## The rule that causes it A readers-writer lock has an obvious failure mode: if arriving readers are always admitted whenever other readers are inside, a busy read path can keep the reader count above zero forever and a writer never gets in. Go's `sync.RWMutex` closes that hole with a documented rule — **while a goroutine is blocked in `Lock`, new `RLock` calls block as well**. Readers already inside finish and leave; readers arriving after the writer queues wait behind it. The writer is therefore guaranteed to acquire the lock in bounded time. The documentation states the consequence explicitly: this prohibits recursive read locking. Taking `RLock` while you already hold `RLock` in the same goroutine is not safe, even though it looks harmless. ## The three-goroutine timeline 1. Goroutine A calls `RLock` and enters the shared section. The reader count is 1. 2. Goroutine B calls `Lock`. A reader is inside, so B parks — and from this moment the lock refuses new readers. 3. Goroutine A, still inside its shared section, calls a helper that also takes `RLock`. That second acquisition is a *new* reader, so it parks behind B. 4. B is waiting for the reader count to reach zero. It never will, because the only reader is A and A is parked waiting for B. The lock is not reentrant and does not try to be. It has no notion of which goroutine holds a read lock — it only counts — so it cannot recognise that the second `RLock` came from a goroutine already inside. ## How it appears in real code Almost never as two `RLock` lines next to each other. It appears as two exported methods on the same type, each correctly taking the lock, and then one calling the other: - `Report()` takes `RLock`, then calls `Count()`, which takes `RLock` too. - A `String()` method takes `RLock` and is invoked from inside another locked method by a formatting call. - A callback the caller supplies is invoked while the lock is held, and the callback calls back into the same type. The last one is the nastiest, because the deadlock is created by a caller you have never seen. ## Why it hides Step 2 must land between step 1 and step 3. With a low write rate that window is small, so unit tests and staging pass and the failure surfaces when write traffic increases or a machine gets busier. The race detector will not help: this is a deadlock, not a data race, and `-race` reports unsynchronised access, not blocked goroutines. The runtime's built-in `all goroutines are asleep - deadlock!` check does not fire either, because that only triggers when *every* goroutine is blocked, and a server always has timers, the network poller and other requests keeping things runnable. What does find it is a goroutine dump — send the process `SIGQUIT`, or read `/debug/pprof/goroutine?debug=2` — where you will see one goroutine parked inside `sync.(*RWMutex).Lock` and a pile parked inside `sync.(*RWMutex).RLock`, with the stack showing your own outer method holding on. ## The fix Adopt the convention that **locking happens at exactly one layer**. Exported methods acquire the lock; unexported helpers assume it is already held and never touch the mutex. A naming convention makes it reviewable — a `Locked` suffix, or a doc comment saying the caller must hold `mu`. Then the outer method takes the lock once and calls helpers freely. Additional discipline that follows from the same rule: - Never invoke a caller-supplied function or an interface method while holding the lock, unless you have documented that it must not re-enter. - Keep the locked region small enough to read in one screen, so a nested lock is visible. ## The corollary: there is no upgrade The same rule explains why `sync.RWMutex` offers no way to turn a held read lock into a write lock: the upgrade would have to wait for other readers who may themselves be waiting to upgrade. Go simply does not provide the operation. If you read under `RLock`, decide a write is needed, `RUnlock` and then `Lock`, the state can change in the gap, so you must **re-check the condition after acquiring the write lock**. If the check and the write must be atomic, take `Lock` from the start and accept that readers are excluded for the whole operation.

  • How do you restructure a method that takes RLock and needs another method that also takes RLock?
    Move the shared body into an unexported helper that assumes the lock is already held — a `Locked` suffix or a doc comment makes the requirement reviewable. Exported methods acquire the lock once and call helpers. The rule to enforce in review is simple: no lock-taking call inside a locked section, including caller-supplied callbacks.
  • Does sync.RWMutex offer any way to turn a read lock you hold into a write lock?
    No. There is no upgrade operation at all. You call `RUnlock`, then `Lock`, and then re-check whatever you decided under the read lock, because another goroutine may have changed the state in the gap. If the decision and the write must be atomic, take `Lock` from the start.
  • Why does this deadlock usually survive testing?
    It needs a writer to arrive in `Lock` between the two read acquisitions. With a low write rate that window is tiny, so tests and staging pass and production fails. `-race` will not flag it either, since it is a deadlock rather than unsynchronised access, and the runtime's all-goroutines-asleep detector stays quiet because a server always has other runnable goroutines.

A queue where the librarian waiting to swap the notice is let in ahead of newcomers. If you are already reading and step out to join the back of the queue for a second look, you are now waiting for a librarian who is waiting for you.

saying these in an interview costs you the question

  • Says Go read locks are reentrant
  • Thinks a reader only ever waits for an active writer
  • Claims RUnlock then Lock keeps the check atomic
  • Expects the runtime deadlock detector to catch it
  • Calls a lock-taking method from inside a locked section