skip to content

Your service swapped sync.Mutex for sync.RWMutex on a hot read path and throughput did not improve. How do you find out why?

level: seniorimportance: should knowfreq 40%

answer

  1. the swap is not free
  2. how long is the read section really?
  3. every reader touches one shared counter
  4. profile which side is waiting before changing anything

basics

~20 s

Measure first: a block profile shows whether goroutines wait on the lock at all and on which side. Usually the read section is too short for the extra bookkeeping to pay, the write rate keeps shutting readers out, or the lock was never the bottleneck.

solid answer

~50 s

I would stop guessing and get evidence. Turn on the block profile with `runtime.SetBlockProfileRate` (or run the benchmark with `-blockprofile`) and see whether goroutines are actually parked in `sync.(*RWMutex).RLock`, in `Lock`, or somewhere else entirely — often the lock was never the bottleneck and a CPU profile shows the time is in allocation or in a downstream call. If readers are waiting, look at the write rate: a blocked writer stops admitting new readers, so frequent or slow writes serialise the read path anyway. If nobody is waiting, the swap simply had nothing to win: `RWMutex` costs an atomic add and subtract on a counter every reader on every core touches, and for a critical section that is one map lookup that is roughly what an uncontended `sync.Mutex` costs. The usual fix is to shorten the exclusive section — build the replacement outside the lock and assign it inside — not to change primitive again.

code

go · 15 lines
go
type Rule struct {
	Path    string
	Handler string
}

func (r *Router) Reload(rules []Rule) {
	next := make(map[string]string, len(rules)) // built with no lock held
	for _, rule := range rules {
		next[rule.Path] = rule.Handler
	}

	r.mu.Lock() // exclusive section is now one assignment
	r.routes = next
	r.mu.Unlock()
}

go deeper

for a junior

Know that RWMutex is not automatically faster than Mutex and that readers share one lock's bookkeeping. Be able to say you would measure before and after rather than assume the swap helped.

for a middle

Explain the cost model out loud: the shared reader counter is atomic work on a cache line every reader touches, so a very short read section has nothing to gain from reader concurrency.

for a senior

Drive the investigation. Enable the block profile, separate time waiting in the read path from the write path, and land the fix that shortens the exclusive section rather than swapping primitives a second time.

for a principal

Frame it as one change against a measurement budget, with a stated threshold for reverting. Decide whether the team's effort belongs on the lock at all or on the work happening inside the critical section.

## Why the swap can buy nothing `sync.RWMutex` gives you reader concurrency and charges you for coordination. Internally it is a mutex plus a reader counter, a reader-wait counter and semaphores. An uncontended `sync.Mutex` acquisition is about one atomic compare-and-swap; `RLock` is at minimum an atomic add, and `RUnlock` an atomic subtract, on **one counter that every reader on every core touches**. That word is a cache line ping-ponging between cores. So the shared mode wins only when the concurrency it unlocks is worth more than the coordination it adds. Three conditions have to hold together: 1. **Reads dominate writes.** Not by a little — by orders of magnitude, because a waiting writer blocks arriving readers too. 2. **The read section is long enough to amortise the bookkeeping.** A full scan, a decode, several dependent lookups: yes. A single map read: usually not. 3. **Enough goroutines actually read concurrently.** On a machine running one or two readers at a time there is no serialisation to remove. If any of those is false, the swap is a no-op or a small regression, and the honest outcome is to revert to `sync.Mutex`. ## The measurement sequence **Start with a CPU profile of the real workload.** Before assuming the lock matters at all, confirm where the time goes. It is common for a hot path to spend most of its time allocating, formatting a log line or waiting on a downstream call, in which case no lock change will move the number. **Then the block profile.** `runtime.SetBlockProfileRate(rate)` makes the runtime sample how long goroutines spend blocked on synchronisation, and the resulting profile attributes that time to stacks. This is the one diagnostic that separates the two sides of a readers-writer lock: waiting inside `sync.(*RWMutex).RLock` means readers are being excluded — by a writer holding the lock, or by a writer queued behind earlier readers. Waiting inside `sync.(*RWMutex).Lock` means writers are queueing behind readers. The split tells you which section to shorten. The profile costs throughput while it is on, so enable it deliberately, on one instance, for a bounded window. **A mutex profile** (`runtime.SetMutexProfileFraction`) complements it by attributing contention to the holders rather than the waiters. **Reproduce it in a benchmark.** `go test -bench` with `b.RunParallel`, `-cpu=1,4,16` to vary parallelism and `-benchmem`, keeping the guarded work identical between the two variants. The crossover point between the two locks moves with the length of the read section and the core count, so a single-core benchmark will tell you nothing useful about a 32-core box. ## What the answers usually are - **The read section is trivial.** Two atomics on a shared counter cost about what the mutex cost. Revert. - **Writes are more frequent than assumed.** Every queued writer closes the door on new readers until it has run, so a write every few milliseconds can serialise a read path almost as thoroughly as a plain mutex. Count the writes before blaming the lock. - **The write section is long.** A rebuild done while holding `Lock` stalls every reader for its whole duration. This is the highest-value fix and it is usually easy: construct the new value with no lock held, then take `Lock` only to assign it. The exclusive section drops from milliseconds to a pointer store. - **Work that does not belong is inside the critical section.** I/O, a log call, a JSON decode, a large allocation. Move it out; the lock should cover the shared memory access and nothing else. - **The lock was never the bottleneck.** The CPU profile said so at step one. ## What to change, in order 1. Shorten the exclusive section — build outside, assign inside. 2. Remove I/O, logging and allocation from both sections. 3. Copy the small values you need out under the lock instead of holding it while the caller works with them. 4. Reduce the write frequency: batch or coalesce updates that arrive in bursts. 5. Only then reconsider the primitive itself, with a benchmark that mirrors the production read/write ratio and core count. ## The discipline behind all of it Lock choice is a measured decision, not a style preference. Default to `sync.Mutex`, switch when a profile shows readers genuinely serialising on a section long enough to matter, and be willing to switch back when the numbers do not move. A change that cannot be shown to help is a change that added complexity for free.

  • What does a block profile dominated by sync.(*RWMutex).Lock rather than RLock tell you?
    That writers are the ones queueing, behind readers. Either the write rate is higher than the design assumed, or the exclusive section itself is long. Look at how much work happens while `Lock` is held: building a replacement outside the lock and assigning it inside usually removes almost all of that wait, without touching the read path.
  • How would you benchmark the two locks fairly?
    `go test -bench` with `b.RunParallel` so the read path is driven from many goroutines, `-cpu=1,4,16` to vary parallelism, and `-benchmem` to catch allocation differences. Keep the guarded work byte-for-byte identical between variants and mirror the production read-to-write ratio, because the crossover moves with both the section length and the core count.
  • The read section turns out to be a single map lookup. What is your recommendation?
    Keep `sync.Mutex`. At that size the atomic add and subtract on a counter shared by every reader cost about what an uncontended mutex costs, so you have paid extra complexity and an extra failure mode — recursive read locking — for no measurable gain. Revert, and spend the change budget on the work inside the section.

saying these in an interview costs you the question

  • Treats RWMutex as a free upgrade over Mutex
  • Swaps primitives without measuring before or after
  • Holds the write lock while rebuilding the whole table
  • Never checks how often writers actually run
  • Leaves I/O or logging inside the critical section