Your team swapped a map behind a sync.RWMutex for a sync.Map and throughput dropped. Why?
answer
- specialised container, not a free upgrade
- untyped API costs an allocation per entry
- the cheap read path must amortise
- an uncontended lock is already nearly free
- measure with the real read write mix
basics
~20 ssync.Map is not a general-purpose faster map. Keys and values are any, so every store boxes and every load asserts, and its cheap read path only pays off for the two access patterns it targets. Otherwise a lock-guarded typed map usually wins.
solid answer
~50 s`sync.Map` buys lock-free reads for entries that have settled, and pays for them twice. First, the API is untyped: keys and values are `any`, so storing a non-pointer value boxes it into an interface — an allocation and an indirection per entry — where a `map[string]*connMeta` behind a lock stores the pointer directly and stays compact in cache. Second, the machinery that makes reads cheap only amortises when entries are written once and read many times, or when goroutines work on disjoint key sets; a workload that keeps inserting or rewriting keys never settles, so you pay the bookkeeping without the payoff. Meanwhile an uncontended mutex is a couple of atomic operations, so there was barely any lock cost to remove. Settle it by measuring: `go test -bench` with `b.RunParallel` and `-benchmem`, under your real read/write ratio and key distribution, comparing ns/op, B/op and allocs/op.
code
go · 9 linesfunc BenchmarkRegistryParallel(b *testing.B) {
b.ReportAllocs()
b.RunParallel(func(pb *testing.PB) {
for pb.Next() {
// perform one operation of the production read/write mix
// against the implementation under test
}
})
}go deeper
Know the headline: sync.Map is a specialised type, not a faster map for everything, and a plain map guarded by a lock is the normal choice for shared state.
Explain the two concrete costs — boxing keys and values into any on every operation, and a cheap read path that only amortises for settled entries — plus the fact that an uncontended lock is already very cheap.
Demonstrate the measurement: a parallel benchmark with -benchmem under the real read/write ratio and key distribution, and a profile confirming the lock was contended before you changed anything. Be willing to revert.
Own the decision process rather than the answer: require that a container swap in shared infrastructure be justified by a profile and a benchmark, and keep the boring typed implementation as the default the team falls back to.
## The assumption behind the regression The swap usually comes from a reasonable-sounding chain: the map is shared, the shared map has a lock, locks are contention, `sync.Map` is the lock-free one, therefore `sync.Map` is faster. Every step is plausible and the conclusion is wrong, because `sync.Map` is a *specialised* container whose advantage is conditional, not a drop-in optimisation. ## Cost one: the API is untyped `sync.Map`'s methods take and return `any`. That has concrete runtime consequences that a typed map does not pay: - **Boxing.** Storing a value that is not already pointer-shaped means putting it in an interface, which generally allocates. A `map[int]int` stores eight bytes inline; `sync.Map` with the same data allocates behind each key and each value. - **Indirection.** Every load returns an interface that must be type-asserted, and the real value lives one pointer hop away, which is worse for cache locality than a compact typed map's contiguous storage. - **Garbage.** Those per-operation allocations become garbage-collector work, which shows up as CPU time far away from the map code and is easy to misattribute. This cost is paid on *every* operation, whatever the concurrency, and it is why a `sync.Map` version of a benchmark often shows more allocations per operation than the locked map even when the two do the same logical work. ## Cost two: the fast path is conditional The reason `sync.Map` can serve reads without taking a lock is that it maintains extra internal structure so that stable entries can be read cheaply, at the cost of bookkeeping when entries are added or changed. The published guidance names two workloads where that trade pays: 1. **Write once, read many** — a cache that only grows, where entries settle and are then read repeatedly. 2. **Disjoint key sets** — several goroutines each working on their own keys, so they rarely touch the same entries. A registry of long-lived client connections, written once when a client arrives and read on every subsequent request, fits the first shape well. A per-request scratch map, a hot counter table where every goroutine updates the same handful of keys, or a workload with a steady stream of brand-new keys fits neither: nothing settles, so you pay the bookkeeping continuously while the reads never get to be cheap. ## Cost three: the lock you removed may have been free An uncontended mutex acquisition is a couple of atomic instructions — tens of nanoseconds, no system call, no goroutine parking. If your critical section is a single map lookup and only a handful of goroutines ever collide, there was almost no contention to eliminate. You removed something nearly free and added boxing to every operation. Contention only becomes the dominant cost when many goroutines hit the same lock at once, or when the critical section is long — and the first fix for a long critical section is to make it shorter, not to change container. ## Measuring it properly ```go func BenchmarkRegistryParallel(b *testing.B) { b.ReportAllocs() b.RunParallel(func(pb *testing.PB) { for pb.Next() { // one op of your real read/write mix } }) } ``` Run it with `go test -bench=. -benchmem` and read three numbers for each implementation: ns/op, B/op and allocs/op. The details that decide whether the benchmark means anything: - **Model the real mix.** A 99%-read benchmark and a 50%-write benchmark will rank the two implementations differently. Use the ratio your service actually has. - **Model the key distribution.** Ever-growing distinct keys, a fixed hot set and a long tail behave completely differently. A uniform random key over a huge space flatters `sync.Map`'s disjoint-keys case in a way production may not. - **Use realistic parallelism.** `b.RunParallel` scales with `GOMAXPROCS`; a result at one core says nothing about a container with eight. - **Confirm the lock was the problem first.** Before changing anything, enable the mutex profile with `runtime.SetMutexProfileFraction` and look at a CPU profile. If lock waiting is not a visible share of the profile, no container swap will show up in latency, and you have spent a change budget on noise. ## The judgment to state out loud The default is a plain typed map with a lock, wrapped in a struct whose methods own the locking — it keeps types, keeps `len`, allocates less, and every reader understands it. Move to `sync.Map` when a measurement shows lock contention *and* the access pattern matches one of the two it targets. "We changed it because it sounded faster" is the answer that costs the interview; "we benchmarked both under our real mix and kept the one that won" is the one that passes. And the same discipline applies to reverting: if the profile says the swap did not help, put it back.
- What does -benchmem show that ns/op alone hides?B/op and allocs/op — the allocation cost of boxing keys and values into `any`. A sync.Map version frequently allocates more per operation than a typed map behind a lock, and that garbage turns into collector CPU time later, in a place that looks unrelated to the map.
- How would you confirm the lock was actually the bottleneck before swapping containers?Enable the mutex profile with `runtime.SetMutexProfileFraction` and take a CPU profile under production-like load. If waiting on that lock is not a visible share, the container is not the problem and no swap will move latency. That check costs an afternoon and often saves the whole change.
- Which workloads genuinely fit sync.Map?Entries written once and then read many times — a registry or a cache that only grows — and goroutines operating on disjoint sets of keys. A hot set of shared keys rewritten constantly, or a steady stream of brand-new keys, fits neither: the structure never settles, so the bookkeeping is paid without the cheap-read payoff.
- If the benchmark is inconclusive, which implementation do you keep?The typed map behind a lock. It keeps compile-time types, keeps len, allocates less and reads more obviously, so it is the cheaper thing to maintain. A specialised container should have to win a measurement to earn its place, not merely fail to lose one.
saying these in an interview costs you the question
- Calls sync.Map a drop-in faster concurrent map
- Ignores the allocation cost of boxing into any
- Swaps implementations without benchmarking either
- Assumes lock-free always beats an uncontended mutex
- Benchmarks with a read mix nothing like production
- Never checks whether the lock was contended at all