skip to content

A Go session store stops scaling as you add cores, with wait time on one sync.Mutex — what do you change?

level: seniorimportance: should knowfreq 44%

answer

  1. more cores, same ceiling
  2. the queue grows, the work does not
  3. cheapest change first
  4. shrink, then partition, then publish
  5. re-measure between every step

basics

~20 s

One lock serialises every request, so extra cores only lengthen the queue behind it. Shrink what runs inside the critical section first, then partition the lock across hashed shards, then publish read-mostly state as a snapshot swapped through an atomic pointer.

solid answer

~50 s

Throughput flat or falling as cores are added, with wait time on one `sync.Mutex`, says the store is serialised on that lock: every added core lengthens the queue and makes handoffs costlier. I attack it in cost order. First, shrink the critical section: if the guarded span covers a `json.Marshal`, a write, a log line or anything else that does not touch the guarded fields, move it out, because that often removes the contention outright. Second, if the section is already tight and traffic is spread across many session ids, partition the lock: an array of shards each with its own mutex and map, routed by a hash of the id, so N ids can proceed at once. Third, for the parts that are read-mostly, publish an immutable snapshot behind an `atomic.Pointer` so readers take no lock at all. Then re-measure between steps: each change moves the bottleneck.

code

go · 11 lines
go
type Store struct {
	mu       sync.Mutex
	entries  map[string]Session // every authenticated request touches this
	lastSeen map[string]time.Time
}

func (s *Store) Touch(id string, now time.Time) {
	s.mu.Lock()
	defer s.mu.Unlock()
	s.lastSeen[id] = now
}

go deeper

for a junior

Know the basic shape of the answer: if only one goroutine at a time can be inside the guarded code, extra cores cannot make that part faster, so the first question is what the lock is being held for.

for a middle

Explain the mechanics: a lock serialises the guarded span, and contention adds parking and cache-line movement on top, which is why the curve can bend downward instead of levelling off.

for a senior

An interviewer expects a diagnosis and an ordered plan, cheapest first, with a re-measurement between steps, plus the honesty to stop when the ceiling turns out to be somewhere else entirely.

for a principal

Own the tradeoff between a five-line critical-section fix and a structural change that every future contributor must understand. Say what evidence would justify the more expensive option and who maintains it afterwards.

## Reading the symptom Throughput that stops rising — or actually falls — when you double the cores is the classic signature of a serialised section. Two things are happening at once. The work under the lock can only be done by one goroutine at a time, so the ceiling is fixed by how long one pass takes, no matter how many cores are available. And the added cores make the *handoff* worse: more goroutines arrive per unit time, each one that fails to acquire the lock parks, and every unlock has to wake a waiter and migrate the lock's cache line to another core. Past a certain arrival rate, more contenders mean less useful work, so the curve bends down instead of merely flattening. A store on the request path — a session store consulted by every authenticated request — reaches that point easily, because its call rate is the service's whole call rate. ## Confirm it is the lock Before changing anything, confirm the diagnosis rather than assuming it: wait time concentrated on one lock, throughput independent of core count, and CPU utilisation that does not rise with the added cores (goroutines are parked, not spinning). If CPU *is* saturated and the wait is elsewhere, this is a different problem and none of the fixes below apply. Also check the obvious environmental explanations first — an artificially low limit on how many goroutines can run Go code at once will flatten throughput too, and no amount of lock surgery fixes that. ## Fix one: make the critical section smaller This is first because it is the cheapest change and the most common cause. Find the `Lock()` and read down to the function's return; in Go a deferred unlock releases at *function* return, so the whole rest of the body is inside the lock. If that span contains encoding, a file or network operation, a log write, a channel operation, or a call into code you do not own, move it out: take the lock, mutate or copy out what you need, release, then do the slow part on the local copy. A store that held its lock for two milliseconds of I/O and now holds it for two hundred nanoseconds of map assignment has just gained four orders of magnitude of headroom, without changing a single data structure. Related: look for work that is done under the lock *per call* but does not need to be, such as recomputing a derived value, formatting a key, or allocating. Move it above the `Lock()`. ## Fix two: partition the lock If the section is already a few map operations and the contention remains, the problem is that unrelated session ids are queueing behind each other. Replace the single mutex and map with a fixed array of shards, each with its own mutex and its own map, and route every operation through a hash of the id. Now up to N unrelated operations proceed at once. This works only if traffic is spread across keys. If one key is hot — a single tenant, a shared token — every request routes to the same shard and nothing improves; that case needs a different answer, such as caching that entry per goroutine or removing it from the shared structure entirely. Sharding also costs you cross-key atomicity: totals and full iterations now span shards, and anything touching two ids needs both locks taken in a fixed order. ## Fix three: get readers off the lock entirely If the access pattern is overwhelmingly reads with occasional bulk updates, the most effective change is to stop taking a lock to read at all. Keep the state as an immutable value published behind an `atomic.Pointer`; readers do one atomic load and then read a value nobody will ever mutate, and writers build a fresh copy and swap the pointer under a writers-only mutex. A lock acquisition, even an uncontended one, *writes* to memory and moves a cache line; an atomic load does not, so the read path finally scales with cores. The price is that each write copies the whole structure and readers may act on a snapshot that is microseconds old. ## The order matters, and so does re-measuring Do these one at a time and re-run the same load between them. Each fix moves the bottleneck, and the second-most-contended thing is rarely what you predicted. It is entirely normal for fix one to make the problem disappear and turn the whole exercise into a five-line diff. It is also normal for fix two to reveal that the real ceiling was an outbound connection pool or the database, in which case stop — you have found the true constraint and further lock work buys nothing. ## What not to do - **Do not add goroutines.** More contenders for the same lock is the problem, not the solution. - **Do not raise the number of runnable processors past the cores you have**; it does not create parallelism, and under contention it can make handoffs worse. - **Do not reach for a lock-free rewrite first.** It is the largest change, the hardest to review, and it is only justified once the cheap structural fixes are exhausted and measured. - **Do not "fix" it by holding the lock for longer to batch work**, unless you have measured that batching genuinely reduces total lock time; it usually raises tail latency badly. - **Do not skip the re-measurement.** A change that looks obviously right and is not verified is how a team ends up with a sharded store, a copy-on-write index, and the same throughput they started with.

  • Why can throughput actually fall, rather than merely flatten, as you add cores?
    Because contention has its own cost. Every failed acquisition parks a goroutine and every release has to wake one and hand the lock's cache line to another core. Adding cores raises the arrival rate at the lock without raising the rate at which guarded work completes, so a growing share of the machine's time goes into handoff and waiting rather than useful work.
  • You shard the lock and nothing improves. What is the most likely explanation?
    The traffic is concentrated on one key, or the hash is not spreading it. Sharding only creates parallelism between *different* keys; if one session id, tenant or token accounts for most operations, every request routes to the same shard and that shard's mutex is exactly as contended as the single lock was. Check the key distribution before adding shards.
  • When is the right answer none of these, and what would tell you?
    When the lock is not the constraint. If CPU is already saturated, or the wait is really on an outbound connection pool, a database, or a syscall, then removing lock contention just moves the queue one layer down. The check is whether throughput responds at all to the first, cheapest change; if it does not move and the added cores stay idle, the ceiling is somewhere else.

saying these in an interview costs you the question

  • Adding goroutines to a system already queueing on one lock
  • Jumping to a lock-free rewrite before shrinking the critical section
  • Sharding without checking whether traffic is spread across keys
  • Assuming more cores must mean more throughput
  • Applying all three fixes at once and never re-measuring
  • Blaming the garbage collector without evidence