skip to content

When is sync.Mutex.TryLock the right tool, and why are correct uses of it rare?

level: seniorimportance: nice to knowfreq 26%

answer

  1. it never waits and never queues
  2. returns a bool you must branch on
  3. best-effort work that may be skipped
  4. a retry loop is just a busy-wait
  5. false means you hold nothing

basics

~20 s

TryLock takes a sync.Mutex only if it is free and reports whether it succeeded, never blocking. It fits best-effort work that is genuinely fine to skip, such as a diagnostic snapshot. Most other uses hide a design problem.

solid answer

~50 s

`TryLock` on a `sync.Mutex` attempts the acquisition without waiting and returns a `bool`. It is legitimate when *not getting the lock is an acceptable outcome*: a periodic metrics or debug snapshot of a hot object that would rather skip a tick than block the writers. It is rare because elsewhere the lock is what makes the operation correct, so a false result leaves two bad options - do the work unsynchronised, or silently skip something that mattered. It is also a poor building block: it never joins the waiter queue, so a retry loop is a busy-wait with no fairness, and `sync.Mutex` has no timed or context-aware acquisition to fall back on. The documentation says as much - correct uses exist but are rare, and reaching for it often signals a deeper problem. When it returns false you hold nothing: do not call `Unlock`, and do not touch the guarded fields.

code

go · 8 lines
go
// Best effort: a metrics tick that would rather be skipped than block writers.
func (s *Scoreboard) TrySnapshot() (wins, losses int, ok bool) {
	if !s.mu.TryLock() {
		return 0, 0, false // no lock held here: do not call Unlock
	}
	defer s.mu.Unlock()
	return s.wins, s.losses, true
}

go deeper

for a junior

Know that it exists and what it returns: a bool saying whether the lock was taken, with no waiting. If it returns false you hold nothing, so there is nothing to unlock.

for a middle

Explain why it is not a building block - no queueing, no fairness, no timed acquisition - and why a retry loop around it is a busy-wait rather than a lock with a deadline.

for a senior

Judge whether skipping the work is genuinely correct for the caller, add a counter for failed attempts, and read a TryLock in a diff as a prompt to look at how long the critical section actually is.

for a principal

Treat its appearance as evidence about the design: decide whether the contended component needs a shorter critical section, a different ownership model, or an explicit queue, rather than letting best-effort skipping spread through the codebase.

## The API ```go func (m *Mutex) TryLock() bool ``` It tries to take the mutex and reports whether it got it. It never blocks. If it returns `true` you hold the lock and owe an `Unlock`; if it returns `false` you hold nothing and must not unlock anything. `sync.RWMutex` has the same pair, `TryLock` and `TryRLock`. ## The rare legitimate uses The question to ask is: **is failing to acquire an acceptable, correct outcome?** If the honest answer is *no, the work has to happen*, `TryLock` is the wrong tool and you should just call `Lock`. Cases where the answer is genuinely yes: - **Best-effort observation.** A periodic exporter that samples a hot, contended object for a dashboard. Missing one sample is invisible; adding a blocking reader to the contended path is not. - **Debug and admin endpoints.** A handler that dumps internal state should not be able to stall the production path because an operator hit refresh. - **Opportunistic maintenance.** A background compaction or cache trim that runs on a ticker; if the object is busy, the next tick will do it. - **Detecting that something is already running**, where the alternative is to skip - though a dedicated flag is usually clearer for that, since the mutex is then being used as a signal rather than as protection for data. Notice what these have in common: the caller has a genuinely correct *other* branch, not a degraded one. ## Why it is discouraged **A false result usually leaves no good option.** If the guarded data is what the caller came for, skipping is a silent wrong answer and proceeding without the lock is a data race. Code that returns a zero value on a failed `TryLock` typically means somebody has swapped a hang for a subtler bug. **It provides no waiting semantics at all.** `TryLock` does not enqueue you. A loop that retries is a busy-wait: it burns a CPU, and because it never joins the wait queue it can be beaten repeatedly by goroutines that do. Building a lock-with-timeout out of it is the classic misuse - `sync.Mutex` has no timed acquisition, and simulating one this way costs CPU and gives no fairness guarantee. **It hides the real problem.** Wanting to not wait is usually a statement about the design: the critical section is too long, the lock covers too much, or the work should not have been sharing state in the first place. Those are structural fixes. `TryLock` lets you avoid making them, which is precisely why the standard library documents correct uses as rare and treats its appearance as a smell worth investigating. **It is invisible to reviewers.** A `Lock` that is held too long shows up as latency you can measure. A `TryLock` that keeps failing shows up as work quietly not happening - no error, no metric, no log line, unless you added one. ## Using it correctly when you do ```go func (s *Scoreboard) TrySnapshot() (wins, losses int, ok bool) { if !s.mu.TryLock() { return 0, 0, false } defer s.mu.Unlock() return s.wins, s.losses, true } ``` The discipline: - **Branch on the bool, always.** Ignoring the result is a compile-time-legal way to write a race. - **Only `defer Unlock` inside the success branch**, after the check - never before it. Unlocking a mutex you did not take is a fatal runtime error that kills the process. - **Report the miss.** If the skipped work matters at all, count it. A counter of failed acquisitions is what later tells you the contention got worse. - **Do not loop on it.** If you find yourself retrying, you wanted `Lock`. ## Version note `TryLock` is not part of the original API: `sync.Mutex.TryLock`, `sync.RWMutex.TryLock` and `sync.RWMutex.TryRLock` were added in Go 1.18. Before that the language deliberately offered no non-blocking acquisition at all, and the absence was itself the argument - the designers' position being that code which needs to not wait usually needs a different structure. That position has not changed; only the escape hatch was added, with a documentation comment warning you about it. ## The interview answer in one line Use it when skipping is correct, never as a way to avoid waiting, never in a retry loop, and never to build a timeout - and always branch on the result before you unlock anything.

  • How would you give a sync.Mutex acquisition a timeout?
    You cannot: `sync.Mutex` has no timed or context-aware `Lock`, and looping on `TryLock` until a deadline is a busy-wait with no fairness - queued waiters can beat you indefinitely. Treat the requirement as a design signal instead: shorten the critical section so waiting is bounded in practice, move the contended work behind a queue you control, or use a cancellable coordination primitive rather than a plain mutex.
  • What must a caller not do when TryLock returns false?
    Two things. Do not call `Unlock` - the goroutine holds nothing, and unlocking an unheld mutex is a fatal runtime error that terminates the process rather than a recoverable panic. And do not read or write the guarded fields on that path; the whole point of the false result is that another goroutine may be mid-update. Take the alternative branch, and count the miss if it matters.
  • Is a metric worth adding around a TryLock that fails?
    Yes, whenever the skipped work has any value. A failed acquisition produces no error, no log line and no latency spike, so a rising failure rate is otherwise invisible - the first symptom is stale data nobody can explain. A simple counter of attempts and misses turns a silent skip into something you can alert on and correlate with contention.

saying these in an interview costs you the question

  • Uses TryLock in a retry loop to emulate a lock timeout
  • Calls Unlock after TryLock returned false
  • Reads the guarded fields anyway when acquisition fails
  • Thinks TryLock joins the queue and gets priority
  • Reaches for TryLock instead of shortening the critical section
  • Skips work on a failed attempt with no counter or log