skip to content

How does singleflight.Group.Do collapse many identical concurrent calls into one execution?

level: middleimportance: should knowfreq 45%

answer

  1. a map of work currently running
  2. keyed by what makes the calls identical
  3. the window is only while it is in flight
  4. the losers get the winner's error too
  5. the entry disappears when the call returns

basics

~20 s

golang.org/x/sync/singleflight keys in-flight work. The first caller for a key runs the function; others arriving with that key while it runs block and receive the same result and error. Nothing is cached once it returns.

solid answer

~40 s

`singleflight.Group` holds a map from key to the call currently running for that key, guarded by a mutex. `Do(key, fn)` looks the key up: with no call in flight the caller becomes the owner, runs `fn`, publishes the result and removes the entry; otherwise it waits and gets the identical `(value, err, shared)` triple. That turns a hundred goroutines that all missed the same lookup into one upstream request. Two properties matter. It is **not a cache** — the key is dropped when the call finishes, so the next caller runs `fn` again. And everyone shares the outcome, **including the error**, so one failed or slow call is delivered to every waiter. `DoChan` returns a channel so a caller can select on its own context, and `Forget(key)` drops the in-flight association.

code

go · 33 lines
go
type call struct {
	done chan struct{}
	val  string
	err  error
}

type group struct {
	mu sync.Mutex
	m  map[string]*call
}

func (g *group) Do(key string, fn func() (string, error)) (string, error) {
	g.mu.Lock()
	if c, ok := g.m[key]; ok {
		g.mu.Unlock()
		<-c.done // a duplicate caller waits, it does not call fn
		return c.val, c.err
	}
	c := &call{done: make(chan struct{})}
	if g.m == nil {
		g.m = make(map[string]*call)
	}
	g.m[key] = c
	g.mu.Unlock()

	c.val, c.err = fn()

	g.mu.Lock()
	delete(g.m, key) // the key lives only while fn runs: this is not a cache
	g.mu.Unlock()
	close(c.done)
	return c.val, c.err
}

go deeper

for a junior

Know the shape: many goroutines want the same answer at once, one of them does the work, the rest wait and receive it. Be able to name what the key represents.

for a middle

Explain the map of in-flight calls, that the entry disappears when the function returns, and that the result and the error are shared identically by every waiter.

for a senior

Talk about the failure modes you have hit: one bad attempt served to a hundred callers, waiters pinned to the leader's latency, a key too coarse to be safe, and where the limiter belongs relative to the coalescing layer.

for a principal

Judge when coalescing is worth the coupling it introduces — it ties otherwise independent callers to one execution's fate — and set the expectation for which shared paths must be coalesced and which must stay independent.

## The problem it solves A moment arrives when many goroutines want exactly the same expensive answer at exactly the same time: an entry expires and every in-progress request needs it, a process starts cold, a dependency comes back after a blip. Without coordination each of those goroutines issues its own identical upstream call. That herd is wasteful when the upstream is merely slow and dangerous when the upstream is metered or already struggling — the exact moment you most want to be gentle is the moment you hit it hardest. **Coalescing** — also called request collapsing or deduplication — means: run the work once, hand the answer to everyone who was waiting for it. ## The mechanism `golang.org/x/sync/singleflight` implements it with about as little machinery as you would expect: - a `Group` owns a mutex and a `map[string]*call`; - `Do(key string, fn func() (any, error)) (v any, err error, shared bool)` locks the map and looks for the key; - **miss**: it inserts a fresh call record, unlocks, runs `fn` itself, stores the value and error in the record, removes the key from the map, and releases every waiter; - **hit**: it unlocks and blocks on the existing record, then returns the value and error stored there. The third return value, `shared`, tells a caller whether the result was delivered to more than one caller — useful mainly as a metric, since a rising share of coalesced calls is direct evidence that the herd exists. `DoChan` is the same thing with a channel instead of a block: it returns `<-chan Result`, where `Result` carries `Val`, `Err` and `Shared`. `Forget(key)` removes the key's in-flight entry so that subsequent callers start a new execution rather than joining the current one. ## What it is not It is **not a cache**, and conflating the two is the most common misunderstanding. The key exists only while the function is running. Two calls a millisecond apart, one starting after the other finished, both run `fn`. Coalescing shrinks a burst of N simultaneous calls to one; it does nothing at all for a steady stream of sequential ones. Deduplication across time is a caching decision and lives elsewhere; deduplication across concurrent callers is what this gives you. ## The sharp edges **Everyone shares the error.** If the single execution fails, every waiter gets that same failure, so one unlucky attempt can be amplified across every caller that happened to arrive during the window. If the failure is transient and per-attempt, you have just made a hundred callers fail for the price of one bad request. Retrying inside `fn`, or calling `Forget` before returning an error so the next arrival starts fresh, are the usual answers. **Everyone shares the latency.** A waiter's own deadline does not shorten the leader's call. `Do` blocks until the in-flight function returns, so a caller with a 100ms budget can be held for as long as the leader takes. `DoChan` plus a `select` on the caller's `ctx.Done()` is how you let an impatient caller leave; note that leaving does not cancel the underlying work, it just stops you waiting on it. It also means the leader's context should be one appropriate to *shared* work rather than one that dies when the first caller gives up. **Choose the key carefully.** The key must identify the work exactly. Too coarse and you serve one caller another caller's answer — a correctness bug, and a nasty one if the answer is tenant-scoped. Too fine and nothing ever coalesces. Include every input that changes the result, tenant identity first among them. **A panic propagates.** If the shared function panics, the package does not swallow it; the panic surfaces in the waiting callers too. Recover inside `fn` if the callers must not die together. ## How it pairs with a rate limiter Coalescing and rate limiting solve different halves of the same overload. A limiter decides **how often** calls may leave; coalescing decides **whether two identical calls need to be two calls at all**. Put the coalescing outermost and the limiter inside the shared function, so the one execution that actually happens is the one that spends a token — the waiters cost nothing and consume nothing. Reversing them makes every duplicate caller pay a token before discovering it did not need to make the call, which is precisely the capacity you were trying to save.

  • The shared function returns an error. What do the duplicate callers receive?
    The same error. Every waiter on that key gets the leader's exact result and error, so a single transient failure is amplified to everyone who arrived during the window. Retry inside the shared function, or call `Forget` on the key before returning a failure so the next arrival starts a fresh execution instead of inheriting the bad one.
  • Why does Do sometimes need to be replaced by DoChan?
    Because `Do` blocks until the in-flight execution finishes, regardless of the waiting caller's own deadline. `DoChan` returns a `<-chan Result` so the caller can select between that channel and its `ctx.Done()`, and give up on its own schedule. Leaving does not cancel the shared work; it only stops that caller waiting for it.
  • What goes wrong if the key is too coarse?
    Callers receive an answer computed for different inputs. If the key omits tenant identity, or a filter, or a version, one caller's result is served to another — a correctness and sometimes a data-exposure bug rather than a performance one. The key must contain every input that can change the result.
  • Where should a rate limiter sit relative to the coalescing call?
    Inside the shared function, so only the single execution that really happens spends a token. If each duplicate caller takes a token before entering the coalescing layer, they consume the very capacity that coalescing exists to protect, and the limiter reports far more traffic than actually left the process.

One person goes to the counter and everyone else in the queue for the same item waits for them to come back. They all get the same answer, including "they were out of stock".

saying these in an interview costs you the question

  • Calls it a cache with an automatic expiry
  • Assumes each waiter re-runs the function on error
  • Thinks a waiter's context cancels the shared call
  • Keys the group by something that omits tenant identity
  • Expects sequential duplicate calls to be deduplicated