skip to content

Coordination Patterns

The compositions you build per unit of work: a pool over a jobs channel, a cancellable pipeline, a merge, a group that stops on the first failure. Each is a few lines that hide a deadlock or a leak.

part ofGo (Golang)overview, primer and where to startread it →
on this pageshow

explore

questions

page 1 of 2

What does golang.org/x/sync/errgroup add over a sync.WaitGroup for concurrent tasks that can fail?

level: juniorimportance: must knowfreq 72%

answer

  1. one of them counts, one of them reports
  2. the task signature is the difference
  3. func() versus func() error
  4. join point that can also say why
  5. first non-nil error plus sibling cancellation

basics

~20 s

errgroup joins goroutines like sync.WaitGroup, but each task is a func() error. Group.Wait blocks until every task has returned and gives back the first non-nil error, and errgroup.WithContext also cancels the siblings when one task fails.

solid answer

~40 s

A `sync.WaitGroup` only counts goroutines. It has no place to put a failure, so you end up wiring a buffered error channel next to it, draining that channel after `Wait`, and deciding yourself which error to report. `errgroup.Group` folds all of that in: `g.Go(func() error { ... })` starts a task, and `g.Wait()` blocks until every task has returned and returns the first non-nil error one of them produced. If you build the group with `errgroup.WithContext(ctx)` you also get back a derived `context.Context` that is cancelled as soon as a task fails, so the siblings can stop instead of finishing work nobody will use. `SetLimit` bounds how many tasks run at once. The cost is that only one error survives `Wait` — if you need all of them you have to record them yourself.

code

go · 17 lines
go
var wg sync.WaitGroup
errs := make(chan error, len(urls)) // buffered: a failing task must never block

for _, u := range urls {
	wg.Go(func() { // Go 1.25+; before that: wg.Add(1) and defer wg.Done()
		if err := fetch(u); err != nil {
			errs <- err
		}
	})
}

wg.Wait()
close(errs)

for err := range errs { // and now you decide what to do with them
	log.Println(err)
}

go deeper

for a junior

Be ready to state the two method names and what they take: Group.Go takes a func() error, Group.Wait blocks and returns the first non-nil error. Then say the one thing a WaitGroup cannot do — report a failure.

for a middle

Explain the boilerplate errgroup replaces line by line: the buffered error channel, its size, the close after Wait, and the manual cancel func. Note that Go 1.25's WaitGroup.Go removes the Add/Done bug but not the missing error path.

for a senior

Show judgment about when the join needs an answer at all. Point out that only one error survives Wait, that cancellation is cooperative, and that tasks writing into distinct slots need no mutex while a shared append does.

for a principal

Frame it as a build-versus-depend call: twenty lines of join logic your team owns and tests, against a line in go.mod that lands in every importer's module graph. Say who in your organisation gets to make that call.

## What a WaitGroup actually gives you `sync.WaitGroup` is a counter with a barrier. You increment it before starting a goroutine, each goroutine decrements it when it finishes, and `Wait` blocks until the counter reaches zero. That is the whole contract. It answers exactly one question — *are they all done?* — and it has no channel through which a goroutine can say *and this one failed*. So the moment your concurrent tasks can fail, a `WaitGroup` alone stops being enough, and you write the same boilerplate every time: - a buffered error channel, sized to the number of tasks so a failing goroutine can always send without blocking (an unbuffered one deadlocks if nobody is receiving yet); - a `close` after `Wait` so the channel can be ranged over; - a drain loop that picks which of the collected errors to return; - and, if you want the siblings to stop when one fails, a `context.CancelFunc` you create, capture and call by hand. That is roughly twenty lines that are easy to get subtly wrong and that you now own and must test. ## What errgroup.Group is `errgroup.Group` from `golang.org/x/sync/errgroup` is that boilerplate packaged. Its zero value is usable, and it has a small surface: - `g.Go(f func() error)` runs `f` in a new goroutine and remembers its error. - `g.Wait() error` blocks until every function passed to `Go` has returned, then returns the first non-nil error any of them produced (or `nil` if none failed). - `errgroup.WithContext(ctx)` returns a `*Group` and a derived `context.Context`; that context is cancelled the first time a task returns a non-nil error, and also once `Wait` returns. - `g.SetLimit(n)` caps how many tasks run concurrently; `g.TryGo(f)` starts a task only if the group is currently below that cap. The crucial difference is the task signature. A `WaitGroup` goroutine is a `func()` — it has nowhere to put a result. An errgroup task is a `func() error`, so the group can do something with failure: record it, and (with `WithContext`) cancel everyone else. ## Note on sync.WaitGroup.Go Go 1.25 added `WaitGroup.Go`, which starts a goroutine and handles the `Add`/`Done` pairing for you. That removes the most common `WaitGroup` bug — an `Add` that happens inside the goroutine instead of before it, or a missing `Done` on an early return — but it does **not** close the gap: `WaitGroup.Go` still takes a `func()` with no error, and `WaitGroup.Wait` still returns nothing. Error collection and sibling cancellation remain yours to build. ## Where each one belongs Use a `WaitGroup` when the goroutines cannot meaningfully fail, or when their results already flow somewhere else — into a channel a consumer is reading, into a shared metric, into distinct slots of a slice. Use an errgroup when the join point genuinely needs an answer: *did all of this work succeed, and if not, stop the rest and tell me.* The canonical fit is a request handler that fans out to several independent backends and cannot render its response unless all of them answer. The group starts one task per backend, each task writes its own piece into a distinct field or slice index (distinct destinations, so no mutex is needed), and `Wait` collapses the whole fan-out into a single `if err != nil` at the call site. ## What errgroup does not do for you - **It keeps only one error.** The first task to return a non-nil error wins; the others are discarded. If the failures are diagnostically interesting, log or record each one inside its own task before returning. - **Cancellation is cooperative.** The context that `WithContext` hands back is only useful if the tasks actually pass it to the calls they make or select on `ctx.Done()`. A task that ignores it runs to completion, and `Wait` waits for it. - **The context dies when Wait returns.** Do not keep using the derived context after the join — derive a fresh one for anything that outlives the group. - **It is a dependency.** `golang.org/x/sync` is maintained by the Go team and has no third-party transitive requirements, but it is still a line in `go.mod` and a decision somebody owns. ## The shape to remember A `WaitGroup` answers "are they finished?". An errgroup answers "are they finished, did any of them fail, and should the rest keep going?" — which is what most real fan-outs actually need.

  • Is the zero value of errgroup.Group usable, or must you always call errgroup.WithContext?
    The zero value works: `var g errgroup.Group`, then `g.Go(...)` and `g.Wait()`. What you lose is cancellation — without `WithContext` there is no derived context, so a failure in one task does nothing to the others and `Wait` simply waits for all of them before returning the first error. Use the zero value when the tasks are short and independent; use `WithContext` when abandoning the rest is worth something.
  • Can several errgroup tasks write into the same result struct without a mutex?
    Yes, provided each task writes to a *different* field or a different slice index and nobody reads those locations until after `Wait`. Distinct memory locations are not a data race, and `Wait` gives the happens-before edge that makes the writes visible to the joining goroutine. Two tasks appending to the same slice, or writing the same field, is a race and needs a mutex — or a redesign into distinct slots.
  • What is wrong with using an unbuffered error channel in the hand-rolled version?
    If nothing is receiving yet, the first goroutine that tries to send its error blocks forever, its `Done` never runs, and `Wait` never returns — a deadlock that only appears on the failure path, which is exactly the path least covered by tests. Buffering the channel to the number of tasks guarantees every sender can deposit its error and exit.

A WaitGroup is a turnstile counter: it only tells you everyone has left the building. An errgroup is a shift supervisor: it tells you everyone has left, hands you the first incident report, and can call the rest back in when one worker hits trouble.

saying these in an interview costs you the question

  • Claiming sync.WaitGroup.Wait returns an error
  • Saying errgroup returns all the errors from Wait
  • Thinking errgroup cancels siblings without WithContext
  • Passing a func() with no error return to Group.Go
  • Using an unbuffered error channel in the hand-rolled join
open as a page

How do you merge several receive-only Go channels into one channel a consumer can range over?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Start one forwarding goroutine per source that copies every value into a single shared output channel, and start one extra goroutine that waits for all forwarders to finish and then closes that output exactly once.

open as a page

In a Go pipeline, why does each stage return a receive-only `<-chan Out` and close only that channel?

level: juniorimportance: must knowfreq 60%

basics

~20 s

A stage owns the one channel it creates and sends on, so only it may close that channel, in a defer when its input runs dry. Returning that channel receive-only makes any downstream close or send a compile error.

open as a page

In golang.org/x/time/rate, what is the difference between Limiter.Allow and Limiter.Wait?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Allow never blocks: it takes a token if one is free and returns true, otherwise false, so the caller must shed or retry. Wait blocks until a token is free, the context is cancelled, or its deadline passes.

open as a page

How does a buffered `chan struct{}` limit how many goroutines do work at the same time?

level: juniorimportance: must knowfreq 60%

basics

~20 s

A buffered channel of capacity N holds N permits. Sending an empty struct takes a permit and blocks once N are outstanding; a receive in a defer gives it back. At most N goroutines run the guarded work.

open as a page

Why start a fixed number of worker goroutines reading one jobs channel instead of one goroutine per job?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A fixed pool caps how much work runs at once, so memory, open sockets and downstream load stay bounded however many jobs arrive. One goroutine per job lets a million jobs become a million concurrent callers.

open as a page

How should a Go producer avoid blocking forever when the bounded channel it feeds stays full?

level: middleimportance: must knowfreq 55%

basics

~20 s

Wrap the send in a select that also waits on ctx.Done() and, if you need a deadline, on a timer channel. The send then either succeeds, is abandoned because the caller gave up, or is abandoned after a bounded wait — and the give-up branch must do something explicit with the event.

open as a page

With errgroup.WithContext, what cancels the derived context, and why might siblings keep running anyway?

level: middleimportance: must knowfreq 66%

basics

~20 s

The context from errgroup.WithContext is cancelled when the first task returns a non-nil error, and when Group.Wait returns. Cancellation is only a signal: a task that never checks ctx.Done() or passes the context down runs to completion.

open as a page

In a first-result-wins fan-out over N mirrors, why must the result channel have one buffer slot per attempt?

level: middleimportance: must knowfreq 55%

basics

~20 s

Because only one value is ever received. With an unbuffered channel the losing goroutines block forever on their send, leaking themselves and everything they hold. One slot per attempt lets every sender complete and exit even though nobody is listening.

open as a page

If a Go pipeline's consumer stops receiving halfway, what happens to the upstream stage goroutines?

level: middleimportance: must knowfreq 68%

basics

~20 s

They block forever on their next send and never return, pinning every value they hold. The fix: thread a context through every stage and write each send as a select against cancellation, so an abandoned stage exits.

open as a page

In golang.org/x/time/rate, what do the 10 and the 100 in rate.NewLimiter(10, 100) mean?

level: middleimportance: must knowfreq 55%

basics

~20 s

The first argument is the sustained rate: 10 tokens refilled per second. The second is burst, the bucket's capacity: after an idle spell up to 100 calls go out back-to-back, then the pace settles to 10 per second.

open as a page

In a worker pool, why must close(results) happen only after wg.Wait() returns?

level: middleimportance: must knowfreq 68%

basics

~20 s

Closing a channel that a worker might still send on panics. wg.Wait() returning is the only proof that every worker has left its loop, so close(results) must follow it. close(jobs) merely tells the workers to finish.

open as a page

In Go, why does appending incoming events to a slice instead of a bounded channel remove backpressure?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Appending to a slice never blocks, so nothing ever slows the producer down. A bounded channel makes the sender wait once the buffer is full; a slice just grows, so the backlog turns into heap until the process runs out of memory.

open as a page

You launch three goroutines to fetch the same file from three mirrors. How does the caller return as soon as the fastest one answers?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Give every attempt the same result channel and receive from it exactly once. That single receive yields the first value any goroutine sends, so the caller returns with it immediately and never waits for the other two attempts to finish.

open as a page

What do errgroup.Group.SetLimit and TryGo do, and when does SetLimit panic?

level: middleimportance: should knowfreq 42%

basics

~20 s

errgroup.Group.SetLimit(n) caps the group at n active tasks, so Group.Go blocks until a slot frees; a negative n means no limit. TryGo starts a task only if there is room, returning false otherwise. SetLimit panics while tasks are active.

open as a page

When several Go channels are merged into one, what ordering does the merged channel guarantee?

level: middleimportance: should knowfreq 46%

basics

~20 s

Values from any one source keep that source's order, because a single goroutine forwards it. Across sources nothing is guaranteed: the merged channel interleaves arbitrarily, with no round-robin, no fairness and no ordering by timestamp.

open as a page

Racing three mirrors for one archive, how do you return the first successful download rather than the first attempt to finish?

level: middleimportance: should knowfreq 42%

basics

~20 s

Receive in a loop instead of once. Return the first result whose error is nil; collect the failures and keep receiving. Only after all attempts have reported do you give up, returning the accumulated errors together.

open as a page

How does singleflight.Group.Do collapse many identical concurrent calls into one execution?

level: middleimportance: should knowfreq 45%

basics

~20 s

golang.org/x/sync/singleflight keys in-flight work. The first caller for a key runs the function; others arriving with that key while it runs block and receive the same result and error. Nothing is cached once it returns.

open as a page

Why send on a semaphore `chan struct{}` before the `go` statement rather than inside the goroutine?

level: middleimportance: should knowfreq 45%

basics

~10 s

Acquiring before the go statement makes the spawning loop block, so only about N goroutines ever exist. Acquiring inside the goroutine starts one per item immediately: it bounds work in flight, not memory.

open as a page

If three worker goroutines each run `for j := range jobs` on the same channel, how many of them see each job?

level: middleimportance: should knowfreq 60%

basics

~20 s

Exactly one. A channel receive is a hand-off, not a broadcast: each value goes to a single receiver, so three workers ranging over one channel split the jobs between them rather than each processing all of them.

open as a page

A Go ingestion service's live heap climbs under steady input — how do you find the unbounded backlog?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Confirm the growth is live data with GODEBUG=gctrace=1, then take a heap profile with inuse_space to see which structure holds the bytes, and check the goroutine profile: if no producer is parked in a channel send, the backlog is absorbing work that a bounded channel should have stalled.

open as a page

Why does errgroup.Group.Wait report only one failure when a fan-out to five upstreams has two of them fail?

level: seniorimportance: should knowfreq 38%

basics

~20 s

errgroup.Group.Wait keeps only the first non-nil error a task returned and discards the rest, so a second failure never reaches the caller. Log each failure inside its own task, or collect per-task errors and combine them with errors.Join.

open as a page

Why do a Go fan-in merge's forwarding goroutines leak when the consumer abandons the merged channel?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Each forwarder is blocked sending into the merged channel. Nothing will ever receive again, so those goroutines park forever, keep their last value and their source alive, and the merged channel never closes. Give every forwarder a cancellation case.

open as a page

A Go pipeline panics with `send on closed channel` in its decode stage — what does that tell you about a downstream stage?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Something other than decode closed decode's output channel while decode was still sending, almost always the next stage closing its input. Only the sending goroutine may close a channel; returning stages as receive-only makes a downstream close a compile error.

open as a page

A worker keeps one rate.Limiter per tenant in a map. Why does that map grow without bound, and what do you do about it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Limiters are created on first sight of a key and never removed, so the map keeps an entry for every tenant ever seen. Bound it: key on a set you control, and evict entries idle longer than the refill time.

open as a page

A bulk uploader's in-flight permit count sits pinned at its semaphore limit with nothing finishing — how do you find the cause?

level: seniorimportance: should knowfreq 35%

basics

~20 s

That shape means permits were taken and never returned. Look for an acquire whose release is not deferred immediately after it, since an early error return above the defer never registers it. Confirm with the goroutine profile.

open as a page

A batch runner sends all 500 jobs before reading from an unbuffered results channel and stops making progress. Why?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Every worker is blocked sending its first result, because the only goroutine that would receive from the results channel is the one still sending jobs. With no worker receiving, the jobs send blocks too, and nothing can move.

open as a page

When a Go ingest pipeline's bounded channel stays full, how do you decide between stalling upstream and dropping events?

level: principalimportance: should knowfreq 30%

basics

~20 s

It is a data-loss decision, not a coding one. Stall only if the upstream can absorb waiting without failing itself; shed only for data whose owner has agreed it may be lost, per class, with a counter and an alert. Decide it before the incident and make it configurable.

open as a page

In a Go service, what does exporting len(jobs) on a bounded intake channel actually measure?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

It measures how many items are sitting in that channel's buffer at the instant you read it — nothing else. It excludes items already taken by workers and producers blocked waiting to send, it is stale immediately, and it is always zero for an unbuffered channel.

open as a page

A Go fan-in merge calls wg.Wait() and close(out) inline before returning out — why does it hang?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Because the forwarders are blocked sending. With an unbuffered output channel nothing can receive until the caller holds the channel, so every forwarder parks on its send, the WaitGroup counter never reaches zero, and Wait never returns.

open as a page

showing 1–30 of 35