skip to content

As module owner, when do you take golang.org/x/sync/errgroup into go.mod instead of hand-rolling the join?

level: principalimportance: nice to knowfreq 26%

answer

  1. two decisions with different owners
  2. who inherits the go.mod line
  3. binary versus imported library
  4. the number describes the downstream
  5. multiply the cap by your replica count

basics

~20 s

Judge it by who bears the cost. In a binary the dependency stops with you, and errgroup beats twenty lines of join logic you must own. In a library it joins every importer's module graph, so the bar is higher.

solid answer

~50 s

The technical comparison is not really in dispute: `errgroup` is a small, Go-team-maintained package under `golang.org/x` with no third-party transitive requirements, and the alternative is roughly twenty lines of `WaitGroup`, buffered error channel and manual cancel that your team then owns, tests and gets wrong on the failure path. The decision is about who pays. For a service binary, the dependency stops with you — take it. For a library other teams import, your `go.mod` requirement joins their module graph and their minimum-version selection, so the bar is higher and "vendor the twenty lines" is defensible. I would also separate the two decisions people conflate: taking the dependency is a one-time module call, while the value passed to `SetLimit` is a capacity decision about a shared downstream that belongs in configuration and is owned by whoever owns that downstream's budget.

go deeper

for a junior

Know that golang.org/x/sync is a real module requirement in go.mod, maintained by the Go team, and that the alternative is writing the join with sync.WaitGroup and a channel yourself.

for a middle

Be able to describe what a direct requirement costs — it enters importers' module graphs and version selection — and why the concurrency limit is a configuration value rather than a constant.

for a senior

Argue the library-versus-binary split concretely and show you would derive the limit from the downstream's capacity, then check what it multiplies to across replicas.

for a principal

Separate the module-graph decision from the capacity decision, name the owner of each, write the standard down in a few lines, and accept that a platform or downstream owner can overrule you on either.

## The two decisions hiding in one question People argue about "should we use errgroup" as though it were a single call. It is two, with different owners: 1. **Does this module take a dependency on `golang.org/x/sync`?** A module-graph decision, owned by whoever owns `go.mod` and constrained by whatever dependency policy the organisation has. 2. **What concurrency does this fan-out get?** A capacity decision about the thing being called, owned by whoever owns that capacity. Answering the first does not answer the second, and a lead who conflates them ends up with a limit hard-coded in a library because it was convenient at the time. ## Decision one: the dependency Arguments in favour are unusually strong here, and worth stating precisely rather than as "it is basically stdlib": - The `golang.org/x` repositories are maintained by the Go team under the same review process as the standard library, and `x/sync` is a small module with no third-party requirements of its own, so the transitive blast radius is nil. - The code it replaces is exactly the code that is easy to get wrong and hard to test: an unbuffered error channel that deadlocks only on the failure path, an `Add` that races the goroutine it counts, a `cancel` that is never called. - It is versioned and reproducible like anything else — minimum version selection picks the lowest version satisfying every requirement, and `go.sum` pins the hashes. Arguments against are situational, and they are where the judgment lives: - **You are writing a library.** Every direct requirement you add appears in the graph of every module that imports you and participates in their version selection. That is not a veto — plenty of good libraries depend on `x/sync` — but it changes who you are spending on. For a widely imported package whose entire use of the dependency is one twenty-line join, keeping the dependency count at zero is a legitimate and common choice. - **The organisation has a real cost per dependency.** Some places require a licence review, an internal proxy entry, a security sign-off, or maintain an allow-list. The engineering cost is trivial; the process cost may not be, and pretending otherwise loses you credibility with the people who run that process. - **You need one narrow thing.** If all you want is a bounded loop with no error propagation and no cancellation, the group buys you little. What should *not* decide it: performance (the difference is unmeasurable next to the work being coordinated), or a general aversion to dependencies applied without regard to who maintains this one. ## Decision two: the limit The concurrency limit is not a property of the code that fans out. It is a property of what is on the other end — a database's connection pool, an upstream's rate limit, a third party's contractual quota. Three consequences: - It belongs in **configuration**, not in a constant, because the right value changes when the downstream is resized and nobody wants a release for that. - A per-process limit is not a system limit. Fifty replicas each allowing eight concurrent calls is a limit of four hundred at the database, and the person who set `8` may never have done that multiplication. If the constraint is genuinely global, a per-process cap is the wrong control and you should say so. - The owner of the downstream can overrule you. If the database team says the service gets sixteen connections, your fan-out's limit is derived from that number, not chosen independently. ## What you would write down A defensible team standard is short: - Service binaries may depend on `golang.org/x/sync` without discussion; the join and cancellation logic is not something we re-implement per service. - Libraries intended for cross-team use justify each direct requirement in review, and a twenty-line internal join is an acceptable alternative there. - Every unbounded fan-out over caller-controlled input is a review defect; a limit is required and its value comes from configuration. - The limit's value is agreed with the owner of the downstream it protects, and reviewed when that downstream is resized. ## How you would be overruled, and why that is fine A platform team's dependency policy, a security review, or the owner of a saturated database can all overturn either decision, and each of them is closer to a cost you cannot see from inside your module. The mark of the level is not winning the argument — it is having separated the two decisions, named the cost each one imposes, and put the number somewhere it can be changed by the person who will need to change it.

  • A platform standard bans new direct dependencies in shared libraries. How do you respond?
    Comply, and scope it. In a library the standard is defending something real — my requirement lands in every importer's graph — so I would inline the join, keep it in one unexported helper with tests for the failure path, and note the duplication. Then I would ask that the standard distinguish libraries from service binaries, since the cost it targets does not exist for a leaf binary.
  • Your fan-out limit is 8 and the service runs 50 replicas. What have you actually configured?
    Up to four hundred concurrent calls at the shared downstream, which is a number nobody chose. A per-process limit only controls a single process; it protects that process's memory and connection pool but says nothing about aggregate load. If the constraint is global, the control has to be global too — a quota enforced at the dependency, or a limit derived from the replica count as part of deployment configuration.
  • How would you decide the limit's initial value with no production data?
    Start from the downstream's stated capacity divided among the callers, not from the machine. For a database that is the connection pool allocated to this service; for a third party, the contracted rate. Pick something conservative, put it in configuration, and treat the first load test as the thing that moves it. A number nobody can change without a release is the failure mode to avoid.

saying these in an interview costs you the question

  • Treating x/sync as equivalent to a random third-party module
  • Adding a direct requirement to a shared library without justifying it
  • Hard-coding the concurrency limit inside a reusable package
  • Choosing the limit from the machine's core count for I/O work
  • Ignoring that a per-process cap multiplies by the replica count