As module owner, when do you take golang.org/x/sync/errgroup into go.mod instead of hand-rolling the join?
answer
- two decisions with different owners
- who inherits the go.mod line
- binary versus imported library
- the number describes the downstream
- multiply the cap by your replica count
basics
~20 sJudge it by who bears the cost. In a binary the dependency stops with you, and errgroup beats twenty lines of join logic you must own. In a library it joins every importer's module graph, so the bar is higher.
solid answer
~50 sThe technical comparison is not really in dispute: `errgroup` is a small, Go-team-maintained package under `golang.org/x` with no third-party transitive requirements, and the alternative is roughly twenty lines of `WaitGroup`, buffered error channel and manual cancel that your team then owns, tests and gets wrong on the failure path. The decision is about who pays. For a service binary, the dependency stops with you — take it. For a library other teams import, your `go.mod` requirement joins their module graph and their minimum-version selection, so the bar is higher and "vendor the twenty lines" is defensible. I would also separate the two decisions people conflate: taking the dependency is a one-time module call, while the value passed to `SetLimit` is a capacity decision about a shared downstream that belongs in configuration and is owned by whoever owns that downstream's budget.
go deeper
Know that golang.org/x/sync is a real module requirement in go.mod, maintained by the Go team, and that the alternative is writing the join with sync.WaitGroup and a channel yourself.
Be able to describe what a direct requirement costs — it enters importers' module graphs and version selection — and why the concurrency limit is a configuration value rather than a constant.
Argue the library-versus-binary split concretely and show you would derive the limit from the downstream's capacity, then check what it multiplies to across replicas.
Separate the module-graph decision from the capacity decision, name the owner of each, write the standard down in a few lines, and accept that a platform or downstream owner can overrule you on either.
## The two decisions hiding in one question People argue about "should we use errgroup" as though it were a single call. It is two, with different owners: 1. **Does this module take a dependency on `golang.org/x/sync`?** A module-graph decision, owned by whoever owns `go.mod` and constrained by whatever dependency policy the organisation has. 2. **What concurrency does this fan-out get?** A capacity decision about the thing being called, owned by whoever owns that capacity. Answering the first does not answer the second, and a lead who conflates them ends up with a limit hard-coded in a library because it was convenient at the time. ## Decision one: the dependency Arguments in favour are unusually strong here, and worth stating precisely rather than as "it is basically stdlib": - The `golang.org/x` repositories are maintained by the Go team under the same review process as the standard library, and `x/sync` is a small module with no third-party requirements of its own, so the transitive blast radius is nil. - The code it replaces is exactly the code that is easy to get wrong and hard to test: an unbuffered error channel that deadlocks only on the failure path, an `Add` that races the goroutine it counts, a `cancel` that is never called. - It is versioned and reproducible like anything else — minimum version selection picks the lowest version satisfying every requirement, and `go.sum` pins the hashes. Arguments against are situational, and they are where the judgment lives: - **You are writing a library.** Every direct requirement you add appears in the graph of every module that imports you and participates in their version selection. That is not a veto — plenty of good libraries depend on `x/sync` — but it changes who you are spending on. For a widely imported package whose entire use of the dependency is one twenty-line join, keeping the dependency count at zero is a legitimate and common choice. - **The organisation has a real cost per dependency.** Some places require a licence review, an internal proxy entry, a security sign-off, or maintain an allow-list. The engineering cost is trivial; the process cost may not be, and pretending otherwise loses you credibility with the people who run that process. - **You need one narrow thing.** If all you want is a bounded loop with no error propagation and no cancellation, the group buys you little. What should *not* decide it: performance (the difference is unmeasurable next to the work being coordinated), or a general aversion to dependencies applied without regard to who maintains this one. ## Decision two: the limit The concurrency limit is not a property of the code that fans out. It is a property of what is on the other end — a database's connection pool, an upstream's rate limit, a third party's contractual quota. Three consequences: - It belongs in **configuration**, not in a constant, because the right value changes when the downstream is resized and nobody wants a release for that. - A per-process limit is not a system limit. Fifty replicas each allowing eight concurrent calls is a limit of four hundred at the database, and the person who set `8` may never have done that multiplication. If the constraint is genuinely global, a per-process cap is the wrong control and you should say so. - The owner of the downstream can overrule you. If the database team says the service gets sixteen connections, your fan-out's limit is derived from that number, not chosen independently. ## What you would write down A defensible team standard is short: - Service binaries may depend on `golang.org/x/sync` without discussion; the join and cancellation logic is not something we re-implement per service. - Libraries intended for cross-team use justify each direct requirement in review, and a twenty-line internal join is an acceptable alternative there. - Every unbounded fan-out over caller-controlled input is a review defect; a limit is required and its value comes from configuration. - The limit's value is agreed with the owner of the downstream it protects, and reviewed when that downstream is resized. ## How you would be overruled, and why that is fine A platform team's dependency policy, a security review, or the owner of a saturated database can all overturn either decision, and each of them is closer to a cost you cannot see from inside your module. The mark of the level is not winning the argument — it is having separated the two decisions, named the cost each one imposes, and put the number somewhere it can be changed by the person who will need to change it.
- A platform standard bans new direct dependencies in shared libraries. How do you respond?Comply, and scope it. In a library the standard is defending something real — my requirement lands in every importer's graph — so I would inline the join, keep it in one unexported helper with tests for the failure path, and note the duplication. Then I would ask that the standard distinguish libraries from service binaries, since the cost it targets does not exist for a leaf binary.
- Your fan-out limit is 8 and the service runs 50 replicas. What have you actually configured?Up to four hundred concurrent calls at the shared downstream, which is a number nobody chose. A per-process limit only controls a single process; it protects that process's memory and connection pool but says nothing about aggregate load. If the constraint is global, the control has to be global too — a quota enforced at the dependency, or a limit derived from the replica count as part of deployment configuration.
- How would you decide the limit's initial value with no production data?Start from the downstream's stated capacity divided among the callers, not from the machine. For a database that is the connection pool allocated to this service; for a third party, the contracted rate. Pick something conservative, put it in configuration, and treat the first load test as the thing that moves it. A number nobody can change without a release is the failure mode to avoid.
saying these in an interview costs you the question
- Treating x/sync as equivalent to a random third-party module
- Adding a direct requirement to a shared library without justifying it
- Hard-coding the concurrency limit inside a reusable package
- Choosing the limit from the machine's core count for I/O work
- Ignoring that a per-process cap multiplies by the replica count