Several teams embed your Go package in long-lived daemons — how do you decide whether it may start goroutines itself?
answer
- a promise every embedder inherits
- convenience for you, shutdown ordering for them
- could a caller build this themselves without loss
- split construction from start
- stop must mean exited, not signalled
basics
~20 sTreat it as a promise every embedder inherits. Default to blocking calls they wrap; start goroutines only when the abstraction cannot exist without them, and then commit to an explicit start, a bounded count, and a stop that means exited.
solid answer
~60 sThe question is who absorbs the cost. Concurrency I start lands in someone else's shutdown ordering, someone else's leak checks, and someone else's on-call — and they cannot opt out. So my default is a blocking core: the embedder writes `go` and owns the goroutine. I start one only when the abstraction genuinely requires background work no caller could supply, and then I treat it as part of the contract: nothing runs in the constructor, one explicit start, one stop whose return means the goroutine has exited, a documented and bounded goroutine count, and errors delivered to something the caller owns rather than logged in my format. If a consuming team wants the convenience form, I add it beside the blocking one with a name that is visible at the call site, so the safe shape stays the default. And I decide it early, because once teams have built shutdown sequences on the current shape, changing whether I start goroutines is a breaking change even though no signature moved.
code
go · 7 lines// Run watches until ctx is cancelled, then returns. The caller
// decides whether to run it on its own goroutine.
func (w *Watcher) Run(ctx context.Context) error
// StartBackground runs Run on a new goroutine. The returned stop
// blocks until that goroutine has exited.
func (w *Watcher) StartBackground(ctx context.Context) (stop func())go deeper
Take away the default: a package usually blocks and lets its caller start goroutines. If a package you import starts one, find out how it stops before you embed it.
Be able to argue both sides for a specific type and name what the package owes if it starts a goroutine: an explicit start, a bounded count, and a stop whose return means the goroutine has exited.
Show that you have integrated a package that got this wrong and what it cost at shutdown, and describe the review checks you apply before adopting one into a long-lived process.
Own the policy across consumers you do not control: which shape is the default, what you commit to when you deviate, how you serve the team that wants the convenience without moving the default, and how you treat a later goroutine addition as a behavioural compatibility event.
## Why this is a decision and not a preference A package embedded in other teams' long-lived processes is not just code they call; it is code that runs inside their lifecycle. Every goroutine it starts must fit into their startup order, their drain, their shutdown deadline, their leak checks and their on-call runbook. They cannot opt out of it, and they usually cannot see it — it does not appear at any call site. That asymmetry is what makes it an ownership call. You get the convenience; they get the failure modes. A senior engineer can pick the right shape for one API. This question is about setting the policy for a package with several consumers you do not control, and being able to defend it when one of them wants the other answer. ## The default: a blocking core Export the work as a call that runs on the caller's goroutine and returns when it is done or cancelled. Consequences worth stating out loud: - Concurrency appears where the embedder wrote it, so their reviewers see it. - Their shutdown sequence already knows how to wait for something they started. - Errors return on the normal path, so they are handled in their idiom, not logged in yours. - Their leak checks stay meaningful, because nothing of yours survives a call. - How much runs at once is theirs to bound, which matters because they know their machine and you do not. If a consumer wants it in the background, that is one keyword in their code. ## When starting one is justified Some abstractions cannot exist without background work. Multiplexing a single handle among many subscribers. Refreshing a credential before it expires. Coalescing writes on a timer. Draining a kernel-level event source and fanning it out. In those cases the goroutine is the product: a purely synchronous version would just push the same goroutine into every consumer, written five slightly different ways. The test I apply: **could a competent caller implement this concurrency themselves from the synchronous API without losing a property?** If yes, do not start it. If no — because the concurrency is bound to a resource only the package holds — start it, and pay for it in the contract. ## What starting one commits you to 1. **Nothing in the constructor.** Construction allocates and validates. A separate, explicitly named call starts the work, so the call site shows where a lifetime began. 2. **One way to stop, and stopping means stopped.** The stop's return is the guarantee that the goroutine has exited and released its handles. "Signals the watcher to stop" is not something an embedder can order a shutdown around. 3. **A bounded, documented count.** "One goroutine per watcher" is a promise. "One per subscription, plus one per reconnect" is a leak that scales with their traffic and appears nowhere in the signature. 4. **Errors go somewhere the caller owns.** A callback they supply, a method they can poll, or a channel with a written contract. Never a log line in your chosen format — you would be picking their observability stack for them. 5. **A panic in your goroutine must not take their process down for a reason they cannot see.** Decide and document what happens, because they cannot recover in a goroutine they did not start. 6. **It is now API.** Adding a goroutine in a later version can break an embedder's leak test and their shutdown ordering, even though no exported signature changed. Removing one can change whether callbacks arrive concurrently. Both are behavioural compatibility problems, and both need a release note that reads like an API change. ## Serving the consumer who wants the convenience When a large consumer asks for the background form, the answer is rarely "no" and should never be "I will change the default". Offer both, with the safe one as the default and the convenient one obviously named at the call site — a blocking `Run(ctx) error` alongside something whose name says it starts something and hands back a stop. Two consequences: a reviewer in their repo can see which was chosen, and you keep exactly one implementation, with the convenience form as a thin wrapper you can reason about. ## Holding the line, and knowing when not to Expect pushback of the form "every consumer writes the same six lines". Sometimes that is a genuine signal that the concurrency belongs in the package, and the six lines are evidence to take seriously — if all five consumers wrote the same wrapper, the abstraction is missing. But check first whether they wrote the *same* six lines or five different ones; divergence usually means the right bounding differs per consumer, which is precisely the decision you must not make for them. The cost of getting this wrong is asymmetric, and that asymmetry is the argument. Too synchronous, and consumers write a little boilerplate you can later absorb. Too concurrent, and you have shipped a lifetime into several production processes that nobody can end without forking you.
- How do you decide the case is strong enough to start a goroutine?Ask whether a competent caller could build the same concurrency on top of the synchronous API without losing a property. If they could, it is theirs. If they could not — because the concurrency is bound to a resource only the package holds, such as one handle fanned out to many subscribers — it belongs inside, and then the contract has to carry an explicit start, a bounded count, and a stop that means exited.
- Why is adding a background goroutine in a later release a compatibility problem?Because embedders build on observed behaviour, not just signatures. A new goroutine can fail their leak checks, change their shutdown ordering, and make callbacks that were serial arrive concurrently. Nothing in the exported surface changed, so nothing warns them. It needs the same care and the same release note as a signature change.
- Five consuming teams all wrap your blocking call the same way. Does that change your answer?It is real evidence that an abstraction is missing, and worth acting on. But check whether it is the same wrapper or five different ones. Identical wrappers argue for shipping it as an opt-in helper beside the blocking core. Divergent ones argue the opposite: the right bounding differs per consumer, which is exactly the decision you should not make on their behalf.
- What do you do about errors from a goroutine your package started?Give them a destination the caller owns: a handler function they supply, a method they can query, or a documented channel. Logging in your own format picks their observability stack for them and buries failures in a place their alerting does not look. If there is no reasonable destination, that is a further argument that the goroutine belongs in their code, not yours.
saying these in an interview costs you the question
- Treats starting goroutines as an implementation detail rather than a contract
- Changes the default from blocking to background because one consumer asked
- Starts background work inside the constructor for convenience
- Ships a stop that only signals and calls the shutdown story complete
- Lets the goroutine count grow with caller traffic without documenting it
- Logs background errors in the package's own format and considers them handled
- Assumes adding a goroutine later is safe because no signature changed