skip to content

As a library maintainer, how do you decide what Close must guarantee about goroutines your Go package started?

level: principalimportance: nice to knowfreq 30%

answer

  1. the guarantee is exported, not internal
  2. ask what the caller's next line does
  3. consider not starting the goroutine at all
  4. a stronger Close trades race for hang
  5. one doc-comment sentence settles it

basics

~20 s

Decide from what callers do next. If they may release resources, end a test, or count goroutines after Close, then Close must block until every goroutine the package started has returned, and say so in the doc comment.

solid answer

~50 s

Treat the guarantee as exported API, not an implementation detail: whatever `Close` does today, callers encode it in their teardown order. The strong contract — Close returns only after every goroutine the package started has returned — lets a caller safely close the file or database handle on the next line, and makes an end-of-test goroutine assertion pass without sleeps. Its cost is that Close can now block, so every blocking operation inside those goroutines must be selectable against the stop signal, or you have traded a leak for a hang. Where in-flight work is genuinely slow, follow the stdlib split: a `Shutdown(ctx)` that drains within a deadline, plus a `Close` that stops immediately. Then write the guarantee in the doc comment and back it with a test, because a promise nobody can see is one the next maintainer will quietly break.

go deeper

for a junior

Take away the habit: if a type in your code starts a goroutine, it needs a Close, and you should read that Close's documentation before writing anything after it in your teardown.

for a middle

Be able to explain the difference between a Close that signals and one that joins, and what each lets the calling code legally do on the next line.

for a senior

Show you would audit the goroutines' blocking operations before promising a joining Close, and that you would back the promise with a goroutine-count test rather than prose.

for a principal

Own the contract as exported API: pick between fire-and-forget, joining, bounded shutdown, and not owning a goroutine at all; justify who pays for each; and land the change with documentation, a test, and a migration story for callers who already depend on today's behaviour.

## Why this is a decision and not a detail A package that starts a goroutine on the caller's behalf has taken on a piece of the caller's lifecycle. The caller cannot see that goroutine, cannot stop it, and cannot join it — the only lever they have is whatever you exported. That makes the behaviour of `Close` an interface commitment: teams will write teardown code that depends on it, and the dependency is invisible in their source. This is the kind of call an API review can and should overrule, which is exactly why it belongs to whoever maintains the package rather than to whoever is writing the goroutine that week. ## The options, and who pays for each **1. Fire-and-forget.** `Close` sets a flag or cancels a context and returns. Cheapest to implement and never blocks. The caller pays: they have no safe point after which the goroutine is quiescent, so any resource they release next is a race, and any test that counts goroutines is flaky. In practice teams paper over it with `time.Sleep` in tests, which is how you know the contract is wrong. **2. Signal and join.** `Close` cancels, then waits for every goroutine the package started to return. The caller gets a real happens-before edge: after `Close` returns, nothing of yours is running. You pay: `Close` can block, so every blocking operation inside your goroutines must have a cancellation case, and you have to be honest about in-flight work that will now be abandoned. **3. Bounded join.** `Shutdown(ctx) error` drains what it can and returns the context's error if the deadline passes; a separate `Close` stops immediately. This is the shape `net/http.Server` uses, and it is the right answer when in-flight work is valuable enough to wait for but slow enough that an unbounded wait is unacceptable. You pay in API surface and in the caller having to decide a deadline. **4. Do not start the goroutine at all.** The strongest option, and the one to consider first: expose a synchronous `Run(ctx) error` that blocks and let the caller decide whether to put it behind a `go` statement. Then the caller already owns the goroutine and the whole question evaporates. Reach for an internal goroutine only when the package genuinely needs one to live across calls — a background refresher, a watcher, a batching writer. ## How to decide Ask what the caller's next line is allowed to be. If a reasonable caller might, immediately after `Close`, close a file the goroutine writes to, return a connection to a pool, unmap a buffer, finish a test, or shut the process down cleanly, the guarantee has to be the joining one. The cost of the weak contract is not paid by you; it is paid at 3am by someone whose teardown raced a goroutine they never knew existed, and it presents as data corruption or a panic in unrelated code rather than as a bug in your package. Ask also what happens to in-flight work. If dropping it is acceptable, join and abandon. If it is not — an audit trail, a partial upload — you need the bounded-join shape so the caller can choose how long to wait, and you need to say what happens when the deadline passes. ## Rolling out a strengthened contract Strengthening `Close` from "signals" to "joins" is a behavioural change even though the signature does not move, and the failure mode it introduces is a hang rather than a race. Before you ship it: - Audit every blocking operation in the goroutines you own and give each one a cancellation path. The one that most often has none is a send on a channel the caller is supposed to be draining — and if the caller has stopped draining precisely because they called `Close`, your new join deadlocks them. - Decide the in-flight policy explicitly and document it, rather than letting the first case that blocks decide it for you. - Ship with a test that proves the guarantee — goroutine count around `New`/`Close`, with no sleep — so the property is enforced rather than asserted in prose. - Consider adding `Shutdown(ctx)` alongside rather than making `Close` unboundedly blocking, so callers who cannot tolerate an indefinite wait have a supported option instead of inventing one. ## Write the sentence The artefact that actually settles this is one line of doc comment on the exported method: *"Close stops the watcher and blocks until the goroutine it started has returned; it is safe to call more than once."* That sentence tells the caller what teardown ordering is legal, tells the next maintainer what they may not silently weaken, and gives a reviewer something concrete to hold a change against. Without it, every caller assumes a different contract, and the ones who assumed the strong version will be right about a third of the time.

  • When is a synchronous Run(ctx) better than a package that starts its own goroutine?
    Almost whenever the work maps to a single call the caller drives. Returning a blocking function hands ownership back: the caller chooses whether to run it concurrently, how to join it, and what to do with its error. Keep an internal goroutine only for state that must live between calls, and then own its shutdown completely.
  • What makes Shutdown(ctx) a better fit than a Close that blocks indefinitely?
    It bounds the wait and makes exceeding it an explicit, returnable error rather than a hang. Callers with a deployment deadline can wait five seconds and move on, which is the split net/http.Server uses: Shutdown drains gracefully within the context, Close stops now.
  • A team reports your package leaks goroutines after they call Close. How do you triage it?
    First establish which contract they assumed. If Close only signals, the goroutine may simply not have returned yet and there is no leak, just a missing join — and that is a documentation and API defect on your side. If Close does join, then some blocking operation in the goroutine has no cancellation path, and that is the bug to find.
  • How do you keep a joining Close from becoming a deadlock for callers?
    Guarantee that every blocking operation inside your goroutines races the stop signal, and never make Close depend on the caller continuing to consume something. If the design genuinely needs the caller to drain, that requirement must be documented on Close itself, and a bounded Shutdown offered as the escape hatch.

saying these in an interview costs you the question

  • Treats Close semantics as an implementation detail callers should not rely on
  • Strengthens Close to block without auditing blocking operations first
  • Starts a background goroutine when a synchronous call would do
  • Ships the guarantee only in a commit message, not the doc comment or a test
  • Offers an unbounded blocking Close with no deadline-based alternative