skip to content

Why is cancelling a Go service's root context.Context not enough to shut it down cleanly?

level: juniorimportance: should knowfreq 55%

answer

  1. a signal, not a join
  2. who is still running when main returns?
  3. the runtime waits for nobody
  4. cancel, then wait, then close

basics

~20 s

Cancelling the root context only signals goroutines to stop; it never waits for them. If main returns first, the process dies mid-work. A clean shutdown cancels, then waits for every goroutine it started, then releases shared resources.

solid answer

~50 s

Cancelling the root `context.Context` closes its `Done` channel, and that is all it does: the cancel func returns immediately, and each goroutine reacts only the next time it selects on `ctx.Done()` or passes the context into a call that watches it. Meanwhile the process ends the moment `main` returns — the runtime does not wait for other goroutines, and their deferred calls never run. So shutdown has three beats, not one: broadcast the cancellation, wait until every subsystem reports that it has actually returned (a `sync.WaitGroup`, or one done channel per subsystem), and only then close shared resources such as a database handle or a log flusher. Skip the waiting beat and a graceful shutdown is just a slightly slower kill: writes are cut in half, buffered records vanish, and it looks random because it depends on scheduling.

code

go · 8 lines
go
func main() {
	ctx, cancel := context.WithCancel(context.Background())
	go worker(ctx) // returns when ctx.Done() is closed

	cancel()
	// main returns here: the process exits at once, and worker
	// may never even observe the cancellation.
}

go deeper

for a junior

Recall the two facts and say them plainly: cancelling a context is only a signal, and the process exits as soon as main returns. Then name the missing step — waiting for the goroutines you started.

for a middle

Explain the mechanism: the cancel func closes the Done channel and returns, goroutines see it only where they select on it, and the join is separate bookkeeping you own with a WaitGroup or done channels.

for a senior

Show where the ordering bites in production: releasing a shared handle or flusher before the join produces failures that depend on scheduling, and any blocking call in the shutdown path needs its own deadline because cancellation cannot reach it.

for a principal

Own the rule as a codebase convention: one root context, one place that joins, one place that releases, and every long-lived goroutine registered at the moment it is started so nobody can add a subsystem the shutdown does not know about.

## The thing people expect cancellation to do A Go service usually threads one root `context.Context` through everything it starts: the HTTP entry point, a queue consumer, a background flusher. When the process is told to stop, something cancels that root, and the intuition is that the program then "stops". It does not. Cancellation and stopping are two separate steps, and the whole discipline of an ordered shutdown exists to bridge them. ## What cancellation actually is `context.WithCancel` hands you a derived context and a cancel func. Calling that func closes the context's `Done` channel and sets its `Err`. That is the entire mechanism. It is a **broadcast**, delivered to anyone who bothers to look, and it returns instantly — it does not know who is watching, cannot count them, and never blocks. A goroutine observes it only where it chose to look: * in a `select` that has a `case <-ctx.Done():`, * inside a library call it passed the context to, which watches on its behalf, * or not at all, if it is sitting in a plain blocking call — a channel send, a mutex acquisition, a network read with no deadline — that knows nothing about contexts. This is the deepest consequence: **Go cannot kill a goroutine.** There is no `goroutine.Kill`, no interrupt, no abort. Cancellation is cooperative. A goroutine that ignores the context is unaffected by cancelling it, forever. ## What ends the process The second half of the surprise is the process boundary. When `main` returns, the program exits immediately. Other goroutines are not given a chance to finish, and their deferred calls — including a `defer flush()` or a `defer rows.Close()` — never run. `os.Exit` is blunter still: it skips deferred calls even in the goroutine that calls it. Put the two facts together and the failure mode is obvious. If shutdown is "cancel the root, then return from `main`", the cancellation and the death of the process race each other, and the process usually wins. Goroutines are cut off wherever they happened to be. ## The three beats A correct shutdown is: 1. **Signal.** Cancel the root context, which propagates to every derived context. Entry points stop accepting new work. 2. **Join.** Wait until each subsystem confirms it has returned. In practice each long-lived goroutine is registered with a `sync.WaitGroup` and does `defer wg.Done()`, and the shutdown path calls `wg.Wait()`; or each subsystem exposes a done channel that closes when its loop exits. 3. **Release.** Only after the join, close the shared things everybody was using: a database handle, an open file, a metrics or log flusher, a network listener. The join is what makes the third beat safe. Closing a flusher while a worker is still writing to it does not fail loudly and predictably; it fails at whatever moment scheduling picks, which is why these bugs reproduce only under load and only in production. ## What a good answer says about waiting Two shortcuts to reject explicitly: * **`time.Sleep` at the end of `main`.** This is a guess. If it is short you still cut work off; if it is long you have made every deployment slower for nothing. It also gives no signal when a subsystem is genuinely stuck. * **"cancel blocks until they finish".** It does not, and no context API does. `context.AfterFunc` runs a function *after* a context is done, which is still not a join of your goroutines. The join must be your own bookkeeping, because only your code knows how many goroutines it started. ## Where the trigger comes from In a real service the cancellation is triggered by the process being asked to terminate — the platform sends a signal, and the program turns that into one call to the root cancel func. The signal-handling machinery is a separate concern from the ordering; what matters here is that from that instant onwards, everything the process does is the three beats above, and the shape of the exit is decided by whether beat two exists. ## The shape to remember One root context, cancelled once. One place that waits. One place that releases, after the wait. If a question about shutdown ever seems to have two answers, it is usually because a subsystem was released before it was joined.

  • Does the cancel func guarantee the goroutine has stopped by the time it returns?
    No. It closes the context's `Done` channel and returns immediately. The goroutine may be mid-request, mid-write, or blocked in a call that never looks at a context. The only proof that it stopped is your own join — a `sync.WaitGroup` it decrements, or a done channel it closes as it returns.
  • What ends a Go process even while goroutines are still running?
    `main` returning ends it immediately: the runtime does not wait for other goroutines, and their deferred calls do not run. `os.Exit` is harsher — it skips deferred calls even in the goroutine that calls it. Both are why the wait has to happen before you reach the end of `main`.
  • Can you force a goroutine that ignores the context to stop?
    No. Go has no way to kill or interrupt a goroutine; cancellation is cooperative. Anything that ignores the context can only be abandoned, by exiting the process while it is still running. That is why blocking calls in a shutdown path need their own deadlines rather than relying on cancellation to reach them.

Cancelling the root context is announcing last orders. Everyone hears it, nobody has left yet, and locking the doors before the room is empty is how you lose people.

saying these in an interview costs you the question

  • Says the cancel func blocks until watching goroutines return
  • Adds a time.Sleep at the end of main to let goroutines finish
  • Thinks deferred calls in other goroutines run when main returns
  • Believes a goroutine can be killed from outside
  • Calls os.Exit as soon as the shutdown signal arrives