A Go service sometimes never exits after shutdown begins. How do you find the stuck subsystem?
answer
- the process can describe its own hang
- near-zero CPU means parked, not spinning
- arm a deadline the moment teardown starts
- dump every goroutine, not just this one
- read the parked states: chan send, semacquire
basics
~20 sArm a watchdog when teardown starts: on its deadline, dump every goroutine's stack with runtime.Stack(buf, true) and exit non-zero. The stacks name the blocked line, usually a send nobody receives or a call with no deadline.
solid answer
~50 sMake the process report on itself instead of guessing. The moment shutdown begins, start a watchdog goroutine that selects between a done channel and a deadline; if the deadline wins, call `runtime.Stack(buf, true)` (or `pprof.Lookup("goroutine").WriteTo(os.Stderr, 2)`), write the dump, and `os.Exit` with a non-zero code so the exit is yours and attributable rather than a silent kill. Then read the dump: goroutines parked in `chan send`, `chan receive`, `select` or `semacquire` are the candidates, and the top frames name the exact line. The usual causes are a teardown run in the wrong order — a producer blocked sending into a channel whose consumer has already returned — a blocking call with no deadline that cancellation cannot reach, or a subsystem waiting on a context that was never derived from the root. For a process already hung, sending it SIGQUIT makes the runtime print goroutine stacks and exit.
code
go · 15 lines// armed the moment shutdown starts, before any teardown runs
done := make(chan struct{})
go func() {
select {
case <-done:
case <-time.After(20 * time.Second):
buf := make([]byte, 1<<20)
n := runtime.Stack(buf, true) // all goroutines, blocked included
os.Stderr.Write(buf[:n])
os.Exit(1)
}
}()
teardown() // cancel, join each stage, release
close(done)go deeper
Know that a hung Go process can print where every goroutine is parked, and that a stack dump beats guessing. Recognise near-zero CPU as blocked rather than looping.
Explain the mechanics: a watchdog selecting between a done channel and a deadline, runtime.Stack with all set to true, and what parked states like chan send or semacquire mean in the dump.
Demonstrate the whole loop under pressure — instrument, capture, read the dump, name the cause as an ordering error or a missing deadline, then fix the sequence and keep the watchdog as a permanent guard.
Own the policy: every service bounds its own teardown, dumps on the deadline and exits with a code that says so, so a slow shutdown surfaces as a stack trace in the logs instead of as a fleet-wide complaint that rollouts are slow.
## The symptom Shutdown starts and the process does not die. CPU is near zero — it is not spinning, it is *parked* — and eventually something external kills it. Sometimes. On some deploys. That intermittency is the tell: the hang depends on what happened to be in flight when the trigger arrived, so it is invisible in tests and shows up as "deployments are slow" long before anyone calls it a bug. ## Instrument, do not guess A hung Go process is one of the friendliest things to debug, because the runtime can print exactly where every goroutine is parked. Two ways in: * **From inside, on a deadline.** When teardown begins, start a watchdog goroutine that waits on either "teardown finished" or a timer. If the timer wins, dump and exit. * **From outside, on an already-hung process.** Sending SIGQUIT makes the Go runtime print goroutine stacks and terminate. `GOTRACEBACK=all` in the environment widens what a crash prints to all goroutines. The in-process watchdog is what you want in a service, because the hang happens on a machine nobody is logged into, and by the time a human looks the process is gone. ### What to dump `runtime.Stack(buf, true)` fills a byte slice with the stacks of **every** goroutine, blocked ones included, and returns the number of bytes written. Size the buffer generously — a megabyte — because a truncated dump usually cuts off the goroutine you needed. `pprof.Lookup("goroutine").WriteTo(w, 2)` gives the same information in the same human-readable form, plus the count. This is the goroutine dump specifically. A CPU profile of an idle hung process samples nothing; a heap profile tells you about allocation, not about who is blocked. Wrong instrument, no answer. ### Exit deliberately After dumping, exit non-zero yourself. Two reasons. You get the dump written and flushed at a moment you chose, and the exit is attributable — an exit code you own says "my shutdown missed its deadline", while being killed externally looks identical to a crash, an OOM, or a healthy-but-slow stop. ## Reading the dump Each goroutine prints its state and its stack. The states that matter: * **`chan send`** — parked trying to send. Almost always the ordering bug: the consumer for that channel has already returned, so nothing will ever take the value. Teardown ran back to front. * **`chan receive`** — parked waiting for input on a channel whose sending side never closed it. Something upstream returned without closing what it owned. * **`select`** — waiting on several cases, none ready. Look at which cases: if a `ctx.Done()` case is present, this goroutine's context is not the one you cancelled. * **`semacquire`** — waiting on a mutex or a `WaitGroup`. A `Wait` that never returns means a counter that was never decremented, often a goroutine that returned early on a path without its `Done`. * **`IO wait`** — a network or file call in progress. Cancellation does not reach these; only a deadline does. One of these will be the subsystem in the teardown sequence; the rest are usually its victims waiting behind it. ## The recurring causes **Teardown in the wrong order.** The most common by far: a stage was joined and released before the stage feeding it, so a producer is stuck mid-send. The dump makes this obvious because the stuck goroutine's stack names the channel's send site. **A blocking call with no deadline.** An outbound request, a database query, a file read on a slow mount. Cancellation is cooperative; a call that ignores the context cannot be shortened, only abandoned. The fix is a deadline on the call, not more waiting. **The wrong context.** A subsystem constructed with its own background context, or one derived before the root was created. It watches a `Done` channel that nobody ever closes, so it waits forever, politely. **A join that can never complete.** A wait on a counter that a goroutine failed to decrement on an error path — the goroutine is long gone, and the wait outlives it. ## Turning the finding into a fix The dump tells you *where*; the fix is almost always *ordering* or *deadlines*. Move the stage's join ahead of the release it was fighting with; put a deadline on the blocking call; derive the subsystem's context from the root. Keep the watchdog afterwards — it is cheap, it turns a future regression from "deployments got slow again" into a stack trace in the logs, and it puts an upper bound on how long the process can take to die. ## What not to do Do not lengthen a sleep until the symptom goes away; you have hidden a hang behind a slower deploy. Do not treat an external kill as the shutdown mechanism — anything the process was going to flush is lost, and you learn nothing about why. And do not reach for the race detector: it finds unsynchronised access, not blocked goroutines, and a hung shutdown is not a data race.
- The dump shows a goroutine parked in 'chan send'. What does that tell you?That the receiving side for that channel is already gone, so the value can never be taken — teardown ran back to front. The stack names the send site, which names the stage that should have been joined later. The fix is ordering: join and close along the direction data flows, not against it.
- Why exit non-zero yourself instead of letting the platform kill the process?Because you control the moment, so the dump is written and flushed, and the exit is attributable: your own code says shutdown missed its deadline. An external kill produces no evidence and is indistinguishable from a crash or an out-of-memory kill in the logs.
- How do you keep the watchdog from firing during a slow but healthy drain?Set its deadline comfortably above whatever budget the service already enforces on its own drain, so the watchdog only ever fires on a genuine hang. Log which teardown step was still outstanding when it fires, so the dump arrives with the sequence position attached rather than as an anonymous wall of stacks.
saying these in an interview costs you the question
- Lengthens a sleep until the symptom disappears
- Reaches for a heap or CPU profile to find a blocked goroutine
- Says the race detector will find a shutdown hang
- Treats the external kill as the shutdown mechanism
- Restarts and files it as a flake
- Assumes cancellation can interrupt a blocking network call