Your Go report service's Server.Shutdown never returns and deploys hang. How do you diagnose it?
answer
- an unbounded context is the enabler
- one connection can never go idle
- dump every stack while it is stuck
- your app's port is already closed by then
- look for the lock holder, not the waiter
basics
~20 sShutdown returns only when every tracked connection has gone idle, so an unbounded Shutdown blocks forever on one handler that never returns. Capture the goroutine profile with full stacks while it hangs, then bound the context and escalate to Close.
solid answer
~50 s`Shutdown` waits for each active connection to finish its request; with `context.Background()` it waits forever, so one handler that never returns hangs every deploy. Diagnose it by taking the goroutine profile with full stacks while the process is stuck — `pprof.Lookup("goroutine").WriteTo(w, 2)` from the shutdown path, or `/debug/pprof/goroutine?debug=2` if you have somewhere to scrape it from. The catch is that the profiling endpoint must be on a *second* `http.Server` on its own port: `Shutdown` closed the main listener as its first act, so scraping the app's own port is impossible by then. Read the dump for handler goroutines parked in `ServeHTTP` frames — typically a channel receive with no sender left, a mutex held by another stuck goroutine, or an outbound call made without a deadline. Then fix both halves: give the handler a bound, and always pass a deadline to `Shutdown` and call `Close` when it expires.
code
go · 9 linesctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := srv.Shutdown(ctx); err != nil {
log.Printf("drain abandoned: %v", err)
// debug=2 prints readable stacks for every goroutine
_ = pprof.Lookup("goroutine").WriteTo(os.Stderr, 2)
_ = srv.Close()
}go deeper
Recall that Shutdown with an unbounded context waits forever, and that passing a deadline is the baseline habit that stops a hung deploy.
Explain why one stuck handler blocks the whole drain, and know that the goroutine profile at debug level 2 prints readable stacks for every goroutine.
Run the whole diagnosis: capture the dump while it hangs, identify the parked handler and the lock holder behind it, and fix the handler as well as the shutdown path.
Make the failure self-reporting across services — diagnostics on their own listener, the abandoned-drain error logged and counted — so the next occurrence is diagnosed from artefacts rather than reproduced by hand.
## The symptom and the mechanism Deploys of a report-download service go out several times a day. One day they start taking minutes each; the pod or process sits there after being asked to stop and eventually gets killed from outside. The shutdown code looks correct: ```go _ = srv.Shutdown(context.Background()) ``` and that line is the whole bug's enabler. `Shutdown` returns when every connection the server tracks has gone idle, or when its context ends. `context.Background()` never ends. So a single handler that never returns pins one connection active forever, and `Shutdown` waits forever with it — a deadlock between the drain and the request it is draining. Nothing in the process logs anything, because from the runtime's point of view nobody is doing anything wrong: one goroutine is parked, one goroutine is polling, all is well. ## Getting the evidence The goroutine profile with full stacks is the right instrument, because it shows every goroutine and precisely where it is parked, including how long it has been blocked. There are two ways to get it, and the choice matters here: **From inside the shutdown path.** This always works, and it is what you should add to the service permanently: ```go ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second) defer cancel() if err := srv.Shutdown(ctx); err != nil { log.Printf("drain abandoned: %v", err) _ = pprof.Lookup("goroutine").WriteTo(os.Stderr, 2) _ = srv.Close() } ``` The `2` is the debug level that prints human-readable stacks for every goroutine rather than a compressed pprof payload. Now a hung deploy leaves the evidence in the logs by itself, which is the difference between diagnosing this once and diagnosing it every time. **By scraping the profile endpoint.** This is where people get caught. If your pprof handlers are registered on the same `http.Server` that is draining, that listener was closed in `Shutdown`'s very first step — the port is gone before you can reach it. Serving diagnostics on a **separate `http.Server` on its own port**, shut down last (or not at all), is the fix, and it is worth doing for reasons well beyond this bug. ## Reading the dump Scan for goroutines whose stack contains your handler and `net/http.(*conn).serve`. Those are the connections keeping the drain open. What is at the top of each stack tells you the cause: - **`chan receive` with no producer** — a handler waiting on a result channel whose producer already returned or panicked. - **`sync.Mutex.Lock` / `RWMutex.RLock`** — the handler is fine, and the goroutine *holding* that lock is the real culprit; find it in the same dump. - **`sync.WaitGroup.Wait`** — a handler joining work that will not finish. - **`net/http.(*persistConn).roundTrip` or a socket read** — an outbound call with no deadline; an upstream that never answers becomes your unbounded drain. - **A long transfer legitimately in progress** — not a bug at all, just a request slower than your patience, which changes the answer from "fix the deadlock" to "choose a deadline". The header on each goroutine also gives its blocked-for duration, which separates "stuck for 40 minutes" from "started two seconds ago". ## The fixes, in the order you apply them 1. **Bound the drain.** Never pass `context.Background()` to `Shutdown`. A deadline turns an indefinite hang into a bounded delay with a logged error. 2. **Escalate.** On the deadline error, call `Close`. Without it, the connections stay open and the process lingers until something outside kills it — the deadline alone changes nothing about the connections. 3. **Bound the handler.** The parked goroutine is a bug in its own right: an unbounded wait in a handler will also leak goroutines under normal load, not just at shutdown. Give outbound calls request-scoped deadlines, give internal waits a `select` with a timeout, and make sure whatever the handler waits on cannot silently disappear. 4. **Make the next occurrence self-diagnosing.** Log the abandoned-drain error with a counter, and dump the goroutine profile on that path as shown above. ## What to say about "but the request was still legitimate" Sometimes the dump shows no deadlock at all — just a 300 MB download that is genuinely still streaming. That is a different conversation: your drain is being held open by real work, and the decision is how long that work is worth delaying a deploy, not how to fix a bug. Distinguishing those two outcomes from the same dump is the point of taking it.
- Why can you not just scrape /debug/pprof/goroutine from the app's own port while it hangs?Because `Shutdown` closes the listeners before it starts waiting, so that port stopped accepting connections at the very beginning of the drain. Diagnostics have to live on a second `http.Server` with its own listener, torn down after the main one, or be dumped from inside the process.
- The dump shows a handler blocked in sync.Mutex.Lock. Where is the bug?Not in that handler. Someone else holds the mutex and is not releasing it — find that goroutine in the same dump and read its stack. A blocked waiter is a symptom; the holder's stack is the diagnosis, and it is often parked on an outbound call with no deadline.
- Does adding a deadline to Shutdown fix the hang on its own?It fixes the hang, not the leak. `Shutdown` returns after the deadline but leaves those connections open and their handlers running, so you still need the `Close` escalation to end them, and you still need to fix whatever made the handler unable to return.
- How would you tell a deadlock apart from a genuinely slow download in the same dump?Read what the goroutine is parked on and for how long. A transfer in progress shows a write on the connection and advances between dumps; a deadlock shows a channel receive, a lock, or a WaitGroup, and its blocked-for duration keeps growing while nothing moves.
saying these in an interview costs you the question
- Passes context.Background() to Shutdown in production
- Tries to scrape pprof on the port that just closed
- Blames the goroutine blocked on the lock rather than the holder
- Adds a deadline but never escalates to Close
- Concludes the runtime is broken because nothing is logged