During shutdown your image service panics writing to a blob store main already closed, so how should main order shutdown?
answer
- a collaborator must outlive its users
- built leaves-first, closed leaves-last
- deferred calls already run backwards
- drain before you close, with a fresh deadline
basics
~20 sClose in reverse construction order, and only once the work using a collaborator has stopped: stop accepting requests, drain in-flight handlers with http.Server.Shutdown, then close the blob store. Deferred calls already unwind in that order.
solid answer
~50 sThe panic says the closing was ordered by convenience rather than by dependency. The graph was built leaves-first — store, then thumbnailer, then server — so it must come down leaves-last: first stop taking new work, then wait for the work already running, and only then close what that work was using. Concretely, `signal.NotifyContext` turns SIGINT and SIGTERM into a cancelled context; on cancellation call `srv.Shutdown(shutCtx)` with a *fresh* context carrying a deadline, because `Shutdown` returns once active handlers have returned. `defer store.Close()` was registered before the server existed, so it runs after `Shutdown` — the correct order, for free. A background worker that is not behind the server needs the same treatment: cancel its context and wait for it to signal that it has returned before closing anything it writes to. If the whole shutdown hangs instead, `GOTRACEBACK=all` plus a SIGQUIT prints every goroutine's stack and shows which one is still parked.
code
go · 22 linesctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
store, err := NewBlobStore("/var/lib/thumbs")
if err != nil {
return err
}
defer store.Close() // registered first, runs last
srv := &http.Server{Addr: ":8080", Handler: NewThumbnailer(store, cache, 4096)}
errc := make(chan error, 1)
go func() { errc <- srv.ListenAndServe() }()
select {
case err := <-errc:
return err
case <-ctx.Done():
}
shutCtx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
defer cancel()
return srv.Shutdown(shutCtx) // returns once in-flight handlers have finishedgo deeper
Know that whoever creates a collaborator is the one that closes it, and that deferred calls run in reverse order of registration, which is why cleanup is deferred right after construction.
Explain what http.Server.Shutdown actually waits for, why the drain context must be a fresh one with a deadline, and how signal.NotifyContext turns SIGTERM into cancellation.
Walk the whole sequence for a real process — signal, stop intake, drain, wait for non-HTTP workers, unwind — and say how you diagnose a hang with a full goroutine stack dump before the supervisor kills you.
Own the shutdown budget as a contract with the platform: how long the process may take, what work it is allowed to abandon, and what the service promises to have flushed before it exits.
## Why the panic happens Startup builds a chain: the blob store exists first, the thumbnailer is built on it, the HTTP server is built on the thumbnailer. Every arrow means "uses". Shutdown has to walk those arrows backwards. Closing the store while a handler is mid-write breaks the invariant that a collaborator outlives every user of it — and the symptom is exactly what you saw: a write to a closed handle, in a handler, during the last second of the process's life. The rule is one sentence: **stop the producers of work, drain the work in flight, then close the resources that work uses — reverse construction order.** ## The shape that gets it right ```go func run() error { ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM) defer stop() store, err := NewBlobStore("/var/lib/thumbs") if err != nil { return err } defer store.Close() // registered first, so it runs last srv := &http.Server{Addr: ":8080", Handler: NewThumbnailer(store, nil, 4096)} errc := make(chan error, 1) go func() { errc <- srv.ListenAndServe() }() select { case err := <-errc: return err case <-ctx.Done(): } shutCtx, cancel := context.WithTimeout(context.Background(), 20*time.Second) defer cancel() return srv.Shutdown(shutCtx) } ``` Four things are load-bearing. **`defer` already gives you reverse order.** Deferred calls run last-in-first-out, and you register each cleanup immediately after its construction succeeded. So the order you close in is derived from the order you built in, automatically, and a partial startup failure only unwinds what actually got created. This is the main reason the wiring lives in a `run` function rather than in `main` behind a `log.Fatal`: `os.Exit` runs no deferred functions at all, so any cleanup you wrote would simply be skipped. **`signal.NotifyContext` is the entry point of shutdown.** It returns a context cancelled on the first matching signal, plus a `stop` function that unregisters the handler. Cancelling that context is the single "we are going down" event; every long-lived part of the program can select on it. **`http.Server.Shutdown` is what "drain" means.** It closes the listeners so no new connections are accepted, closes idle keep-alive connections, and then waits for active requests to return before it returns. That is exactly the wait you need before the store closes. Note what it does *not* do: it does not cancel the request contexts of handlers already running, so a handler that ignores deadlines can hold shutdown open for as long as it likes — which is why you give `Shutdown` a bounded context. **The shutdown context must be a fresh one.** Passing the already-cancelled signal context to `Shutdown` makes it return immediately with that context's error and abandon in-flight requests — the same class of bug in a different disguise. Create `context.WithTimeout(context.Background(), …)` with a budget slightly under whatever grace period your supervisor gives the process between SIGTERM and SIGKILL. ## The parts that are not behind the server A thumbnailing service usually also has something not driven by HTTP: a worker draining a queue of images, a periodic sweeper of the store. Those are the ones that produce the panic you asked about, because nothing waits for them by default. Give each the same shutdown context, and give `run` a way to wait: ```go done := make(chan struct{}) go func() { defer close(done) worker.Loop(ctx) // returns when ctx is cancelled }() // after Shutdown: <-done ``` Only once both the server and the worker have returned is it safe for `store.Close()` to run. If the worker holds a large `[]byte` image mid-encode, "safe" means it finished or abandoned that unit of work — which it can only do if it actually checks the context between units. ## When shutdown hangs instead of panicking The sibling failure is a process that receives SIGTERM and never exits, and gets SIGKILLed by its supervisor. The cheapest diagnostic is the runtime's own: with `GOTRACEBACK=all` set, sending SIGQUIT (Ctrl-\\ at a terminal) makes the runtime print every goroutine's stack and exit. Read it top-down and you almost always find one of three things: a handler blocked on a collaborator with no deadline, so `Shutdown` is still waiting; a worker goroutine parked on a channel receive because it never selects on the cancellation; or `Close` itself blocked on flushing to something unreachable. Each has a different fix, and the stack dump tells you which one you have instead of leaving you to guess. ## Ownership, stated once Whoever constructed a collaborator closes it. A type that was handed a store never closes that store, because it cannot know whether someone else is still using it — the wiring function can. Keeping that rule means the shutdown order is always readable in one function, next to the construction order it mirrors.
- Shutdown never completes and the supervisor kills the process. How do you find where it is stuck?Set `GOTRACEBACK=all` and send the process SIGQUIT: the runtime prints every goroutine's stack and exits. The dump usually names the culprit directly — a handler blocked on a call with no deadline so the drain never finishes, a worker parked on a channel receive because it never selects on the cancellation, or a Close flushing to something unreachable.
- Why not pass the already-cancelled signal context straight to http.Server.Shutdown?Because Shutdown honours that context: given one that is already done, it returns almost immediately with the context's error and stops waiting, abandoning requests that are still running. Build a fresh `context.WithTimeout` instead, with a budget a little under the grace period between SIGTERM and SIGKILL, so you drain what you can and still exit before you are killed.
- What about a background worker that is not behind the HTTP server?It needs the same discipline explicitly: hand it the cancellation context, have it check that context between units of work, and have the wiring function wait for it — typically a `done` channel the goroutine closes as it returns — before any collaborator it writes to is closed. Skipping that wait is precisely how a write lands on a closed store.
saying these in an interview costs you the question
- Closes collaborators in construction order instead of reverse
- Closes the store first so in-flight work fails fast
- Reuses the cancelled signal context as the drain deadline
- Assumes Shutdown cancels the contexts of running handlers
- Puts cleanup after a log.Fatal or os.Exit, where it never runs
- Lets a type close a collaborator it was handed rather than constructed