A worker's drain deadline fires with items still in flight. What should the process report, and how do you find what was still running?
answer
- the bad state is still live
- two branches must not log the same sentence
- count it so it becomes a rate
- dump every live stack before exiting
- bound the straggler, not the budget
basics
~20 sReport it loudly: log the abandoned item count, emit a metric, and exit with a status reserved for an incomplete drain. Before exiting, write the goroutine profile to standard error so the straggler's stack is captured.
solid answer
~50 sTwo jobs: make the outcome visible, and capture the evidence while it still exists. Visible means the deadline branch never logs the same thing as the clean branch — log the count and ideally the ids of the items being abandoned, increment a counter so "drains that ran out of budget" is a graphable rate, and return a distinct non-zero status so a fleet timing out on every deploy is visible. Evidence means dumping the goroutine profile right there, before the process exits: `pprof.Lookup("goroutine").WriteTo(os.Stderr, 1)` groups every live goroutine's stack with a count; debug level 2 dumps them individually. The straggler is usually obvious in it — parked on an outbound call with no deadline of its own, on a receive from a channel nobody will send to, or waiting for a lock the handler still holds. The fix is nearly always to bound that one operation, not to lengthen the drain budget.
code
go · 12 linesselect {
case <-workersDone:
log.Print("drain complete")
return nil
case <-drainCtx.Done():
if errors.Is(drainCtx.Err(), context.DeadlineExceeded) {
log.Printf("drain incomplete: %d items in flight, %d goroutines",
inflight.Load(), runtime.NumGoroutine())
_ = pprof.Lookup("goroutine").WriteTo(os.Stderr, 1)
}
return errDrainIncomplete
}go deeper
Know that the timeout branch of a shutdown needs its own log line and its own non-zero exit, and that the number of items still in flight is worth printing.
Explain how to capture live stacks from inside the process with the goroutine profile, what the two debug levels print, and why checking for context.DeadlineExceeded distinguishes a genuine timeout from a context that was cancelled by its parent.
Walk the full loop: capture the dump at the instant of the deadline, read the straggler's shape out of it, bound that operation rather than the budget, and then watch the incomplete-drain metric fall. Say why the profiling endpoint is usually already gone.
Decide what an incomplete drain means for the service — accepted loss to alert on by rate, or a correctness event that pages — and make the dump-plus-metric behaviour standard so every service in the estate produces the same evidence at the same moment.
## The deadline branch is a diagnosis point, not just an exit When the drain budget expires with work in flight, the process still has a few milliseconds of full control over a live program that is in exactly the bad state. Everything you want to know is currently in memory and will be gone shortly. Treat the branch as the moment to record it. ## What to report **A distinct log line.** The clean branch and the deadline branch must not produce the same message. Include the number of items still in flight, and their identifiers if you track them, so the queue side can be checked for redelivery. **A metric.** A counter incremented on the deadline branch turns "the drain sometimes runs out" into a rate you can put on a dashboard and alert on. Deploys stop being anecdotes: you can see that a change made drains slower. **A distinguishable exit status.** Return an error from the run function so the process exits non-zero, and reserve a status for "drain incomplete". A fleet where every instance exits with that status on every deploy is telling you the budget is fiction, and only a non-zero status makes that visible to whatever supervises the process. **Never the success message.** The single most common defect here is a shutdown routine that logs "shutdown complete" after the select regardless of which branch it took. ## Confirm which deadline actually fired Before blaming the budget, check the error: ```go if errors.Is(drainCtx.Err(), context.DeadlineExceeded) { // the budget really did run out } ``` If it reports `context.Canceled` instead, the budget did not expire at all — it was cancelled, almost always because it was derived from the already-cancelled shutdown context. That is a wiring bug rather than slow work, and it needs the opposite fix. ## Capturing what was still running The goroutine profile is the right instrument, because the question is literally "which goroutines are still alive and where are they parked": ```go _ = pprof.Lookup("goroutine").WriteTo(os.Stderr, 1) ``` - **debug level 1** prints each distinct stack once with a count of the goroutines sharing it. This is the readable form and usually enough: five hundred goroutines parked on the same line is an immediate answer. - **debug level 2** prints every goroutine individually with its full stack and its wait state, which is the form you want when the stragglers differ from each other. Writing to standard error means the dump lands wherever the process's output already goes, with no extra endpoint to reach and no dependency on a network that may already be draining away. If the service exposes profiling over HTTP, note that the listener is usually gone by this point in the shutdown — which is exactly why the in-process dump at the deadline is worth having. `runtime.NumGoroutine()` in the same log line is a cheap companion: a number in the log survives even if the dump is truncated. ## Reading the dump The stragglers almost always fall into a handful of shapes: - **An outbound call with no deadline of its own.** The stack sits in a read on a connection. The item cannot finish because the thing it is waiting for has no bound, and the drain budget was the only clock in the system. - **A receive that will never be satisfied.** A goroutine parked on a channel receive whose sender already exited during shutdown. - **A lock held by something else that is also stuck.** Two stacks, one holding, one waiting. - **A retry loop that ignores cancellation.** The stack is in a sleep between attempts, and the loop checks nothing. - **A goroutine leak that predates the shutdown.** The count in the dump is large and growing across deploys; the drain merely exposed it. ## The fix is usually not a bigger budget Each of those shapes has its own repair: give the outbound call a deadline, make the retry loop select on cancellation, close the channel or drop the receive, fix the lock ordering. Raising the drain budget only widens the window in which the same straggler still fails to finish, while making every deploy slower. Lengthening the budget is the right answer only when the dump shows ordinary work that simply needs a little more time — and even then, only if the platform's grace period leaves room. ## Closing the loop After the fix, the metric is the check: the rate of incomplete drains should go to roughly zero and stay there. If it does not, the dump is still the instrument, and the next deploy will hand you a fresh one at exactly the same instant.
- Why write the goroutine profile to standard error instead of scraping it over HTTP?By the time the drain deadline fires the service's listeners are usually already closed, so there is nothing left to scrape, and the moment of interest is gone microseconds later. An in-process write to standard error needs no network and lands in the same log stream the shutdown is already writing to.
- What is the difference between debug level 1 and 2 on the goroutine profile?Level 1 aggregates identical stacks and prints each with a count, which is compact and usually enough to spot a straggler. Level 2 prints every goroutine separately with its full stack and wait state, which is what you want when the stuck goroutines differ from one another.
- The dump shows a goroutine blocked reading an outbound response. What do you change?Give that call its own deadline so it cannot outlive the item it belongs to, rather than extending the drain budget. A per-call bound fixes the straggler in normal operation too, not just at shutdown, and keeps the drain budget a bound on ordinary work instead of a substitute for missing timeouts.
- Should an incomplete drain page someone?That depends on what an abandoned item costs. If the queue redelivers and handlers are idempotent, an occasional incomplete drain is an accepted loss to alert on by rate. If abandoning an item means a duplicate side effect or a silent drop, it is a correctness event and belongs on a page.
saying these in an interview costs you the question
- Logs shutdown complete on both the clean and deadline branches
- Raises the drain budget without looking at what was stuck
- Exits zero after abandoning in-flight work
- Tries to scrape a profiling endpoint that shutdown already closed
- Never distinguishes DeadlineExceeded from an inherited cancel
- Treats a growing goroutine count across deploys as normal