How do you put a time bound on sync.WaitGroup.Wait when draining a Go service at shutdown?
answer
- Wait takes no deadline and returns nothing
- turn finishing into a channel event
- close a channel, then select on it
- the losing branch leaves a goroutine parked
- expiring the wait stops nothing
basics
~20 sWait takes no deadline, so run it in its own goroutine that closes a done channel when it returns, then select over that channel and a deadline channel such as ctx.Done() or time.After. The winning branch tells you whether the drain finished.
solid answer
~50 s`sync.WaitGroup.Wait` has no timeout parameter and no return value, so the bound has to come from a `select`. Bridge the wait onto a channel: start a goroutine that calls `wg.Wait()` and then `close(done)`, and `select` over `<-done` and a deadline channel — `ctx.Done()` from a `context.WithTimeout`, or `<-time.After(d)` if there is no context to carry. If `done` wins, the drain completed and you can flush and release resources; if the deadline wins, return an error such as `ctx.Err()` so the caller and the logs record an incomplete drain. Two caveats: the bridging goroutine stays parked in `Wait` when the deadline wins, which is harmless at process exit but a leak in a long-lived process or a test; and the expiry is not cancellation — the workers keep running unless you also cancel the context they were given.
code
go · 15 linesfunc (w *Writer) drain(ctx context.Context) error {
done := make(chan struct{})
go func() {
w.wg.Wait()
close(done)
}()
select {
case <-done:
return nil // every counted record finished
case <-ctx.Done():
// stragglers are still running; do not close the output yet
return ctx.Err()
}
}go deeper
Remember that Wait has no timeout parameter, and that the way to bound any blocking operation in Go is to turn it into a channel receive and select over it alongside a deadline channel.
Write the bridging goroutine and the select from memory, explain why the channel is closed rather than sent on, and say what state the process is in when each branch wins.
Talk about the branch nobody writes: what must not be released while stragglers run, what gets logged, what error propagates, and how you cancel the workers rather than merely stopping the wait.
Own where the number comes from. The drain budget has to fit inside the platform's termination allowance with room for the final flush, and it should be configurable rather than a constant buried in the shutdown path.
## Why a bound is needed at all An unbounded drain hands control of your shutdown to whatever is slowest in flight. One record whose write is stuck on a wedged network connection keeps `wg.Wait()` blocked forever, and the process sits there until something external kills it — usually with a hard signal that gives you no chance to flush, log or report. A bounded drain converts that into a decision you make: wait up to N seconds, then proceed with whatever state you have and *say so*. ## Wait has no bound to give The method is `func (wg *WaitGroup) Wait()`. No argument, no return, no context. That is a deliberate design: `WaitGroup` is the minimal counter, and composing it with time is left to `select`, which is where every timing decision in Go lives. So the bound has to be built from the outside, and the standard construction is to turn "the wait finished" into a channel event: 1. Make a `done := make(chan struct{})`. 2. Start one goroutine whose entire body is `wg.Wait()` followed by `close(done)`. 3. `select` over `<-done` and a deadline channel. Closing is the right signal rather than sending, because a close is broadcast and never blocks — the bridging goroutine does not care whether anyone is still listening when the deadline has already won the race. ## Which deadline channel If the shutdown path already receives a `context.Context` — and it usually should, because whoever calls it knows the overall budget — use `ctx.Done()` and return `ctx.Err()` when it fires. That keeps one deadline for the whole teardown rather than a private timer per component, and it lets a caller cancel the drain early if the situation changes. If there is no context, `case <-time.After(d):` is the compact form. It allocates a timer that is not stopped when the other branch wins; for a once-per-process shutdown that is irrelevant, and it would only matter in a hot loop. ## What each branch means **`done` won.** Every counted item finished. This is the branch where you flush buffers, close the output, and return `nil`. **The deadline won.** Some items are still running. Three things must happen here, and skipping any of them is the common defect: - Return a non-nil error. A drain that expired is not a clean shutdown, and the process's exit status and logs should be able to say so. - Do not release resources the still-running workers are using. If you close the output file the moment the deadline fires, the stragglers panic or write into a closed handle instead of merely being abandoned. - Record how much was left. A count of outstanding items in the shutdown log is the difference between "we lost some records last Tuesday" and "we lost 135 records last Tuesday". ## The parked goroutine When the deadline wins, the bridging goroutine is still blocked inside `Wait`, and it stays blocked until the counter reaches zero — possibly forever. At process exit that costs nothing. Two situations where it does cost something: a long-lived process that drains a *subsystem* repeatedly rather than the whole program, where those goroutines accumulate; and tests, where a leak checker or the goroutine profile will show one parked waiter per expired drain. The honest description in a review is "this goroutine outlives the deadline by design, and here is why that is acceptable here". ## A bound is not cancellation This is the point most often missed. Giving up on the wait changes nothing about the workers — they are still running, still holding connections, still allocating. If you want the in-flight work to actually stop, the work itself must be watching a cancellable context that you cancel, and the operations inside it must be context-aware. A common composition is: cancel the workers' context at the deadline, then wait a short second window for them to unwind, and only then give up. Without that, a "5 second drain timeout" only bounds how long *shutdown* waits, not how long the process keeps working. ## Where the number comes from The drain budget is not a free choice: it has to fit inside whatever external deadline the process runs under, with room left over for the flush and for writing the shutdown log. Choosing it is a separate discussion from building the mechanism, but the mechanism should take the budget as a parameter — a hard-coded constant deep inside the shutdown function is the version that nobody can tune when the platform's grace period changes.
- What happens to the goroutine that was calling Wait when the deadline branch wins?It stays blocked inside Wait until the counter reaches zero, which may be never. At process exit that is free. In a long-lived process that drains a subsystem repeatedly, or in tests with a leak check, those parked waiters accumulate and show up in the goroutine profile, so the leak has to be acknowledged rather than ignored.
- Does the deadline firing stop the in-flight work?No. Giving up on the wait only ends the waiting; the workers keep running, holding connections and allocating. Real cancellation requires the work itself to watch a cancellable context that you cancel, with context-aware operations inside. A common shape is to cancel at the deadline, then allow a short second window for the workers to unwind.
- What should the shutdown function return when the drain does not finish in time?A non-nil error, typically the context's error, plus a logged count of what was still outstanding. An expired drain is not a clean shutdown, and swallowing the error means the exit status and the logs claim success while records were abandoned.
saying these in an interview costs you the question
- Passing a timeout argument to Wait
- Closing the output the moment the deadline fires
- Treating an expired drain as a successful shutdown
- Assuming the deadline cancels the running work
- Sending on the done channel instead of closing it