skip to content

In a Go service, why must Submit stop accepting new records before shutdown calls sync.WaitGroup.Wait?

level: juniorimportance: must knowfreq 60%

answer

  1. shutdown is two phases, not one
  2. a quiet moment is not an empty service
  3. Wait returns the instant the counter is zero
  4. refuse first so zero means finished
  5. count on the accept side, not in the worker

basics

~20 s

A sync.WaitGroup's Wait returns the moment its counter reaches zero, which can happen in a lull between two submissions. Refusing new work first makes the counter only fall, so Wait returning really means nothing is left in flight.

solid answer

~50 s

Draining is two phases and the order matters. First flip a flag that makes the accept path refuse and return an error to callers; only then wait for what was already accepted. `sync.WaitGroup.Wait` blocks until the counter is zero and returns the instant it gets there, so while submissions keep arriving the counter can dip to zero during a quiet moment, `Wait` returns, and the next record is accepted by a service that has already declared itself drained — that record dies with the process. Concretely: `Submit` checks the closing flag, calls `wg.Add(1)` before it reports the record accepted, and the worker runs `defer wg.Done()`. `Shutdown` sets the flag, then calls `wg.Wait()`. The `sync` package also requires that an `Add` lifting the counter off zero happen before a `Wait`, so late admissions are misuse, not merely a lost record.

code

go · 21 lines
go
type Writer struct {
	closing atomic.Bool
	wg      sync.WaitGroup
}

func (w *Writer) Submit(r Record) error {
	if w.closing.Load() {
		return errors.New("audit writer is draining")
	}
	w.wg.Add(1) // counted before the caller is told "accepted"
	go func() {
		defer w.wg.Done()
		w.write(r)
	}()
	return nil
}

func (w *Writer) Shutdown() {
	w.closing.Store(true) // phase 1: refuse new records
	w.wg.Wait()           // phase 2: wait for accepted ones
}

go deeper

for a junior

Be ready to say the two steps out loud in order: stop accepting, then wait for what was accepted. Know that Wait returns the moment the counter hits zero and takes no timeout.

for a middle

Explain why the counting must happen where work is accepted rather than inside the goroutine, and what the accept path returns to callers during the drain. Mention that raising the counter off zero while a Wait is running is documented misuse.

for a senior

Show that you know what the counter does not cover: queued or buffered items nobody has started yet. Talk about flushing those after the wait and about reporting an incomplete drain rather than logging a clean exit.

for a principal

Frame the drain as a contract with callers and with the platform: what an in-flight record is promised, what an error from the accept path means to the caller, and how much of the termination budget the service is allowed to spend before it is killed.

## What "in flight" means A long-running Go service — say an audit-log writer that other components hand records to — has three populations of work at any instant: records nobody has submitted yet, records that have been *accepted* and are being written, and records already durably written. Draining is the act of getting the middle population to zero before the process exits, without letting the first population grow into it. Go's usual counter for the middle population is a `sync.WaitGroup`. It holds an integer. `Add(n)` raises it, `Done()` lowers it by one, and `Wait()` blocks the calling goroutine until the counter is zero. That is the entire contract, and two properties of it drive this whole question: 1. `Wait` returns **as soon as** the counter reaches zero. It has no memory of what came before and no opinion about what comes after. 2. `Wait` takes no argument and returns nothing. There is no "wait for whatever is submitted in the next second" mode. ## Why waiting alone is not draining Suppose shutdown just calls `wg.Wait()` while `Submit` keeps admitting records. The service handles bursty traffic, so it is normal for every worker to finish and the counter to sit at zero for a few microseconds. `Wait` sees that zero and returns. Shutdown moves on: it closes the file handle, flushes nothing, calls `os.Exit`. Meanwhile another component calls `Submit`, the counter goes 0 → 1, a worker starts — and the process disappears underneath it. The record is gone, and the service reported a clean shutdown. The drop is the visible symptom. There is a second, uglier problem: raising a `WaitGroup` counter from zero while some goroutine is inside `Wait` is documented misuse. The `sync` package requires that any `Add` which takes the counter off zero *happen before* a `Wait` on that group, precisely so `Wait` can be implemented without racing the reuse case. Violate that and the behaviour is undefined; `sync` detects some occurrences and panics reporting WaitGroup misuse. So "just call `Wait`" is wrong twice: it loses work, and it uses the primitive outside its contract. ## The two-phase shape The fix is to make the counter monotonically non-increasing before anyone waits on it: **Phase 1 — refuse.** Set a flag the accept path consults. The accept path is whatever function admits work: an exported `Submit`, a handler, a loop reading a queue. Once the flag is set, `Submit` returns an error instead of admitting anything, so callers learn immediately that the service is going away and can retry elsewhere, block, or write the record to disk themselves. Silently returning `nil` while dropping the record is the worst possible outcome, because it converts a shutdown problem into a data-integrity problem that shows up in an audit years later. **Phase 2 — wait.** Now `wg.Wait()` is meaningful: nothing can raise the counter, so reaching zero is a terminal state, not a lull. ## Where the Add goes The counting must happen on the accept side, not inside the worker. If `Submit` starts a goroutine and that goroutine calls `Add(1)` as its first statement, there is a window where the record is accepted but uncounted: `Wait` can return while the goroutine has not yet been scheduled. The `go` statement guarantees the goroutine starts, not that it starts *soon*. So the rule is: `Add` before the record is reported accepted, `Done` in the work itself (via `defer`, so an early return or a recovered failure still decrements it). Then "counter zero" and "nothing accepted is unfinished" mean the same thing. ## What this shape still does not cover Two gaps are worth naming so the design is not oversold. First, the check-the-flag-then-`Add` pair is not itself atomic; a submission can slip between the two steps unless they are guarded together. Second, a `WaitGroup` counts only work that has been *started and counted*. Records sitting in an in-memory batch or in a channel's buffer, waiting for a worker to pick them up, are invisible to it — the counter can reach zero with the batch still unflushed. Those need their own flush after `Wait` returns, and their own counter if you want to report honestly on what was drained. ## The interview shape of the answer Name the two phases, name their order, say *why* the order matters in terms of `Wait`'s exact semantics (returns at zero, no memory), and say what `Submit` returns to callers during the drain. That is a complete answer at this level.

  • What should the accept path return to its caller once the service has started draining?
    An explicit error saying the service is shutting down. The caller can then retry against another instance, block, or persist the record itself. Returning nil while quietly discarding the record turns a shutdown detail into silent data loss, which is exactly what this design exists to prevent.
  • Why call wg.Add(1) in Submit rather than as the first line of the worker goroutine?
    The go statement guarantees the goroutine will run, not that it runs promptly. If the Add lives inside it, there is a window where the record is accepted but the counter has not moved, so Wait can see zero and return while that record is still unwritten. Counting on the accept side closes the window.
  • Once wg.Wait returns, is every record safely written?
    Only the ones the group counted. Records still sitting in an in-memory batch or a channel buffer were never added to the counter, so the drain has to flush that buffer after Wait returns and before releasing the output. A WaitGroup at zero means "nothing started is unfinished", not "nothing is pending".

saying these in an interview costs you the question

  • Calling wg.Wait alone is a graceful shutdown
  • New work may keep arriving while the drain runs
  • Wait blocks until every goroutine ever started exits
  • Silently dropping late submissions is acceptable
  • Adding to the counter inside the worker goroutine is equivalent