A `go worker()` call returns no value, so how does a supervisor goroutine learn that the worker exited and why?
answer
- a go statement is not an expression
- something has to be blocked, waiting
- hand the return value back over a channel
- buffer it by one so the last send never blocks
- select over the exit channel and ctx.Done()
basics
~20 sA goroutine cannot return anything to whoever started it. The usual answer is a channel: a small wrapper goroutine calls the worker function and sends its returned error on a chan error, and the supervisor waits on that channel.
solid answer
~40 sA `go` statement is not an expression and yields no value, so the worker's return has to be handed back over a channel. The idiom is a wrapper goroutine, `go func() { exits <- run(ctx) }()`, where `exits` is a `chan error` buffered by one so the worker never blocks on the send if the supervisor has already given up. The supervisor then sits in a `for` loop with a `select` over `ctx.Done()` and `exits`. When a value arrives it carries the exit reason — nil for a clean stop, non-nil for a failure — and the supervisor logs it and relaunches the worker. Before relaunching it checks `ctx.Err() != nil`, because a worker that returned because shutdown started must not be started again.
code
go · 19 lines// exits carries the worker's return value. It is buffered by one so the
// worker goroutine can finish even if the supervisor has already returned.
func superviseSession(ctx context.Context, dial func(context.Context) error) {
exits := make(chan error, 1)
launch := func() { go func() { exits <- dial(ctx) }() }
launch()
for {
select {
case <-ctx.Done():
return
case err := <-exits:
log.Printf("device session exited: %v", err)
if ctx.Err() != nil {
return // shutting down: do not relaunch
}
launch()
}
}
}go deeper
Be ready to say out loud that a go statement returns nothing and that a channel is how a result comes back. Sketch the wrapper goroutine that sends the worker's error, and the supervisor's select over that channel and ctx.Done().
Explain why the exit channel is buffered by one, and what leaks if it is not. Expect to be asked what the supervisor does with the error it receives, and how it tells a clean stop from a crash.
Show the shape that survives production: one identified exit value per worker, a shutdown check before relaunching, and a log line that carries the wrapped reason rather than a bare 'worker restarted'.
Own the convention across the codebase: what a long-lived worker function's error return means, whether nil is reserved for orderly stops, and how exits are surfaced so an operator can see which worker is flapping.
## The `go` statement gives you nothing back `go f()` starts `f` in a new goroutine and the statement completes immediately. It is a statement, not an expression: it has no value, there is no handle object, and there is no way to "join" a goroutine and read its result the way you might await a task in other runtimes. The language deliberately offers exactly one way to get information out of a goroutine — communication over a channel (or a `sync` primitive plus shared memory, which is the same idea with more foot-guns). That is the whole reason a supervisor exists as a separate goroutine: something has to be blocked, waiting to be told that the worker is gone. ## The exit channel The standard shape is a wrapper goroutine whose only job is to call the worker and forward its return value: ```go exits := make(chan error, 1) go func() { exits <- dial(ctx) }() ``` `dial` here is the long-lived work — for a session manager, open a socket to a remote device, read frames from it until the connection dies, and return the error that ended it. The wrapper turns "the function returned" into "a value appeared on a channel", which is something a `select` can wait on alongside everything else the supervisor cares about. The error value is the report. Return `nil` for an orderly stop (the context was cancelled, the peer closed cleanly) and a non-nil error otherwise; wrap it with `fmt.Errorf("...: %w", err)` so the supervisor can log a reason a human can act on. ## Why the channel is buffered by one An unbuffered send blocks until a receiver is ready. If the supervisor has returned — because shutdown finished, or because it gave up on this worker — nobody will ever receive, and the worker goroutine blocks on the send forever. That is a goroutine leak: it holds its stack, and anything it references, for the life of the process. Giving the channel a buffer of one lets the final send complete unconditionally, and the value is simply discarded when the channel becomes unreachable. ## The supervisor loop ```go for { select { case <-ctx.Done(): return case err := <-exits: // worker is gone; decide what to do about it } } ``` Two cases, two reasons to wake up. `ctx.Done()` is the shutdown path: the supervisor stops looping and returns; the worker sees the same cancelled context and unwinds on its own. The `exits` case is the interesting one: the worker died on its own terms and the supervisor now owns the decision to relaunch. The check that beginners miss sits inside that second case. After a shutdown starts, in-flight workers return almost immediately — with an error, because their socket read was interrupted. If the supervisor relaunches blindly it will start a fresh worker during shutdown, and possibly loop doing so. Test `ctx.Err() != nil` (or `errors.Is(err, context.Canceled)`) and return instead. ## Many workers, one channel With a session per remote device you do not want a channel per worker in the `select`; `select` cases are fixed at compile time. Send a small struct instead: ```go type exit struct { device string err error } ``` One `chan exit` buffered to the number of workers lets a single supervisor loop learn which worker died and why, and relaunch just that one. ## What not to do Writing the result into a shared variable and reading it from the supervisor is a data race unless something synchronises the two goroutines; a channel send and the matching receive give you that synchronisation for free, which is why the idiom is a channel and not a package-level `var lastErr error`. A `sync.WaitGroup` is also not a substitute here. It tells you that everything you counted has finished; it carries no value and cannot tell you *which* worker stopped or *why*. A `WaitGroup` is how you wait for a shutdown to complete; a channel is how you learn about an individual exit.
- Why buffer the exit channel by one instead of leaving it unbuffered?So the worker's final send can always complete. If the supervisor has already returned — after shutdown, or after giving up on this worker — an unbuffered send has no receiver and the worker goroutine blocks on it forever, leaking its stack and everything it references. With a buffer of one the send succeeds and the value is dropped when the channel becomes unreachable.
- The worker returns an error immediately after you cancel the supervisor's context. Should the supervisor relaunch it?No. That error is the shutdown, not a crash: the worker's socket read was interrupted by cancellation. The supervisor should check `ctx.Err() != nil`, or `errors.Is(err, context.Canceled)`, in the exit branch and return instead of relaunching. Without that check a shutdown turns into a burst of freshly started workers that immediately die again.
- How do you supervise fifty device sessions without fifty select cases?Use one channel carrying a struct that identifies the worker, such as `struct{ device string; err error }`, buffered to the number of workers. Every wrapper goroutine sends its device id along with the returned error, and the single supervisor loop learns which session died and relaunches only that one. `select` cases are fixed at compile time, so the identity has to travel in the value.
A go statement is like posting a letter with no return address. If you want to hear back, you have to hand the worker a self-addressed envelope first — that envelope is the channel.
saying these in an interview costs you the question
- Thinks a go statement evaluates to the function's return value
- Stores the worker's error in a shared variable instead of sending it
- Sends the exit on an unbuffered channel nobody will read
- Relaunches the worker even after the context was cancelled
- Uses a WaitGroup and expects it to say which worker failed