skip to content

How should a long-running net/http streaming handler notice the client disconnected and stop?

level: seniorimportance: should knowfreq 35%

answer

  1. the loop needs a second way out
  2. waiting on one channel is the bug
  3. the request carries the signal
  4. a successful write proves less than you think

basics

~20 s

Select on r.Context().Done() inside the write loop and return when it fires - the server cancels the request context when the client goes away. Treat an error from Write or Flush as a second, later signal.

solid answer

~50 s

A streaming handler must be written as a loop with two exits: the next unit of data, and cancellation. `r.Context()` is cancelled when the client disconnects, so a `select` with a `<-r.Context().Done()` case returns promptly instead of blocking forever on a source that will never be read. Never block on a bare channel receive or a blocking read in a streaming loop - that is how a handler goroutine, the connection, and whatever file or subscription it holds survive the user closing their laptop. An error from `Write` or `Flush` is also a disconnect signal, but a lagging one: bytes go into the kernel socket buffer, so several writes can succeed after the peer is gone. Use `defer` in the handler to close the tail source, and remember the same context is cancelled when the handler returns, so work that must outlive the response cannot use it.

code

go · 18 lines
go
func tail(w http.ResponseWriter, r *http.Request, lines <-chan string) {
	w.Header().Set("Content-Type", "text/plain; charset=utf-8")
	rc := http.NewResponseController(w)
	for {
		select {
		case <-r.Context().Done():
			return // client hung up: normal, not an error
		case line, ok := <-lines:
			if !ok {
				return
			}
			fmt.Fprintln(w, line)
			if err := rc.Flush(); err != nil {
				return
			}
		}
	}
}

go deeper

for a junior

Know that the request carries a context which is cancelled when the client disconnects, and that a streaming loop should return when it fires rather than looping forever.

for a middle

Explain why the select needs the cancellation case even while waiting for data, and why cleanup belongs in defers in the handler so that returning is sufficient.

for a senior

Show that you have operated one of these: the write error is a lagging signal, disconnects are not failures, slow consumers need an explicit policy, and you verify the fix by killing a client and watching the resource count fall.

for a principal

Set the house rule for streaming endpoints - bounded per-client buffering, a defined slow-consumer policy, and a distinct outcome label for hang-ups - so per-connection cost is capped by design rather than by whoever writes the next handler.

## The failure this prevents A live log-tail endpoint holds resources for as long as it streams: a handler goroutine, a connection, an open file or a subscription to a fan-out. If the loop only ever exits when the data source ends, then a support engineer who closes the tab leaves all of that running - and on a busy day you accumulate hundreds of them. This is the classic streaming leak, and it is invisible in tests because tests always read to the end. ## The prompt signal: the request context The server cancels the context returned by `r.Context()` when the client goes away. So the streaming loop should be a `select` with a cancellation case: ```go for { select { case <-r.Context().Done(): return case line, ok := <-lines: if !ok { return } fmt.Fprintln(w, line) if err := rc.Flush(); err != nil { return } } } ``` Two properties matter. First, the cancellation case is checked on every iteration, including while waiting - a bare `for line := range lines` has no such case and blocks indefinitely on a dead client. Second, returning is enough: the server finalises the response, and `defer`red cleanup in the handler closes the file or unsubscribes. Any goroutine the handler started for this request must also be told to stop, either by the same context or by a channel the handler closes on the way out. ## The lagging signal: write errors It is tempting to rely on `w.Write` failing once the client is gone. It does eventually, and you should still check it - a `Flush` or `Write` error means stop, immediately, with no point in logging it as a server fault. But it is not prompt. A write only has to reach the kernel's socket buffer to succeed, so a handler can write and flush several times after the peer has vanished before an error surfaces. If your only exit is a write error, the loop can spin for a long time producing data nobody will read; if the source is expensive, that is the whole cost of the leak with extra steps. ## Do not confuse cancelled with failed When the loop exits because the context was cancelled, nothing went wrong. The client hung up; that is a normal end for a stream. Log it at debug level if at all, do not increment an error counter, and do not try to write an error into the response - there is nobody there. Treating client disconnects as 5xx is a common way to make a streaming endpoint's dashboards useless. ## The other end of the loop: pacing A source that produces faster than the client consumes will back up. Over a socket, a slow client eventually makes `Write` block once the kernel buffer and the peer's window are full, which turns into a handler that appears wedged. Decide deliberately what should happen: drop intermediate records and send the newest, buffer a bounded number and disconnect if the client falls too far behind, or block and let backpressure propagate. A bounded buffer with an explicit disconnect is usually the honest answer for a tail - it puts a ceiling on per-client memory and makes the failure visible instead of silently unbounded. ## Cleanup discipline Write the handler so every resource it acquires is released by a `defer` in the same function. Close the file, cancel the subscription, return the buffer. Then the only thing you need the loop to do on disconnect is `return`, and there is exactly one place where cleanup lives. If cancellation-driven cleanup is spread across several goroutines, the leak will come back the first time someone adds a case to the `select`. ## Verifying it The cheap check is to start the stream with `curl -N`, kill the client, and confirm the server-side goroutine count and any per-stream metric fall back to their baseline within a second or two. A stream that only cleans up when the source ends will hold flat, and that is exactly the shape of the bug you are looking for.

  • Why is an error from Write not a good enough disconnect signal on its own?
    Because a write only has to reach the kernel's socket buffer to succeed. After the peer disappears, several writes and flushes can still return nil before an error surfaces, so a loop whose only exit is a write error keeps producing data for a client that is gone. Check the error and stop on it, but drive the loop from the request context.
  • The client is much slower than the source. What are your options?
    Three, and you should pick one explicitly: drop intermediate records and always send the newest; buffer a bounded number per client and disconnect anyone who falls too far behind; or block and let backpressure reach the producer. For a live tail, a bounded buffer plus disconnect is usually right - it caps per-client memory and makes the slow consumer visible rather than silently unbounded.
  • Should a stream that ends because the client disconnected count as a failed request?
    No. A disconnect is the normal end of a stream, so it should not increment error counters or produce a 5xx in your metrics - you cannot send a status anyway, the headers went out long ago. Record it as a distinct outcome, with the duration and bytes streamed, so dashboards can tell a hang-up apart from a source failure.

saying these in an interview costs you the question

  • Ranges over the source channel with no cancellation case
  • Relies only on a Write error to detect the disconnect
  • Counts every client hang-up as a server-side error
  • Starts per-request goroutines with no way to stop them
  • Buffers unboundedly for a slow consumer