skip to content

A relay streaming uploads through io.Pipe leaks a goroutine per request. Why, and how do you fix it?

level: seniorimportance: should knowfreq 42%

answer

  1. someone is parked, not spinning
  2. the writer waits for a reader that left
  3. the collector cannot reclaim a blocked goroutine
  4. closing the read end is the fix
  5. io.ErrClosedPipe is what unblocks it

basics

~10 s

The producing goroutine is blocked writing to a PipeWriter nobody will read, because the consumer abandoned the PipeReader. Closing the PipeReader makes that Write return io.ErrClosedPipe so the goroutine can finish.

solid answer

~50 s

In the classic relay you create `pr, pw := io.Pipe()`, start a goroutine that copies the inbound upload into `pw`, and hand `pr` to the outbound request. If the consuming side returns early — the request fails to build, the outbound call errors, a validation check rejects the upload, the caller's `context.Context` is cancelled — nobody reads `pr` again. Since `io.Pipe` has no buffer, the producer's `Write` parks forever, and with it the goroutine and everything it references, including the inbound body. Garbage collection does not save you: a blocked goroutine is a GC root, so nothing is collected either. The fix has two halves. On the consuming side, `defer pr.Close()` on every exit path, which makes the blocked `Write` return `io.ErrClosedPipe` and lets the producer unwind. On the producing side, always finish with `pw.CloseWithError(err)` so a producer failure reaches the reader instead of looking like a clean end of stream.

code

go · 12 lines
go
func relay(dst io.Writer, src io.Reader) error {
	pr, pw := io.Pipe()
	defer pr.Close() // unblocks the producer on every early return

	go func() {
		_, err := io.Copy(pw, src)
		pw.CloseWithError(err) // nil err = clean end of stream
	}()

	_, err := io.Copy(dst, pr)
	return err
}

go deeper

for a junior

Remember that io.Pipe holds no bytes, so a write waits for a reader. If nobody reads and nobody closes, that goroutine never returns. Be able to point at the two Close calls a correct pipe relay needs.

for a middle

Explain the mechanics: the Write parks until reads consume it or an end closes; closing the read end makes the pending Write return io.ErrClosedPipe. Show that you know CloseWithError propagates a producer failure while a bare Close looks like a clean end.

for a senior

Diagnose it from the outside in: a climbing goroutine count, a goroutine profile full of identical parked stacks, memory held because each blocked goroutine is a GC root. Then argue ownership — who closes each end on every path — and add a regression test that exercises the failure paths.

for a principal

Own the pattern rather than the incident: make pipe ownership a reviewable rule, decide whether relays get a shared helper that encapsulates both closes, and set the expectation that leak checks run in tests rather than being discovered by a memory graph a week later.

## The shape that leaks The relay looks innocent: Create a pipe, start a goroutine that copies the source into the write end, hand the read end to whatever consumes it. It streams at constant memory and reads beautifully. Then the goroutine count climbs all week and never comes down. The reason is the property that makes `io.Pipe` useful: it has no buffer. A `Write` on the `PipeWriter` returns only once readers have consumed the bytes it handed over, or once one of the ends is closed. If the consumer stops reading and simply drops the `PipeReader` on the floor, the producer's `Write` blocks forever. Every early return on the consuming side is a way to drop it: a request that fails to build, an outbound call that errors after the first few kilobytes, a validation failure, a `context.Context` cancelled by the caller hanging up, a `panic` recovered further up the stack. In each case the deferred cleanup you did write covers the consumer, and the goroutine you started is still parked. ## Why the garbage collector does not clean it up A common wrong answer is "the pipe becomes unreachable and gets collected". It does not. Every live goroutine is a root for the collector, so a goroutine blocked in a `Write` keeps alive its stack, the pipe, the source it was copying from, and any buffers on the way. That is why a leaked goroutine usually shows up first as growing memory: each one pins a working set. There is no finalizer that unblocks a pipe, and nothing times it out. ## Diagnosis The symptom is a goroutine count that only rises. Two ways to pin it down: - **In a test.** Sample the goroutine count before the operation, run it, let things settle, and sample again — the ones that never returned are the leak, and asserting on this in a test is what stops the bug from coming back. Exercise the *failure* paths, not just the happy one: the happy path usually drains the pipe correctly and leaks nothing, which is exactly why this survives code review. - **In production.** Fetch the goroutine profile. The stacks tell you immediately: dozens or thousands of goroutines parked inside the pipe's write, all with the same call site above them, pointing straight at the relay. A CPU profile shows nothing here — the goroutines are parked, not spinning — and the race detector is irrelevant, because there is no race, only a wait that never ends. ## The fix, both halves **Consumer side: close the read end on every path.** `defer pr.Close()` in the function that owns the pipe. A blocked `Write` then returns `io.ErrClosedPipe` and the producing goroutine unwinds. Note the ordering trap: if you hand `pr` to something that takes ownership and closes it for you, that close is what saves you — but if the object is never handed over (the construction failed, or the call was never made), nothing closes it and you must. **Producer side: always close the write end, with the error.** Finish the producing goroutine with `pw.CloseWithError(err)` rather than a bare `pw.Close()`. A plain `Close` tells the reader the stream ended normally — so a producer that died halfway delivers a truncated payload the consumer happily accepts as complete, which is a silent corruption bug rather than a visible failure. `CloseWithError(err)` delivers that error to the reader's next read instead. Passing a nil error is equivalent to a plain `Close`, so the single line handles both the success and the failure path. Symmetrically, `pr.CloseWithError(err)` delivers a chosen error to the blocked writer rather than `io.ErrClosedPipe`, which is useful when the producer logs what stopped it. ## Why not a timeout instead A tempting non-fix is to wrap the copy in a timeout or a `select`. It converts an unbounded leak into a slow one, and it papers over an ownership bug: the question is not "how long should the producer wait" but "who closes this pipe, on every path". Answer that in the code and the timeout becomes unnecessary. ## The rule to carry away Every `io.Pipe` has exactly two owners and each owner has exactly one obligation: the producer must close the write end when it stops producing, and the consumer must close the read end when it stops consuming — both on every path, including the ones that return an error. Write those two `defer`s at the moment you write `io.Pipe()`, not afterwards.

  • How would you catch this leak in a test rather than in production?
    Wrap the operation in a test that samples the goroutine count before and after, with a short settle window, and fails if it did not return to the baseline. Crucially, drive the *error* paths — an outbound call that fails midway, a cancelled `context.Context` — because the success path drains the pipe and leaks nothing.
  • Which error does a Write on an io.PipeWriter return after the read end is closed?
    `io.ErrClosedPipe`, if the reader was closed with a plain `Close`. If the reader used `CloseWithError(err)`, that error is delivered to the writer instead. Either way the blocked `Write` returns promptly with a short count, which is what lets the producing goroutine unwind.
  • What does the consumer see if the producer dies halfway and calls a plain pw.Close()?
    A clean end of stream — the consumer believes it received the whole payload and happily commits a truncated object. That silent-corruption outcome is why the producing goroutine should end with `pw.CloseWithError(err)`: a nil error still means a normal close, and a non-nil one surfaces at the consumer's next read.
  • Would adding a timeout around the copy be an acceptable fix?
    It bounds the damage but does not fix it: the leak becomes slow rather than absent, and it hides an ownership bug. The real question is who closes each end on every path. Once the consumer's `defer pr.Close()` and the producer's `CloseWithError` are in place, the timeout has nothing left to catch.

The producer is standing at the counter holding a parcel for a clerk who went home. Nobody will take it, and shouting at the parcel does not help — someone has to close the counter.

saying these in an interview costs you the question

  • Expects the garbage collector to reclaim a blocked goroutine
  • Blames Go's goroutine scheduler rather than the blocked Write
  • Closes the PipeWriter only on the success path
  • Adds a timeout instead of closing the read end
  • Looks at a CPU profile for goroutines that are parked
  • Tests only the happy path, where nothing leaks