skip to content

A Go proxy's io.Copy to the client fails mid-body after http.ResponseWriter.WriteHeader — what does the returned error let you do that stack unwinding would not?

level: seniorimportance: nice to knowfreq 33%

answer

  1. the failing hop knows what the top does not
  2. the status line is already gone
  3. how many bytes reached the client
  4. retry is off the table once n is above zero

basics

~20 s

It lets the hop that failed act while it still has the context: stop the copy, close the upstream, and record how many bytes io.Copy relayed. Unwinding to one handler loses where the failure happened.

solid answer

~50 s

Once `WriteHeader` has run, the status line is gone — you cannot turn a 200 into a 502 and the client already holds part of the body. The only useful actions are local ones, and the returned error is what makes them possible: `io.Copy` hands back both an `error` and an `int64` byte count, so the failing hop knows whether the upstream died on the first byte or after most of the body, can close the upstream reader, can decide that a retry is now impossible, and can log a truncation the client will otherwise experience silently. A top-level handler that merely catches has none of that — it knows something failed, not where or how far. The cost is that the decision is written at every hop, and a hop that logs without returning keeps writing into a half-broken response.

code

go · 9 lines
go
func forward(w http.ResponseWriter, upstream io.Reader) {
	w.WriteHeader(http.StatusOK)
	n, err := io.Copy(w, upstream)
	if err != nil {
		// the 200 is already on the wire; n says how far we got
		log.Printf("relayed %d bytes then failed: %v", n, err)
		return
	}
}

go deeper

for a junior

Focus on the mechanics first: io.Copy returns a byte count and an error, and the error must be checked with a return before anything else happens.

for a middle

Explain why the status code is fixed once the headers are written, and what the byte count means when the copy fails part-way.

for a senior

Show the operational reasoning: which decisions only the failing hop can make, why a retry stops being safe once bytes have reached the client, and what a log line must carry so the truncation is not silent.

for a principal

Own where recovery belongs across a proxy fleet — how much per-hop handling is worth its repetition, and what the standard log and metric shape is so incidents are comparable across services.

## The scenario A proxy sits in front of an upstream. A request arrives, the proxy dials the upstream, gets a response, writes the status and headers to the client, then streams the body through with `io.Copy(w, upstreamBody)`. Somewhere in the middle of a large body the copy fails — the upstream connection dropped, or the client went away. This is the moment where Go's choice to return failures rather than throw them pays, and it is worth being precise about why. ## What is already irreversible `http.ResponseWriter.WriteHeader` commits the status line. After it runs: - the status code cannot be changed; a later call is ignored, - the client has already received the headers and whatever body bytes were flushed, - there is no buffer to rewind and no way to replay the request to another upstream and pretend the first attempt did not happen. So the fantasy of "catch it at the top and return a 502" is not available regardless of the error model. What remains is a set of *local* decisions, and those are exactly what a returned error supports. ## What the returned error buys the failing hop 1. **Position.** The hop that called `io.Copy` knows which upstream it was talking to, which request this is, and how long the copy had been running. A handler ten frames up knows only that something failed somewhere below. 2. **Progress.** `io.Copy` returns `(int64, error)`. The byte count is meaningful even on failure and it changes the diagnosis completely: zero bytes points at the upstream refusing or resetting immediately, while a large count near the expected content length points at a slow client, an idle timeout, or a client that disconnected. Two incidents with identical error text are different incidents once you have `n`. 3. **Cleanup and containment.** The hop can close the upstream body immediately rather than letting it dangle until something else notices, and it can stop writing rather than continuing to push bytes into a response the client has abandoned. 4. **A decision it is qualified to make.** Only this hop knows whether a retry was ever possible. Before the first byte reached the client a retry against a second upstream might have been legitimate; after `n > 0` it is not, because the client already has a partial body. That distinction is invisible to a top-level handler. ## Why a chain of hops still works A Go proxy is usually a chain: an outer handler that wraps an inner one, that wraps the transport call. Each layer is ordinary code, and each returns or propagates an ordinary value. That means a layer can inspect the failure and add what it knows — which route matched, which upstream was chosen, how long the attempt took — without any special mechanism. The chain composes because errors are values, not because a runtime is carrying them. The distinction from unwinding is not that unwinding is impossible; it is that unwinding takes the decision *away from the frames that had the context* and hands it to one place that has none. Go's model keeps the decision where the knowledge is. ## The cost, and the failure mode to watch Two costs are real and you should name them. First, repetition: the same shape of check appears at every hop, and a reader has to scan past it. Second — and much more damaging — a hop that *observes* the failure without *stopping*. In this streaming context that is not a cosmetic bug: logging the copy error and then continuing to write means the proxy keeps pushing bytes into a response whose upstream is gone, or dereferences a response value that is nil because the dial failed. The visible chain only helps if every check ends the function. ## What to say about telling the client A truncated body is the one thing you cannot repair from here, and being honest about that is part of a good answer: the client sees a short read, your logs see the byte count, and the reconciliation between the two happens in observability rather than in the response. What you owe the operator is a log line carrying both the error and `n`, so the truncation is not a silent event.

  • Why can't the proxy send a 502 once io.Copy has failed part-way through the body?
    Because `WriteHeader` already committed the status line and the client holds part of the body. A later status write is ignored — there is nothing to rewind. All that remains is to stop copying, release the upstream, and log the truncation with the byte count so an operator can see it happened.
  • What does io.Copy's byte count change about the diagnosis?
    It separates two incidents that share the same error text. Zero bytes means the upstream failed essentially immediately — refused, reset, or a bad response. A large count means the stream was healthy for a long time and then died, which points at the client disconnecting or a timeout firing at one of the layers rather than at the upstream being unavailable.
  • What is the honest cost of handling this at every hop rather than once at the top?
    Repetition, and a new failure mode: a hop that logs the error and forgets to return keeps writing into a response whose upstream is already gone, or touches a nil response value. The visible chain only pays if every branch ends the function, which is a discipline a reviewer has to enforce.

saying these in an interview costs you the question

  • Thinks the status can be changed after WriteHeader
  • Treats one top-level handler as equal to per-hop handling
  • Ignores io.Copy's byte count when reporting the failure
  • Logs the copy error and keeps writing to the client
  • Wants to retry the upstream after bytes reached the client