skip to content

A byte-counting http.ResponseWriter wrapper made io.Copy of a large file to the client slower. Why?

level: seniorimportance: nice to knowfreq 26%

answer

  1. io.Copy asks the destination a question first
  2. the destination is now your struct
  3. a 32 KiB buffer appears from nowhere
  4. the writer underneath still has the fast path
  5. declare ReadFrom and count what it returns

basics

~20 s

net/http's own response writer implements io.ReaderFrom, and io.Copy uses that fast path. The wrapper's type does not, so io.Copy fell back to a 32 KiB buffered loop of Write calls. Declare ReadFrom on the wrapper, delegate, and add the returned count.

solid answer

~50 s

`io.Copy` first checks whether its destination implements `io.ReaderFrom`, and `net/http`'s response writer does — which on a plain HTTP/1 connection can push the bytes down to the connection's own `ReadFrom` and, for an `*os.File` source, reach the kernel's zero-copy send path. Once a middleware substitutes `&recorder{ResponseWriter: w}`, the destination's dynamic type is `*recorder`, which satisfies only `io.Writer`, so `io.Copy` allocates a 32 KiB buffer and loops read-write-read-write through user space. Nothing is broken and nothing logs, but a large-file endpoint gains copies, syscalls and allocations, visible in a benchmark with `-benchmem` or a CPU profile. The fix is to give the wrapper its own `ReadFrom` that type-asserts the wrapped writer to `io.ReaderFrom`, delegates, and adds the returned count to the byte counter — falling back to `io.Copy` on the wrapped writer when the assertion fails. `http.ResponseController` does not help here: it has no `ReadFrom`.

code

go · 11 lines
go
func (r *recorder) ReadFrom(src io.Reader) (int64, error) {
	if rf, ok := r.ResponseWriter.(io.ReaderFrom); ok {
		n, err := rf.ReadFrom(src)
		r.nbytes += n
		return n, err
	}
	// copy onto the wrapped writer, not onto r, or this recurses
	n, err := io.Copy(r.ResponseWriter, src)
	r.nbytes += n
	return n, err
}

go deeper

for a junior

Take away the headline: io.Copy has a faster route into net/http's writer, and a wrapper that only forwards Write hides it.

for a middle

Explain the two interface checks io.Copy performs, the 32 KiB fallback buffer, and how declaring ReadFrom on the wrapper puts the destination back on the fast path.

for a senior

Demonstrate the diagnosis: a benchmark with -benchmem or a CPU profile pinning extra copying on the wrapper, then a fix that keeps the byte count correct on the new path.

for a principal

Weigh whether the throughput matters for your traffic mix at all, and argue for paying this complexity once in a shared wrapper rather than letting each service rediscover it.

## What the fast path was `io.Copy(dst, src)` is not a plain loop. Before allocating anything it asks two questions: 1. Does `src` implement `io.WriterTo`? If so, call `src.WriteTo(dst)`. 2. Does `dst` implement `io.ReaderFrom`? If so, call `dst.ReadFrom(src)`. Only if both fail does it allocate a 32 KiB buffer and loop `Read` into it, then `Write` out of it. `net/http`'s response writer implements `io.ReaderFrom`. Under the right conditions — an ordinary HTTP/1 body with no compression or chunk-rewriting layer in between, and an `*os.File` on the reading side — that path can hand the copy to the connection itself and let the operating system move the bytes without ever bringing them into the process. Even when it cannot, delegating to the writer's own `ReadFrom` avoids the extra buffer and the extra pair of copies per chunk. This is not an exotic path: `http.ServeContent` and `http.ServeFile` copy through it, so any endpoint serving files or large payloads is on it by default. ## What the wrapper did A recording wrapper embeds the `http.ResponseWriter` interface, which promotes exactly `Header`, `Write` and `WriteHeader`. The wrapper does not implement `io.ReaderFrom`, so step 2 above now fails and `io.Copy` takes the slow branch. Per response you gain a 32 KiB allocation, a read syscall and a write syscall per chunk, and a full user-space copy of every byte. On a hot large-file endpoint that turns up as higher CPU per byte and a jump in `B/op` and `allocs/op` under `go test -bench . -benchmem`, or as `syscall.Write` and memory-move frames dominating a CPU profile taken with `go tool pprof`. The reason it is hard to notice is that nothing failed. The bytes are correct, the status is correct, the tests pass. Only the throughput of one class of endpoint changed, and the change arrived with a middleware whose stated purpose was to count bytes. ## The fix Give the wrapper a `ReadFrom` so its own type satisfies `io.ReaderFrom` again: ```go func (r *recorder) ReadFrom(src io.Reader) (int64, error) { if rf, ok := r.ResponseWriter.(io.ReaderFrom); ok { n, err := rf.ReadFrom(src) r.nbytes += n return n, err } n, err := io.Copy(r.ResponseWriter, src) r.nbytes += n return n, err } ``` Three details matter: - **Add the returned `n` to the counter.** This is the bug that hides inside the fix: once `ReadFrom` exists, body bytes no longer flow through your `Write` override, so a wrapper that forwards `ReadFrom` without counting will log large responses as zero bytes — the same class of wrong-number defect as recording status 0. - **Fall back with `io.Copy` on the wrapped writer, not on the wrapper.** Passing `r` back into `io.Copy` would re-enter your own `ReadFrom` and recurse. - **Remember the implicit status.** The delegated write triggers `net/http`'s implicit 200 without calling your `WriteHeader` override, so the wrapper still needs a sane initial status. ## What does not fix it `http.ResponseController` is the standard answer for lost capabilities, but its surface is `Flush`, `Hijack` and the read and write deadlines. There is no `ReadFrom` on it, and adding `Unwrap` to your wrapper does not change the wrapper's own method set, so `io.Copy` still sees only an `io.Writer`. This is one of the few capabilities you genuinely have to forward by declaring the method. Making your own `Write` use a bigger buffer does not help either — the buffer belongs to `io.Copy`, not to you — and setting `Content-Length` changes framing, not the copy strategy. ## Is it worth doing? Measure before you decide. For an API serving small JSON bodies, the `ReadFrom` path is irrelevant and the extra method is noise. For a service that streams files, media or large exports, the difference is real and it is exactly the kind of regression that gets attributed to "the new observability layer" without anyone being able to say why. Because the wrapper is shared, the cost of forwarding `ReadFrom` is paid once and every endpoint keeps its original throughput — which is the general argument for putting careful, well-tested wrapping in one library rather than in each service.

  • After forwarding ReadFrom, why might the logged byte count drop to zero for large responses?
    Because the bytes no longer pass through the wrapper's `Write` override — `io.Copy` hands the whole source to `ReadFrom` instead. If that method delegates without adding the `int64` it returns to the counter, every file response logs zero bytes. Forwarding a capability always means re-establishing whatever the wrapper was measuring on the old path.
  • Would implementing Unwrap and using http.ResponseController restore this fast path?
    No. `http.ResponseController` exposes `Flush`, `Hijack` and the read and write deadline setters; it has no `ReadFrom`, and `Unwrap` does not add methods to your wrapper's type. `io.Copy` inspects the destination value it was handed, so the only way back onto the fast path is for the wrapper itself to implement `io.ReaderFrom`.
  • How would you confirm the regression rather than assume it?
    Write a benchmark that serves a large payload through the handler with and without the wrapper and run it with `-benchmem`: the slow path shows a 32 KiB-per-call allocation and more time per byte. A CPU profile of the live endpoint shows the same thing as user-space copying and extra write syscalls that disappear once `ReadFrom` is forwarded.

saying these in an interview costs you the question

  • Assumes embedding keeps the wrapped writer's io.ReaderFrom
  • Says http.ResponseController can restore ReadFrom
  • Calls io.Copy on the wrapper itself and recurses forever
  • Forwards ReadFrom without adding the returned count
  • Blames the handler or the file system rather than the wrapper
  • Tries to fix it by enlarging the wrapper's own Write buffer