How does io.TeeReader let you hash a stream while something else consumes it?
answer
- observation rides along with the read
- a hash is just an io.Writer
- nothing flows until the consumer pulls
- an unread tail is an unhashed tail
basics
~20 sio.TeeReader(r, w) returns a Reader that writes every byte it reads from r into w before returning it. Pass a hash as w and the digest builds itself as the consumer reads - but only over bytes actually read.
solid answer
~40 s`io.TeeReader(r, w)` gives back a reader that forwards each `Read` to `r` and, for every byte it obtains, writes those bytes into `w` before the `Read` returns. Since `hash.Hash` is itself an `io.Writer`, passing `sha256.New()` as `w` means the digest is computed in flight, with no second pass and no full copy in memory. The consumer keeps its existing code — it just reads from the tee instead of the original. Two consequences follow from it being pull-driven: nothing reaches the hash until somebody reads, so `Sum` before draining gives you the digest of an empty input; and if the write side fails, the error surfaces out of the tee's `Read`, not silently. The tee also has no `Close` and no locking, so lifetimes and single-consumer discipline stay with your code.
code
go · 7 linesh := sha256.New()
tee := io.TeeReader(src, h)
early := h.Sum(nil) // the digest of zero bytes: no read has happened yet
// ... consume tee to io.EOF ...
final := h.Sum(nil) // now the digest covers everything the consumer readgo deeper
Be ready to state what the two arguments are and which direction bytes flow: read from the first, and everything read is copied into the second. Knowing that a hash can be used directly as the writer is the practical payoff.
Explain the ordering — the write happens inside Read, before it returns — and why the digest is empty until something drains the stream. Say where a write-side error surfaces and why that is the deliberate design.
Show the failure you would actually hit in production: a consumer that stops early leaving a digest that covers a prefix, or an observing sink slow enough to throttle the main transfer because there is no concurrency hiding it.
Argue about where observation belongs: instrumentation on the read path costs latency on every byte, so decide what genuinely must be inline (integrity) versus what can be sampled or moved off the hot path, and who owns that budget.
## Observing a stream without copying it twice A recurring need in stream handling is to compute something *about* the bytes — a checksum, a byte count, a copy for an audit log — while the bytes are already on their way somewhere else. The naive approach is to read everything into memory, compute the digest, then hand the buffer to the consumer. That doubles memory and bounds you by the largest payload anyone sends. `io.TeeReader` solves it in one pass. ### The contract `io.TeeReader(r io.Reader, w io.Writer) io.Reader` returns a reader that, on each `Read`, first reads from `r` and then writes exactly the bytes it obtained into `w`. The write happens *before* `Read` returns, on the same goroutine, so by the time the consumer has the bytes in hand, `w` has already seen them. If the write to `w` fails, that error is returned from the tee's `Read` — the tee refuses to let an observation failure pass unnoticed. If the underlying read yields some bytes together with an error, those bytes are still handed to `w`. Because `hash.Hash` embeds `io.Writer`, a digest is a legal tee target with no adapter at all. So is a `bytes.Buffer` (keep a copy of a small body), a counter you wrote yourself, or an `io.MultiWriter` combining several of those. ### The trap: it is pull-driven Nothing happens in a tee until somebody reads. `io.TeeReader` starts no goroutine and does no work of its own; every byte the hash sees is a byte the consumer already pulled. Two failure modes come from this. The first is calling `Sum` too early. Constructing the tee and immediately asking the hash for its digest gives you the digest of zero bytes — a value that looks plausible and is completely wrong. The digest is only meaningful after the tee has been read to `io.EOF`. The second is a consumer that stops early. If your consumer decodes a value and stops, or gives up on the first error, the tee only ever saw the prefix that was read, and the digest covers that prefix rather than the whole stream. When you need the digest of the *entire* input, you must make sure the tee is drained to EOF — read whatever the consumer left behind before you call `Sum`. Verifying a checksum against a partly-consumed stream is a genuine correctness bug, and it usually shows up as an intermittent mismatch on exactly the inputs where the consumer short-circuits. ### Where the observation goes Because the tee's second argument is an ordinary `io.Writer`, this is the hook point for anything you want to know about a stream you do not want to rewrite. A byte counter is a five-line type. A rolling sample for debugging is a bounded buffer. Several observers at once is `io.MultiWriter` in the tee slot — but note that the strict fan-out rule applies there too, so a failing observer becomes a read error on the main path unless you wrap it in something tolerant. One caution: the tee sits on the *read* path, so everything it does is on the critical path of the transfer. Hashing is cheap; writing each chunk to a slow sink is not, and it will show up directly as reduced throughput because there is no concurrency hiding it. ### Lifetimes and concurrency The returned value is an `io.Reader` and nothing more. It has no `Close`, so it never closes `r`, and it does not close or flush `w` either — if `w` is something that needs finishing, you finish it yourself after the drain. It contains no mutex, so a tee is for a single consuming goroutine; sharing one across goroutines races on both the underlying reader and the observer. ### Reader side versus writer side It is worth being clear about the direction. `io.TeeReader` observes bytes flowing *out of* a source, so it is the right tool when you own the reader end — an incoming request body, a file being uploaded, a stream you are decoding. When you own the destination end instead, you wrap the writer, either by hand or with `io.MultiWriter`. Same idea, opposite side of the pipe, and choosing the wrong side usually means you end up wrapping something you do not actually control.
- The consumer stops reading after the value it needs. What is wrong with the digest?It only covers the prefix that was actually read, because the tee is driven entirely by the consumer's reads. If you need the digest of the whole input, drain the tee to `io.EOF` yourself before calling `Sum` — read and discard the remainder — or accept explicitly that the digest describes the consumed prefix and nothing more.
- What happens if the write side of an io.TeeReader returns an error?That error comes back out of the tee's `Read` call, so the consumer sees the transfer fail rather than quietly losing the observation. It is a deliberate choice: a checksum or audit copy that silently stopped being written would be worse than a visible failure. If an observer is genuinely optional, wrap it in a writer that swallows its own errors and reports a full write.
- Can several goroutines read from the same io.TeeReader?No. The tee holds no lock, so concurrent reads race on both the underlying reader and the observing writer, and the bytes handed to the observer can interleave arbitrarily. A tee belongs to exactly one consuming goroutine; if you need fan-out across goroutines, do the hand-off explicitly with a channel or a pipe and keep each end single-consumer.
A tee is a meter spliced into a pipe: it measures only while liquid is flowing, so if nobody opens the tap at the far end the meter honestly reads zero.
saying these in an interview costs you the question
- Calls Sum before the tee has been read to EOF
- Thinks the tee writes in the background on its own goroutine
- Assumes a write failure on the observing side is ignored
- Believes io.TeeReader closes or flushes either side
- Shares one tee across several reading goroutines