What does io.Copy(dst, src) do in Go, and what do its two return values mean?
answer
- two interfaces, one loop
- memory does not grow with payload
- the count is what reached the destination
- a fixed scratch buffer, about 32 KB
basics
~20 sio.Copy streams bytes from a source reader to a destination writer through one small reusable buffer, so memory does not grow with the payload. It returns an int64 count of bytes written and the error that stopped the transfer.
solid answer
~50 s`io.Copy(dst, src)` repeatedly reads from `src` and writes what it read to `dst` until the source is exhausted, returning `(written int64, err error)`. The key property is that it is a *streaming* transfer: it moves data through a fixed scratch buffer of about 32 KB that it allocates internally, so relaying a 5 GB upload costs roughly the same memory as relaying 5 KB. That is why a service that receives an upload on one connection and forwards it to another should hand both ends to `io.Copy` rather than materialising the payload first. Because the parameters are the one-method interfaces `io.Writer` and `io.Reader`, the same call works for files, sockets, hashes, compressors and `io.Discard`. Note the count is bytes *written*, not bytes read, so a short count together with a non-nil error tells you exactly how far the transfer got before it broke.
code
go · 12 linesf, err := os.Open(name)
if err != nil {
return 0, err
}
defer f.Close()
h := sha256.New()
n, err := io.Copy(h, f) // n = bytes written into the hash
if err != nil {
return 0, err
}
fmt.Printf("%d bytes, %x\n", n, h.Sum(nil))go deeper
Be ready to state the signature and say what each return value is: bytes written as an int64, then the error. Say out loud that it streams rather than loading everything, and name two concrete reader/writer pairs you have used it with.
Explain the read-write loop, why only the bytes returned by Read get written, and that the scratch buffer is a fixed ~32 KB allocated per call. Be able to contrast the O(1) memory of a copy with the O(payload) memory of collecting first.
Turn it into a capacity argument: per-request memory is a buffer, not a payload, so concurrency times buffer size is a number you can defend. Mention io.ErrShortWrite as evidence you have debugged a writer that quietly accepted less than it was given.
Frame streaming versus staging as a policy for the codebase: which service boundaries are allowed to hold a whole payload, what the ceiling is when a remote peer picks the size, and where a bounded copy belongs instead of an unbounded one.
## What `io.Copy` is The signature is: `func Copy(dst Writer, src Reader) (written int64, err error)` Both parameters are single-method interfaces. `io.Reader` declares `Read(p []byte) (n int, err error)` — fill my slice with up to `len(p)` bytes and tell me how many you managed. `io.Writer` declares `Write(p []byte) (n int, err error)` — consume all of my slice, or explain why you could not. Because those interfaces are tiny, an enormous amount of the standard library satisfies one or both: `*os.File`, `net.Conn`, `*bytes.Buffer`, `*strings.Reader`, a `hash.Hash`, a `gzip.Writer`, `http.ResponseWriter`, `io.Discard`. `io.Copy` is the single function that connects any reader to any writer, which is why it shows up in almost every Go program that touches a stream. ## The loop it runs Conceptually `io.Copy` is a loop over a scratch buffer: 1. call `src.Read(buf)`, getting back `nr` bytes; 2. if `nr > 0`, call `dst.Write(buf[0:nr])`; 3. add the bytes the writer accepted to the running total; 4. repeat until the source signals it has no more data, or either side reports an error. Two details of that loop matter in review. First, only the bytes actually returned by `Read` are written — a reader is allowed to fill less than the whole buffer, and code that writes `buf` instead of `buf[:nr]` is a classic bug that `io.Copy` saves you from. Second, if a writer accepts fewer bytes than it was given without reporting an error, `io.Copy` stops and returns `io.ErrShortWrite`, because a writer that silently drops data is broken. ## The two return values `written` counts bytes that reached `dst`. It is deliberately *not* the number of bytes read: if the destination fails halfway, the count tells you how much of the stream landed, which is what you need for logging, metrics, resume logic, or a `Content-Length` you now cannot honour. `err` is the first error that stopped the transfer. Reaching the natural end of the source is not treated as a failure — a transfer that ran to completion returns a nil error, so the ordinary check is simply `if err != nil`. ## Memory: a fixed buffer, not the payload Unless a faster path is available, `io.Copy` allocates a scratch buffer of about 32 KB per call and reuses it for the entire transfer. This is the property that makes it the right tool for streams of unknown or hostile length: the memory cost is O(1) in the size of the data. Contrast that with any approach that first collects the whole stream into a slice or a string — there the memory cost is O(payload), set by whoever is on the other end of the socket rather than by you. The practical consequence: a relay that streams an upload straight from an inbound request to an outbound one holds ~32 KB per in-flight request, which is a capacity number you can multiply by your concurrency limit and reason about. The same relay written to collect first has no such number. ## Its close relatives - `io.CopyN(dst, src, n)` copies at most `n` bytes and reports whether it got all of them — the bounded form you want when a length is known in advance. - `io.CopyBuffer(dst, src, buf)` is the same copy but with the scratch buffer supplied by you, so a hot path can reuse one instead of allocating per call. - `io.WriteString(w, s)` writes a string to a writer, using the writer's own `WriteString` method when it has one (the `io.StringWriter` interface) instead of allocating a byte slice copy of the string. - `io.Discard` is an `io.Writer` that accepts and throws away everything, which is how you drain a source you do not need, or measure a producer in a benchmark without paying for storage. ## What to say in an interview Name the signature, say "streams through a small fixed buffer", and be explicit that the `int64` is bytes written. Then show you know why it matters: the alternative reads the whole payload into memory and lets a remote peer choose your heap size.
- How would you copy at most a known number of bytes instead of the whole source?Use `io.CopyN(dst, src, n)`, which copies exactly `n` bytes and returns `(written int64, err error)`. The error is nil only when it wrote the full `n`; if the source ran out first it reports `io.EOF` along with the short count, so a truncated source is visible rather than silently accepted.
- What is io.Discard, and why copy into it?`io.Discard` is an `io.Writer` that accepts every write and throws the bytes away. Copying into it lets you consume a stream you do not want to keep — for example to count how many bytes a producer emits, or to run a decoder in a benchmark — without allocating storage proportional to the stream.
- Why prefer io.WriteString(w, s) over w.Write([]byte(s))?`[]byte(s)` allocates and copies the whole string. `io.WriteString` first checks whether the writer implements `io.StringWriter` (a `WriteString(string) (int, error)` method) — `*bytes.Buffer`, `*strings.Builder`, `*bufio.Writer` and `*os.File` all do — and calls that directly, skipping the conversion. It falls back to the conversion only when the writer has no such method.
It is a bucket brigade with exactly one bucket: however big the lake, you only ever hold one bucketful at a time.
saying these in an interview costs you the question
- Says io.Copy reads the whole source into memory first
- Thinks the returned count is bytes read, not bytes written
- Believes io.Copy needs to know the source's length up front
- Reaches for a read-it-all helper on a stream of unknown size
- Writes the whole scratch buffer instead of only the bytes read