A heap profile of your Go frame reader shows one 900 MB []byte built from a peer's declared length. What went wrong?
answer
- who chose that number
- four bytes bought a gigabyte
- the allocation happens before the data
- compare against a maximum, then make
- a lying length has broken the framing
basics
~20 sThe code passed a peer-supplied length prefix straight to make, so the peer chose the allocation size: a few bytes on the wire bought 900 MB of heap. Validate the declared length against a maximum before allocating.
solid answer
~40 sThat profile is the signature of a trusted length prefix. The read path takes a fixed-size header, decodes a length, and does `body := make([]byte, size)` before a single byte of the body has arrived, so a hostile or simply buggy peer spends four bytes and makes you spend a gigabyte. The slice is live while you wait for data that may never come, which is why it tops the inuse_space profile. Hardening is layered: decide a maximum frame size, reject the frame with a protocol error if `size` exceeds it, and only then allocate and `io.ReadFull` the body. Wrap the connection in `io.LimitReader` as a backstop, and close the connection on a violation rather than skipping the declared bytes, since a lying length has desynchronised the framing.
code
go · 12 linesvar hdr [4]byte
if _, err := io.ReadFull(conn, hdr[:]); err != nil {
return err
}
size := binary.BigEndian.Uint32(hdr[:])
if size > maxFrameBytes {
return fmt.Errorf("frame of %d bytes exceeds limit %d", size, maxFrameBytes)
}
body := make([]byte, size)
if _, err := io.ReadFull(conn, body); err != nil {
return err
}go deeper
Take away the rule: never pass a number that came off the network straight into make. Compare it against a limit you chose first.
Explain the sequence precisely — fixed header with io.ReadFull, decode, compare against a maximum, then allocate and read the body — and why the allocation precedes any body bytes arriving.
Diagnose from the profile and harden in layers: the explicit check, io.LimitReader as a backstop, closing on violation, and a regression test that sends a lying header.
Own the numbers together: the per-message cap and the concurrency limit multiply, so state the worst-case memory the service can reach and defend it against the process memory budget.
## Reading the profile A heap profile's `inuse_space` view ranks allocation sites by bytes currently live. One `[]byte` at the top, far larger than anything else, allocated at the line where a frame body is created, is not a leak in the usual sense — it is one allocation that is exactly as large as it was asked to be. The question is who did the asking. In a length-prefixed protocol the shape is always the same: a fixed-size header carries the body length, and the reader does ```go size := binary.BigEndian.Uint32(hdr[:]) body := make([]byte, size) ``` At that moment the process has committed `size` bytes of heap on the strength of four bytes from a stranger. `0xFFFFFFFF` is a legal uint32. The peer does not even have to send a body: the allocation is already made, the slice is live while the read blocks, and the attack costs the sender essentially nothing. That asymmetry — four bytes in, gigabytes out — is what makes this a first-class availability bug rather than a mere sizing mistake. It is worth noting that this does not need an attacker. A version-skewed client with a different header layout, a byte-order mistake, or a corrupted frame produces exactly the same picture. ## The hardening, in order **1. Decide a maximum, and write it down.** Pick a `maxFrameBytes` your service is willing to hold, and make it part of the protocol's documented contract, not a magic number buried in the reader. **2. Validate before allocating.** The check must sit between decoding the length and calling `make`: ```go var hdr [4]byte if _, err := io.ReadFull(conn, hdr[:]); err != nil { return err } size := binary.BigEndian.Uint32(hdr[:]) if size > maxFrameBytes { return fmt.Errorf("frame of %d bytes exceeds limit %d", size, maxFrameBytes) } body := make([]byte, size) if _, err := io.ReadFull(conn, body); err != nil { return err } ``` Note what each piece does. `io.ReadFull` on the header handles short reads and distinguishes a clean close (`io.EOF`) from a header cut in half (`io.ErrUnexpectedEOF`). The comparison is the trust boundary. `io.ReadFull` on the body is bounded because *you* sized the slice. **3. Wrap the source as defence in depth.** Putting the connection behind `io.LimitReader` for the duration of a frame means that even if the check is wrong, one message cannot pull more than the ceiling. Remember that reaching that ceiling is reported as an ordinary end of stream, so the wrapper is a backstop, not the primary check. **4. Close the connection on violation.** The tempting alternative is to read and discard `size` bytes to "stay in sync". Do not. A peer that declared an impossible length has by definition broken the framing, so the bytes after that header are not a message you can skip past — you have no idea where the next header starts. Draining also hands the peer a free way to keep your goroutine and your bandwidth busy. Close it. **5. Bound concurrency too.** A per-frame cap is only half the arithmetic. Worst-case memory is roughly the cap times the number of frames that can be in flight simultaneously, which for a connection-per-goroutine server means the connection limit. A 4 MiB cap with ten thousand connections is 40 GB, so the two numbers have to be chosen together. ## Confirming the diagnosis - The heap profile at `inuse_space` names the allocating line directly; that is usually enough. - Compare `alloc_space` against `inuse_space`: a single huge live allocation looks very different from steady churn, and tells you this is one bad frame rather than a slow leak. - The goroutine profile shows the reader parked in a read while holding that slice, which is the "allocated and waiting" state. - Reproduce it deliberately: a five-byte test that sends a header declaring a gigantic body and nothing else should now produce a rejection and a closed connection, and that test belongs in the suite permanently. ## The general principle Any value that determines an allocation size must be validated before it is used, and "validated" means compared against a bound you chose, not merely checked for being non-zero. In Go the moment is unusually easy to spot in review, because it is always a `make` whose argument traces back to bytes from the network.
- Why not read and discard the declared bytes instead of closing the connection?Because a peer that declared an impossible length has broken the framing: you do not know where the next header begins, so skipping is guesswork. Draining also lets the peer keep a goroutine and bandwidth busy for free. Closing is cheap, unambiguous, and gives the client an honest signal that its frames are invalid.
- The cap is 4 MiB and the profile still shows the heap exhausted. What did the sizing miss?Concurrency. Worst-case memory is the per-frame cap times the number of frames in flight, so with a goroutine per connection the connection limit multiplies the cap. Four MiB across ten thousand connections is forty gigabytes. The per-message cap and the concurrency limit have to be chosen as one number, not separately.
- Where does io.LimitReader fit if the length is already validated?As a backstop. Wrapping the connection for the duration of a frame means a bug in the comparison, a new code path, or a second header format cannot become unbounded again. It is deliberately not the primary check, because hitting its ceiling is reported as an ordinary end of stream rather than as a rejection.
- Does this only happen with a malicious peer?No. A version-skewed client with a different header layout, a byte-order mismatch, or a corrupted frame all produce the same enormous declared length. That is a good argument for treating the check as a correctness invariant rather than a security feature, so it stays in place even on links you consider trusted.
saying these in an interview costs you the question
- Trusts a length prefix because the peer is on an internal network
- Adds the size check after the make call rather than before it
- Reads and discards the declared bytes to stay in sync after a violation
- Sizes the per-message cap without multiplying by concurrent connections
- Blames the garbage collector for a slice that is still live
- Assumes a read deadline prevents the allocation