skip to content

In Go, how do you cap how much data a gzip.Reader can produce from an untrusted .gz file?

level: seniorimportance: should knowfreq 40%

answer

  1. the danger is the ratio, not the input
  2. wrap the thing that expands
  3. the limit reports a plain EOF
  4. ask for one byte more than you allow
  5. the size field is written by the attacker

basics

~20 s

Wrap the gzip.Reader, not the compressed input, in io.LimitReader with your budget plus one byte, then reject any stream that reaches that extra byte. io.LimitReader signals a plain io.EOF at its limit, so the caller must check the count.

solid answer

~50 s

The limit belongs on the decompressed side: `io.LimitReader(zr, max+1)` around the `*gzip.Reader`, because the danger is the expansion ratio and a few kilobytes of input can produce gigabytes of output. The `+1` matters because `io.LimitReader` does not report an error when it runs out — it just returns `io.EOF`, which is indistinguishable from a genuinely finished stream. So you copy, then check whether the byte count exceeded your budget and fail explicitly if it did. Do not use the uncompressed-size field in the gzip trailer as a guard: it is written by whoever produced the file, it is only accurate mod 2^32, and it sits at the *end* of the stream anyway. Also cap the work, not just the bytes, if you decompress concurrently — a bounded reader per stream still lets many streams run at once.

code

go · 15 lines
go
const maxDecompressed = 64 << 20 // 64 MiB

zr, err := gzip.NewReader(src)
if err != nil {
	return err
}
defer zr.Close()

n, err := io.Copy(dst, io.LimitReader(zr, maxDecompressed+1))
if err != nil {
	return err
}
if n > maxDecompressed {
	return errors.New("gzip stream exceeds the 64 MiB decompressed limit")
}

go deeper

for a junior

Know that decompressing input you did not produce can expand enormously, and that io.ReadAll over a gzip.Reader has no ceiling of its own.

for a middle

Be ready to place io.LimitReader correctly around the gzip.Reader and to explain why a limit on the compressed bytes proves nothing about the output.

for a senior

Demonstrate the detectable-overflow idiom with the extra byte, refuse the trailer's size field as a guard, and bound concurrency alongside per-stream size.

for a principal

Own where the ceiling is configured, what the rejection tells the on-call engineer, and how the limit is chosen from real producer sizes rather than a number someone liked.

## The threat DEFLATE can compress a highly repetitive input by a ratio of roughly a thousand to one, and nested or crafted streams do far better. A few hundred kilobytes of `.gz` that arrived from somewhere you do not control can expand into many gigabytes. If your code does `io.ReadAll(zr)` or `io.Copy(buf, zr)` with no ceiling, the process allocates until the kernel kills it — or, on a node with a memory limit, takes the workload next to it down with it. This is a decompression bomb, and the defence is a hard ceiling on **output** bytes. ## Why the limit goes on the outside of the decompressor Capping the compressed input does not bound the output at all; that is the whole point of the attack. The reader you must limit is the `*gzip.Reader` itself, because that is the one whose `Read` produces expanded bytes: ```go const maxDecompressed = 64 << 20 // 64 MiB zr, err := gzip.NewReader(src) if err != nil { return err } defer zr.Close() n, err := io.Copy(dst, io.LimitReader(zr, maxDecompressed+1)) if err != nil { return err } if n > maxDecompressed { return errors.New("gzip stream exceeds the 64 MiB decompressed limit") } ``` A limited reader stops the copy at the ceiling, so memory and time stay bounded whatever the input claims. ## The +1, and why it is not a stylistic detail `io.LimitReader(r, n)` returns a reader that yields at most `n` bytes and then reports `io.EOF`. It does **not** return a distinguishable error. If you limit at exactly your budget and the copy stops there, you cannot tell 'the file happened to be exactly 64 MiB and is fine' from 'the file is 40 GiB and I truncated it', and silently truncating attacker-supplied data is usually worse than rejecting it — you end up storing a half-parsed record and calling it success. Reading `max+1` makes the test unambiguous: any count above `max` means there was more. `io.CopyN` with an explicit `io.EOF` check is an equivalent formulation. ## What not to trust The gzip trailer carries ISIZE, the uncompressed length modulo 2^32. It is tempting as a pre-flight check and it is useless as one: the producer writes it, so an attacker sets it to whatever is convenient; it wraps at 4 GiB, so a 5 GiB stream can declare 1 GiB; and it is physically at the end of the stream, so you would have to decompress everything to read it. Go's `gzip.Reader` does verify it — together with the CRC32 — but only when the caller reads all the way to `io.EOF`, and by then the damage would already be done. Its job is integrity, not admission control. The same reasoning applies to any size hint in an envelope around the data. ## Choosing the ceiling Pick it from what the legitimate producer actually emits, with headroom, and make it configuration rather than a constant so it can be raised at 3 a.m. without a release. Log the rejection with the stream's identity and the byte count so the on-call engineer can tell a bomb from a genuinely oversized batch — those two need opposite responses, and an error that says only 'limit exceeded' forces a guess. ## Related hardening in the same code path - Close the `gzip.Reader`. It does not close the underlying reader, but it releases the decompressor state, and pairing it with a fully consumed stream is what lets the checksum verification actually run. - If the file legitimately contains several concatenated members, the limit must cover the whole thing, since Go's reader concatenates them transparently by default. - Bound concurrency as well as size. A 64 MiB ceiling per stream and unbounded parallel decompressions is still an unbounded memory footprint; the two limits are separate decisions. - Stream to a file or a bounded consumer rather than into memory where you can, so the ceiling protects against latency and disk as well as heap. ## The one-sentence version Limit the reader that produces expanded bytes, budget one byte more than you will accept so the ceiling is detectable, and never let the stream describe its own size.

  • Why is io.LimitReader's behaviour at the limit awkward, and what is the alternative?
    It reports `io.EOF`, exactly like a stream that finished naturally, so on its own it silently truncates. Reading `max+1` and comparing the returned count makes the overflow explicit. `io.CopyN(dst, zr, max+1)` is equivalent: a `nil` error means there were more than `max` bytes, and `io.EOF` means the stream ended within budget.
  • Does closing a gzip.Reader verify the checksum for you?
    No. `(*gzip.Reader).Close` releases the decompressor and does not touch the underlying reader; the CRC32 and length are checked when `Read` reaches the end of the member. A reader that stops early — including one stopped by your limit — never runs that check, so a rejected stream tells you nothing about its integrity.
  • Your ceiling is 64 MiB per stream and the process still OOMs. What did you miss?
    Concurrency. Per-stream bounds multiply by however many streams you decompress at once, so the real footprint is limit times parallelism. Bound the number of concurrent decompressions with a semaphore or a fixed set of worker goroutines, and size the two limits together against the memory you actually have.

A size limit on the compressed file is a weight limit on a folded parachute: it tells you nothing about how much space it takes once it opens.

saying these in an interview costs you the question

  • Limits the compressed input instead of the output
  • Trusts the uncompressed-size field in the gzip trailer
  • Uses io.LimitReader without noticing it returns plain EOF
  • Calls io.ReadAll on an untrusted gzip.Reader
  • Bounds each stream but not the number of concurrent streams