What does gzip.Reader.Multistream(false) change when reading a file of concatenated gzip members?
answer
- cat two .gz files together
- the default hides the seam
- each member has its own header
- EOF at the boundary, then Reset
- the source must give up single bytes
basics
~10 sBy default Go's gzip.Reader reads concatenated gzip members as one continuous stream. Multistream(false) makes Read stop with io.EOF at the end of each member, so you can handle members individually and advance with Reset.
solid answer
~50 sA gzip file may legally hold several complete members back to back — `cat a.gz b.gz > both.gz` is a valid gzip file — and Go's reader hides that by default: you read straight through and never see the boundary. `zr.Multistream(false)` turns the boundary into an `io.EOF` from `Read`, which is what you need when each member is a separate record, has its own `Header.Name` or `ModTime`, or is followed by non-gzip data in the same file. To advance you call `zr.Reset(r)`, which starts the next member and returns `io.EOF` if there is none; `Reset` also re-enables multistream mode, so call `Multistream(false)` again each round. One requirement is easy to miss: the underlying reader must implement `io.ByteReader`, otherwise the reader may over-read and cannot leave the position exactly after the member — wrapping the file in a `bufio.Reader` satisfies it.
code
go · 20 linesbr := bufio.NewReader(f) // must be an io.ByteReader
zr, err := gzip.NewReader(br)
if err != nil {
return err
}
defer zr.Close()
for {
zr.Multistream(false)
name := zr.Header.Name
if _, err := io.Copy(dst, zr); err != nil {
return err
}
record(name)
if err := zr.Reset(br); errors.Is(err, io.EOF) {
return nil // no further member
} else if err != nil {
return err
}
}go deeper
Know that a gzip file may contain several members back to back and that Go reads them as one stream unless told otherwise.
Explain what the flag changes at the boundary: Read returns io.EOF per member, and Reset starts the next one or reports io.EOF when the file is done.
Show the full loop, including re-arming the flag after every Reset and the io.ByteReader requirement, and say when per-member reading is genuinely needed rather than default.
Decide whether per-member metadata belongs in the archive at all, or whether an explicit container with its own index is the honest format for what downstream consumers must do.
## Concatenation is part of the format RFC 1952 permits a gzip file to consist of a sequence of complete members, each with its own header, DEFLATE body and trailer. Decompressing the file means decompressing each member and concatenating the results. This is why appending to a `.gz` archive works at all: a log rotator can compress each rotation separately and append the bytes, and every standard tool reads the result as one logical file. Go implements that default faithfully. `gzip.NewReader` gives you a reader that walks from one member into the next without telling you, and returns `io.EOF` only when the last member ends. ## What the flag changes `(*gzip.Reader).Multistream(ok bool)` selects between the two behaviours. With `false`, `Read` returns `io.EOF` at the end of the *current* member. The reader is now a per-member reader, and the file becomes a sequence you iterate. You advance with `Reset(r)`, which re-initialises the reader over the same underlying source. It returns `io.EOF` when there is no further member, which is the loop's exit condition, and any other error means a malformed next header. Crucially, `Reset` restores the default multistream behaviour, so a loop must set the flag again on every iteration. ```go br := bufio.NewReader(f) // gzip.Reader needs an io.ByteReader here zr, err := gzip.NewReader(br) if err != nil { return err } defer zr.Close() for { zr.Multistream(false) name := zr.Header.Name if _, err := io.Copy(dst, zr); err != nil { // stops at this member's end return err } use(name) if err := zr.Reset(br); errors.Is(err, io.EOF) { return nil } else if err != nil { return err } } ``` ## The io.ByteReader requirement The documentation is explicit that in this mode the underlying reader must implement `io.ByteReader` for the position to be left just after the gzip stream. The reason is buffering: a decompressor reading in chunks would normally consume past the member's last byte, and then the caller could not pick up the next member — or the non-gzip bytes that follow — from the right offset. Byte-at-a-time access at the boundary makes the position exact. An `*os.File` does not implement `io.ByteReader`; `*bufio.Reader` and `*bytes.Reader` do. Wrapping in `bufio.NewReader` is the normal fix, and you must then pass that same `*bufio.Reader` to `Reset`, not the original file, or you will restart from a different position. ## When you actually need it - Each member is one rotation, and you want its `Header.Name` and `ModTime` to route or timestamp what is inside. Those fields belong to a member, so reading transparently across boundaries loses all of them but the first. - A container format embeds gzip members among other data — an index, a manifest, a length-prefixed frame — and the reader must stop exactly where the member does. - You are validating an archive member by member, so a corrupt member can be reported individually instead of failing the whole file. ## When you do not If all you want is the bytes, leave the default alone. Reading transparently is simpler, verifies every member's checksum on the way through, and needs no `io.ByteReader`. Reaching for `Multistream(false)` because the file 'has several parts' and then re-implementing what the default already does is a common over-engineering. ## The failure it prevents The operational version of this is a shipper that appends a member per rotation and a consumer that reads only `Header.Name` from the file it opened. With the default reader that name is the *first* member's, so every batch is attributed to the earliest rotation and the mislabelling only surfaces when someone at 3 a.m. is trying to work out which hour a line came from. Switching to per-member reading is what makes the metadata usable.
- Why must you call Multistream(false) again after every Reset?`Reset` returns the reader to the state it would have had from `gzip.NewReader`, and that default is multistream mode. Forgetting the second call means the very next `io.Copy` runs transparently to the end of the file, so you process member one individually and everything after it as a single blob — a bug that looks like 'only the first boundary works'.
- What breaks if the reader you pass to gzip.NewReader is not an io.ByteReader?In per-member mode the decompressor may consume bytes past the end of the member while filling its own buffer, so the source is no longer positioned at the next header. `Reset` then fails or misreads, and any non-gzip data following the member is unreachable. Wrapping the source in a `bufio.Reader` and passing that same value to `Reset` fixes it.
saying these in an interview costs you the question
- Thinks concatenated gzip members are an invalid file
- Expects the default reader to stop at each member
- Calls Multistream(false) once and loops with Reset
- Passes the raw os.File where an io.ByteReader is required
- Reads Header.Name once and applies it to the whole file