How does gzip.Writer's Flush differ from its Close, and when is calling Flush worth the cost?
answer
- a checkpoint, not an ending
- no trailer, stream keeps going
- the block boundary is the cost
- there is another buffer below you
- latency traded against ratio
basics
~20 sFlush pushes everything compressed so far to the underlying writer and ends the deflate block, so a reader can decompress it immediately; no trailer is written and the stream continues. Close finishes the member. Each flush costs ratio.
solid answer
~50 s`(*gzip.Writer).Flush` ends the current DEFLATE block with an empty stored block, which makes the output land on a byte boundary and lets a reader on the far end reconstruct every byte written so far. The member is not finished: no CRC32, no length, and you keep writing. `Close` does the final flush *and* appends the trailer, ending the member for good. Flush earns its keep when a consumer is waiting on a live stream and latency matters more than a few percent of ratio — an upload or a socket someone is tailing. It costs you: closing a block early throws away the compressor's ability to code across the boundary, so flushing per log line can leave output barely smaller than the input. It also only flushes the gzip layer; a `bufio.Writer` or the socket underneath still has its own buffer to flush.
code
go · 19 linesbw := bufio.NewWriter(conn)
zw := gzip.NewWriter(bw)
for batch := range batches {
if _, err := zw.Write(batch); err != nil {
return err
}
// Checkpoint: the peer can decompress everything so far.
if err := zw.Flush(); err != nil {
return err
}
if err := bw.Flush(); err != nil {
return err
}
}
if err := zw.Close(); err != nil {
return err
}
return bw.Flush()go deeper
Know that Flush lets a reader see data mid-stream while Close ends the stream, and that only Close makes the file complete.
Be ready to explain the sync flush: the block is terminated, the output aligns to a byte boundary, and the compressor loses cross-block context, which is where the ratio goes.
Show that you pick flush granularity from the consumer's latency need and measure the resulting size, and that you always flush the layer beneath the compressor too.
Frame it as a delivery-latency versus bandwidth tradeoff with a named owner, and prefer whole-member boundaries when downstream systems need independently valid artefacts.
## Three different meanings of 'the data is out' When you write through `compress/gzip` there are at least three buffers between your bytes and the destination: the DEFLATE compressor's window inside the `gzip.Writer`, any `bufio.Writer` you wrapped around the destination, and the destination's own buffering (the OS socket or file buffer). `Flush` and `Close` speak only to the first of these. ## What Flush does `(*gzip.Writer).Flush` flushes any pending compressed data to the underlying writer and does not return until that write has happened. Mechanically it performs a DEFLATE *sync flush*: it terminates the current block and emits an empty stored block, which forces the output to end on a byte boundary. That matters because DEFLATE is a bit stream — without the alignment, the last partial byte could not be handed over, and a decompressor could not produce the tail of your data. After a `Flush`, a reader that has received everything written so far can decompress all of it. The stream is still open; the next `Write` starts a fresh block. ## What Close does `Close` flushes, terminates the DEFLATE stream properly, and then appends the gzip trailer: the CRC32 of the uncompressed bytes and the uncompressed length mod 2^32. Only after `Close` is the member complete and verifiable. `Close` does not close the underlying `io.Writer`; that stays your job. So the difference is not 'Close flushes harder'. The difference is that `Flush` is a checkpoint inside a stream and `Close` is the end of the stream. ## The cost of a flush DEFLATE compresses well because it can refer backwards across a long window and because Huffman codes are chosen per block over as much data as possible. Ending a block early gives up both: the code tables are rebuilt for the next block, and the empty stored block adds a handful of bytes of pure overhead. On small writes this is dramatic. A log-shipping agent that flushes after every line of a few dozen bytes can emit output larger than the input, because each line becomes its own block with its own overhead. Flushing once per batch, or once per few hundred kilobytes, gets nearly all of the latency benefit at a fraction of the ratio cost. The honest framing for an interview: `Flush` trades compression ratio for delivery latency, and the right flush granularity is a batching decision, not a per-write reflex. ## The layer below A very common bug: the code flushes the compressor and assumes the bytes are on the socket. ```go bw := bufio.NewWriter(conn) zw := gzip.NewWriter(bw) // ... writes ... if err := zw.Flush(); err != nil { return err } if err := bw.Flush(); err != nil { return err } // without this, bytes sit here ``` `gzip.Writer.Flush` writes into `bw`, and `bw` holds them until its own buffer fills or it is flushed. Whatever the reader is waiting for, it does not arrive. The same applies at shutdown, where the order is `zw.Close()`, then `bw.Flush()`, then close the connection or file. ## When you actually want Flush - A long-lived upload or streaming endpoint where the consumer must see progress before the stream ends. - A framed protocol of your own where each frame must be independently readable at the point it was written. - Interactive tailing, where an operator at 3 a.m. is watching output and a five-minute buffering delay is indistinguishable from a hang. When you are compressing a rotated file to disk and uploading it afterwards, you never need `Flush` at all — a single `Close` at the end is both correct and optimal, because nothing reads the file until it is complete. ## A related trap Do not reach for `Flush` as a way to 'make sure the data is safe' before a crash. It does not fsync anything, and it does not make a partial file valid: without a trailer, a reader still ends in `io.ErrUnexpectedEOF`. If you want restartable output, write whole members and concatenate them — a gzip file may contain any number of complete members, and Go's reader will read them as one stream by default.
- Does calling Flush make a half-written gzip file readable by a tool that expects a complete file?No. A flush only guarantees that everything written so far can be decompressed by a reader consuming the stream; the member has no trailer, so anything that reads to the end still hits `io.ErrUnexpectedEOF` and no checksum is verified. If you need a valid artefact at intervals, close one member and start another rather than flushing.
- Why can flushing after every short write produce output bigger than the input?Each flush ends the deflate block, so the compressor rebuilds its Huffman tables and loses the chance to code across the boundary, and the empty stored block adds fixed overhead. With writes of a few dozen bytes, that overhead exceeds the savings and the stream grows. Batch writes and flush per batch instead.
saying these in an interview costs you the question
- Says Flush and Close do the same thing
- Thinks Flush writes the CRC32 trailer
- Flushes after every line and never measures the ratio
- Assumes flushing the compressor flushes the bufio.Writer below
- Treats Flush as a durability guarantee like fsync