skip to content

A bufio.Writer around your output stream delayed records by minutes - when does buffering cost more than it saves?

level: seniorimportance: nice to knowfreq 30%

answer

  1. fewer syscalls, paid for in latency
  2. a slow stream is the bad case
  3. multiply the default by every connection
  4. the sink may already be batching
  5. a large write skips the buffer entirely

basics

~20 s

Buffering trades latency, memory and a crash-loss window for fewer syscalls. It pays when writes are small and frequent; it costs more than it saves on slow streams, on many concurrent streams, and over sinks that already batch.

solid answer

~50 s

A `bufio.Writer` only reaches the underlying stream when its buffer fills or `Flush` is called, so on a stream producing a few hundred bytes a minute a record can sit in memory for a very long time - the consumer sees nothing, and a kill loses it. There is a memory cost too: 4 KiB per writer and per reader by default, which at fifty thousand concurrent streams is hundreds of megabytes of buffers. A write error now surfaces at `Flush`, far from the record that caused it. The win is real only when you would otherwise pay a syscall per small write; over an in-memory sink, over one that already batches, or when every write already exceeds the buffer - an oversized write skips an empty buffer entirely - you bought an allocation and nothing else. Fix a slow stream with a timed flush under the writer's lock, not a bigger buffer.

go deeper

for a junior

Know the basic trade in one line: buffering swaps immediacy for fewer, larger writes. If a consumer needs the data now, something has to flush.

for a middle

Be able to name the concrete costs - the default buffer size multiplied by the number of streams, the delay before a record becomes visible, and the error appearing only at flush time.

for a senior

Show that you would measure rather than assume, and that you can spot the no-win cases: an in-memory sink, a sink that already batches, and writes already larger than the buffer. Explain how a timed flush is made safe.

for a principal

Frame the buffer as a policy the whole fleet inherits: how much memory per stream the service can commit, how stale downstream data may be, how many records a hard kill may lose, and where those defaults are enforced.

## What buffering actually buys Every `Write` on an `*os.File` or a network connection is a syscall. A log-ingest agent emitting one normalised record per line pays one syscall per record; wrapped in a `bufio.Writer`, seventy-odd records leave in one. When the writes are small and frequent, that is a large win and it is why the wrapper exists. The mistake is to treat it as free and apply it everywhere. It is a trade with four distinct prices. ## Price 1: latency and visibility Buffered bytes are invisible. Nothing downstream - a consumer tailing the output, an operator running a diagnostic, a health check - can see a record until the buffer fills or someone flushes. On a stream producing a few hundred bytes a minute, a 4 KiB buffer means minutes of delay, and on a nearly idle stream it can mean forever. The wrong fix is a bigger buffer, which makes the delay longer. The right fix is a **time bound**: flush on a ticker, or flush at the end of each batch you were going to complete anyway. Note that `bufio.Writer` has no internal locking, so a flush from a timer goroutine must take the same mutex the writing path takes - otherwise you have swapped a latency bug for a data race. ## Price 2: memory per stream The default buffer is 4096 bytes for `bufio.NewReader` and 4096 for `bufio.NewWriter`. Wrapping both ends of a connection is therefore 8 KiB of long-lived allocation per connection. At fifty thousand concurrent connections that is roughly 400 MB that exists only to reduce syscalls - frequently the single largest line in the heap profile of a connection-heavy service. The mitigations are smaller buffers via `bufio.NewReaderSize` and `bufio.NewWriterSize`, pooling and reusing writers with `Reset`, or simply not buffering the direction that already writes in large chunks. ## Price 3: errors move away from their cause Because the underlying write is deferred, so is its failure. A disk-full or broken-pipe error appears from `Flush`, or from a much later `Write` once the writer has recorded it, not from the call that produced the offending record. You keep correctness only if you check `Flush`'s error - and you lose attribution regardless. Flushing more often narrows the window at the cost of the syscall savings you were buying. ## Price 4: the crash-loss window Whatever sits in the buffer is gone if the process is killed. The buffer size is, quite literally, the number of bytes of output you have decided you can afford to lose. That is a product decision dressed as a constructor argument, and it should be made deliberately rather than inherited from the default. (Even `Flush` only reaches the kernel; surviving a machine failure needs `Sync` on the file, which is far more expensive again.) ## When there is no win at all - **The sink is already in memory.** Wrapping a `bytes.Buffer` or a `strings.Builder` in a `bufio.Writer` adds a copy and an allocation to save syscalls that were never happening. - **The sink already batches.** Layering a second `bufio.Writer` over one, or over a component that assembles its own frames, buys nothing and doubles the flush obligations - now two things must be flushed in the right order. - **The writes are already large.** `bufio.Writer` has a fast path: when its buffer is empty and the incoming write is bigger than the free space, it hands the slice straight to the underlying writer instead of copying it. If every record is 64 KiB, the buffer is never doing work; you paid for the allocation and the extra indirection. - **The reader is a `bytes.Reader` or `strings.Reader`.** Same argument on the read side: a `bufio.Reader` over data already in memory adds a copy per read. ## How to decide Measure rather than assume. A benchmark that writes N realistic records with and without the wrapper, run with `-benchmem`, shows both the time and the allocation difference; on a syscall-bound path the gap is dramatic and on an in-memory path it is negative. Then apply the shape that fits: buffer where writes are small and frequent, bound the buffer's age with a timed flush where the stream is slow, size it consciously where streams are numerous, and skip it entirely where the sink already does the batching.

  • How do you stop a low-rate buffered stream from sitting on records?
    Bound the buffer's age rather than its size: flush at the end of each batch, or from a ticker goroutine. Because bufio.Writer has no internal locking, the timed flush and the writing path must share one mutex. Choose the interval from how stale the consumer tolerates its data being, not from throughput.
  • What does a bufio.Writer with an empty buffer do with a write larger than that buffer?
    It skips the buffer and passes the slice straight to the underlying writer, avoiding a pointless copy. That is why wrapping a stream whose every write is already large gains nothing: the buffer sits idle and you have paid only for the allocation and an extra layer.
  • How would you show a colleague that adding the wrapper was worth it?
    Benchmark the realistic write pattern both ways with -benchmem: the same N records, same sizes, same sink. On a per-line write to a file the buffered version wins by a large multiple because it collapses syscalls; over an in-memory sink it loses, because there were no syscalls to collapse.

It is carpooling for syscalls: excellent when riders arrive constantly, terrible when one rider arrives an hour and has to wait for the van to fill.

saying these in an interview costs you the question

  • Says buffering is always faster
  • Wraps every writer in bufio.Writer as a default habit
  • Fixes a latency complaint by enlarging the buffer
  • Assumes buffered records survive a hard kill
  • Ignores per-connection buffer memory at high connection counts
  • Flushes from a timer without holding the writer's lock