skip to content

When do you reach for bufio.Reader's ReadString or Peek instead of bufio.Scanner?

level: middleimportance: should knowfreq 44%

answer

  1. one is the convenience layer over the other
  2. who keeps the newline byte
  3. look at the next bytes, do not eat them
  4. no ceiling means you own the memory
  5. a partial tail arrives with the error

basics

~10 s

Use bufio.Scanner for ordinary token-at-a-time reading with a size cap. Use bufio.Reader when you need the delimiter kept, no token ceiling, a look at the next bytes without consuming them, or rune-by-rune decoding.

solid answer

~50 s

`bufio.Scanner` is the convenience layer: it hands you one token per `Scan`, strips the line ending, and refuses tokens above its maximum size. `bufio.Reader` is the lower level, and you drop to it when the scanner's shape does not fit. `ReadString('\n')` returns the line **including** the delimiter and has no token ceiling, so an unterminated final line comes back as data plus a non-nil error - you must handle a non-empty result with an error rather than discarding it. `Peek(n)` shows you the next n bytes without consuming them, which is how you sniff and skip a UTF-8 BOM or a header before deciding how to parse. `ReadRune` reassembles a multibyte character even when its bytes straddle a buffer refill. `Discard` and `UnreadByte` give you the pushback a tokenizer needs. The price is that framing, size limits and the end-of-input case are now yours to write.

code

go · 5 lines
go
br := bufio.NewReader(f)
if b, err := br.Peek(3); err == nil && bytes.Equal(b, []byte("\xef\xbb\xbf")) {
	br.Discard(3) // Peek consumed nothing; Discard skips the BOM
}
line, err := br.ReadString('\n') // line still ends with '\n'

go deeper

for a junior

Know that both types read text but only one hands you clean tokens for free. Be able to say that Peek looks ahead without consuming and that ReadString keeps the delimiter it stopped on.

for a middle

Explain the trade concretely: no token ceiling, the delimiter retained, the partial tail returned alongside the error, and the framing and size limit becoming your responsibility.

for a senior

Show the judgment - which reader you pick for untrusted or huge records, how you bound memory once the library stops doing it, and how look-ahead lets you choose a parse before committing to it.

for a principal

Argue the house style: when a codebase should standardise on the simple tokenizer and when a hand-rolled reader loop is worth the review burden, given that every such loop re-implements framing and end-of-input handling.

## Two layers, not two options `bufio.Scanner` is built **on top of** the same idea as `bufio.Reader`: read into a buffer, hand out pieces. The difference is who decides where a piece ends and what happens when it does not fit. `bufio.Scanner` answers both questions for you. A split function - `bufio.ScanLines` by default - decides the boundary, the line terminator is stripped, and a token above the maximum size stops the scanner. That is exactly what you want for the ninety per cent case of "read a text file line by line". `bufio.Reader` answers neither. It is a buffer with reading primitives on it, and every policy is yours. ## What bufio.Reader gives you that the scanner does not **The delimiter is kept.** `ReadString('\n')` returns everything up to and including the newline. When you are re-emitting records verbatim, or a protocol distinguishes a terminated record from a truncated one, that byte matters. The scanner has already thrown it away, along with an optional preceding carriage return. **There is no token ceiling.** A 900 MB line is returned as a 900 MB string. That is either the feature you needed or a memory hazard you now own; there is no library-supplied cap to lean on. **The end-of-input case is explicit.** If `ReadString` hits the end of the stream before finding the delimiter, it returns the bytes it did read **and** a non-nil error. A final line with no trailing newline arrives that way, so the loop must be written to process a non-empty result even when the error is set: ```go for { line, err := br.ReadString('\n') if line != "" { record(line) } if err != nil { break } } ``` Discarding the data whenever the error is non-nil is the classic way to lose the last record of every file. **Look-ahead without consumption.** `Peek(n)` returns the next n bytes and leaves the reader's position untouched. That is how you decide *how* to read before you start reading: ```go if b, err := br.Peek(3); err == nil && bytes.Equal(b, []byte("\xef\xbb\xbf")) { br.Discard(3) } ``` A UTF-8 byte order mark at the head of a log file otherwise becomes part of your first field, and the mismatch is invisible in most terminals. `Peek` is bounded by the reader's buffer size - 4096 bytes by default, or whatever `bufio.NewReaderSize` was given - and asking for more than that returns `bufio.ErrBufferFull`. Near the end of input it returns fewer bytes together with the error explaining why. **Correct rune decoding across refills.** `ReadRune` returns one rune and its size in bytes, filling the buffer as needed so a UTF-8 sequence split across two underlying reads is still decoded as one character. Reading raw bytes into a fixed-size slice and converting to a string does not have that property: a multibyte character straddling the slice boundary becomes replacement characters at both ends. If you are counting or classifying characters rather than lines, that difference is the whole reason to use a `bufio.Reader`. **Pushback.** `UnreadByte` and `UnreadRune` step back one item, and `Discard(n)` skips ahead without allocating. Hand-written tokenizers need both. ## What you give up Everything the scanner was doing for you: - **Framing.** You write the loop and the delimiter handling, including CRLF if you care. - **A size limit.** You have to bound the record yourself if the input is not trusted. - **The clean-exit convention.** There is no single `Err()` after the loop; the error comes back from each call, and you must decide which errors mean "stream ended" and which mean "failure". ## A practical rule Start with `bufio.Scanner`. Move to `bufio.Reader` when one of these is true: you need the delimiter, you need look-ahead before choosing a parse, records may legitimately be huge and you want to stream them rather than materialise them, you are decoding runes rather than lines, or you are writing a tokenizer that needs pushback. Mixing the two on the same stream is fine only if you stop using one before starting the other - both hold their own buffer, and whatever one has already read is not available to the other.

  • Why does bufio.Reader.ReadRune work when a UTF-8 sequence straddles a buffer refill?
    Because the reader refills its buffer until it holds a complete encoding before decoding, so the character is reassembled across the boundary. Reading raw bytes into a fixed slice and converting to a string has no such logic: a multibyte character cut in half becomes replacement characters on both sides of the split.
  • What bounds how many bytes bufio.Reader.Peek can return?
    The reader's own buffer, 4096 bytes by default or whatever bufio.NewReaderSize was given. Asking for more than the buffer size returns bufio.ErrBufferFull immediately. Close to the end of the stream Peek returns the bytes that remain plus the error explaining why it returned fewer than requested.
  • Can you switch from a bufio.Reader to a bufio.Scanner mid-stream?
    Only if you hand the scanner the bufio.Reader itself rather than the original source. Each wrapper holds its own buffer, so bytes already pulled into the reader are invisible to a scanner built on the underlying stream, and you silently lose whatever was buffered.

The scanner is a bread slicer: uniform pieces, crusts trimmed, and it jams on a loaf that is too big. The reader is the knife - nothing is trimmed, nothing jams, and every cut is your decision.

saying these in an interview costs you the question

  • Treats bufio.Scanner and bufio.Reader as interchangeable
  • Thinks ReadString strips the delimiter it searched for
  • Believes Peek advances the read position
  • Discards the returned data whenever the error is non-nil
  • Assumes bufio.Reader caps line length for untrusted input