skip to content

A bufio.Scanner loop stops partway through a log file with no error - what limit causes that?

level: middleimportance: should knowfreq 52%

answer

  1. the loop exit looks like a clean end
  2. a per-token size ceiling exists
  3. 64 KiB unless you say otherwise
  4. one method after the loop tells them apart
  5. Buffer must come before the first Scan

basics

~20 s

A line longer than the scanner's maximum token size, 64 KiB by default, makes Scan return false with bufio.ErrTooLong stored in Err. A loop that never calls Err after Scan sees that failure as a clean end of input.

solid answer

~40 s

`bufio.Scanner` grows its buffer only up to `bufio.MaxScanTokenSize`, which is 64 KiB. Hit a longer line and `Scan` returns `false` and records `bufio.ErrTooLong`. Because `for sc.Scan()` exits the same way it exits at end of input, a loop that never checks `sc.Err()` afterwards reports success and quietly drops the rest of the file - an hour of records vanishing with no log line and no non-zero exit. Two fixes go together: always check `sc.Err()` after the loop, and if long records are legitimate, call `sc.Buffer(make([]byte, 0, 64*1024), maxYouCanAfford)` before the first `Scan` - it panics if called later. Do not set the cap to something enormous; that just converts a corrupt line into a large allocation.

code

go · 8 lines
go
sc := bufio.NewScanner(f)
sc.Buffer(make([]byte, 0, 64*1024), 4*1024*1024) // start 64 KiB, cap 4 MiB
for sc.Scan() {
	ingest(sc.Bytes()) // valid only until the next Scan
}
if err := sc.Err(); err != nil {
	return fmt.Errorf("tail %s: %w", path, err) // ErrTooLong lands here
}

go deeper

for a junior

Remember the two-line habit: check the scanner's Err after the loop, every time. Know that a very long line is a real failure mode and not something the library hides from you.

for a middle

Explain the mechanics you are expected to own here: a 64 KiB default maximum token, the buffer growing up to it, Scan returning false with ErrTooLong, and Buffer having to be called before the first Scan.

for a senior

Talk about how you would have caught this in production - the missing-records symptom, the test that feeds an over-long line, and the deliberate choice between raising the cap, skipping the record, or failing loudly.

for a principal

Own the cap as a policy question: how much memory per stream the ingest fleet can commit, whether an over-long record should be dropped or should stop the pipeline, and how that decision is enforced consistently rather than per author.

## The shape of the bug A log-ingest agent tails a newline-delimited file: ```go sc := bufio.NewScanner(f) for sc.Scan() { ingest(sc.Bytes()) } return nil ``` It runs for weeks, then one day an hour of records is missing downstream and the agent logged nothing. What happened is that one process upstream wrote a single enormous line - a stack dump, a base64 blob, a JSON document with an embedded payload - and the scanner stopped there. ## Why a bufio.Scanner has a limit at all `bufio.Scanner` is a tokenizer. It reads into an internal buffer and hands out one token per `Scan` - by default a line, using the `bufio.ScanLines` split function. To return a whole token it must hold that whole token in memory, so an unbounded scanner would let a single line of input dictate the process's memory use. The library refuses to do that: the buffer starts small, grows as needed by doubling, and stops at `bufio.MaxScanTokenSize`, a constant equal to 64 * 1024. When a token would exceed that limit, `Scan` returns `false` and stores `bufio.ErrTooLong` ("token too long"). It does **not** panic, does **not** truncate the token and hand you a prefix, and does **not** skip the line and continue. The scanner is finished. ## Why nobody notices `Scan` returning `false` is the loop's normal exit condition. At a clean end of input `Scan` also returns `false`, and in that case `Err()` is `nil`. The only thing distinguishing "the file ended" from "I gave up" is the value of `Err()`, and the idiom that omits it looks completely reasonable: ```go for sc.Scan() { ... } // wrong: indistinguishable from success ``` That is the single most common `bufio.Scanner` defect, and it fails silently by construction. The corrected form is three lines longer: ```go for sc.Scan() { ... } if err := sc.Err(); err != nil { return fmt.Errorf("read %s: %w", path, err) } ``` ## Raising the limit `Buffer(buf []byte, max int)` sets both the initial buffer the scanner uses and the maximum token size it will grow to: ```go sc := bufio.NewScanner(f) sc.Buffer(make([]byte, 0, 64*1024), 4*1024*1024) // start at 64 KiB, cap at 4 MiB ``` Two constraints. It must be called **before the first `Scan`** - calling it afterwards panics with "Buffer called after Scan". And the `max` you choose is a real memory commitment: it is per scanner, so an agent with one scanner per tailed file multiplies it by the number of files. Choosing a huge cap to "be safe" trades a silent truncation for an allocation a hostile or corrupt producer controls. ## Proving it in a test The cheapest way to keep the fix from regressing is a unit test that feeds exactly the failing input: ```go long := strings.Repeat("x", bufio.MaxScanTokenSize+1) + "\n" sc := bufio.NewScanner(strings.NewReader(long)) if sc.Scan() { t.Fatal("expected Scan to fail") } if !errors.Is(sc.Err(), bufio.ErrTooLong) { t.Fatalf("Err = %v", sc.Err()) } ``` Once that test exists, the behaviour you chose - a bigger cap, a skip-and-continue path, or a hard failure - is pinned. ## Two more scanner details worth knowing **`Bytes()` is borrowed.** The slice it returns points into the scanner's own buffer and is only valid until the next `Scan`. Retaining it - appending it to a batch, stashing it in a map - hands you corrupted records later. Copy it, or use `Text()`, which allocates a fresh string. **The split function decides the token.** `bufio.ScanLines` is the default and strips the trailing newline **and** an optional preceding carriage return, so CRLF input yields clean tokens. `bufio.ScanWords`, `bufio.ScanRunes` and `bufio.ScanBytes` are the other supplied ones, and the same size limit applies to whatever they produce. ## When to stop using bufio.Scanner If over-long records are normal rather than exceptional, the scanner is the wrong tool: it is designed around "the whole token fits in memory". Reading with a `bufio.Reader` lets you consume a long line incrementally and decide per record whether to process, truncate or skip it, which is often what a log pipeline actually wants.

  • Why not just set the maximum token size to something enormous and forget about it?
    Because the cap is what stops one line of input from deciding your memory use. A corrupt or hostile producer that emits a single 500 MB line would then allocate 500 MB per scanner, and an agent tailing many files multiplies that. Pick a cap you can afford per concurrent stream and treat exceeding it as a data-quality signal worth reporting.
  • How long is the slice returned by Scanner.Bytes valid?
    Only until the next Scan call - it points into the scanner's own buffer, which the next token overwrites. Batching those slices gives you records that mutate under you. Copy the bytes, or call Text(), which allocates a new string per token at the cost of an allocation.
  • What does the default split function do with a CRLF line ending?
    bufio.ScanLines strips the trailing newline and one optional carriage return before it, so a line written as "a,b\r\n" yields the token "a,b". That is why scanning a Windows-authored file usually needs no special handling, while reading the same file with a raw delimiter search leaves the \r attached.

saying these in an interview costs you the question

  • Assumes Scan returning false always means the input ended
  • Believes bufio.Scanner has no line-length limit
  • Tries to call Buffer inside the scanning loop
  • Blames the producer without checking Err
  • Retains the slice from Bytes across iterations
  • Sets an unbounded maximum token size to make the error go away