A bufio.Scanner loop stops partway through a log file with no error - what limit causes that?
answer
- the loop exit looks like a clean end
- a per-token size ceiling exists
- 64 KiB unless you say otherwise
- one method after the loop tells them apart
- Buffer must come before the first Scan
basics
~20 sA line longer than the scanner's maximum token size, 64 KiB by default, makes Scan return false with bufio.ErrTooLong stored in Err. A loop that never calls Err after Scan sees that failure as a clean end of input.
solid answer
~40 s`bufio.Scanner` grows its buffer only up to `bufio.MaxScanTokenSize`, which is 64 KiB. Hit a longer line and `Scan` returns `false` and records `bufio.ErrTooLong`. Because `for sc.Scan()` exits the same way it exits at end of input, a loop that never checks `sc.Err()` afterwards reports success and quietly drops the rest of the file - an hour of records vanishing with no log line and no non-zero exit. Two fixes go together: always check `sc.Err()` after the loop, and if long records are legitimate, call `sc.Buffer(make([]byte, 0, 64*1024), maxYouCanAfford)` before the first `Scan` - it panics if called later. Do not set the cap to something enormous; that just converts a corrupt line into a large allocation.
code
go · 8 linessc := bufio.NewScanner(f)
sc.Buffer(make([]byte, 0, 64*1024), 4*1024*1024) // start 64 KiB, cap 4 MiB
for sc.Scan() {
ingest(sc.Bytes()) // valid only until the next Scan
}
if err := sc.Err(); err != nil {
return fmt.Errorf("tail %s: %w", path, err) // ErrTooLong lands here
}go deeper
Remember the two-line habit: check the scanner's Err after the loop, every time. Know that a very long line is a real failure mode and not something the library hides from you.
Explain the mechanics you are expected to own here: a 64 KiB default maximum token, the buffer growing up to it, Scan returning false with ErrTooLong, and Buffer having to be called before the first Scan.
Talk about how you would have caught this in production - the missing-records symptom, the test that feeds an over-long line, and the deliberate choice between raising the cap, skipping the record, or failing loudly.
Own the cap as a policy question: how much memory per stream the ingest fleet can commit, whether an over-long record should be dropped or should stop the pipeline, and how that decision is enforced consistently rather than per author.
## The shape of the bug A log-ingest agent tails a newline-delimited file: ```go sc := bufio.NewScanner(f) for sc.Scan() { ingest(sc.Bytes()) } return nil ``` It runs for weeks, then one day an hour of records is missing downstream and the agent logged nothing. What happened is that one process upstream wrote a single enormous line - a stack dump, a base64 blob, a JSON document with an embedded payload - and the scanner stopped there. ## Why a bufio.Scanner has a limit at all `bufio.Scanner` is a tokenizer. It reads into an internal buffer and hands out one token per `Scan` - by default a line, using the `bufio.ScanLines` split function. To return a whole token it must hold that whole token in memory, so an unbounded scanner would let a single line of input dictate the process's memory use. The library refuses to do that: the buffer starts small, grows as needed by doubling, and stops at `bufio.MaxScanTokenSize`, a constant equal to 64 * 1024. When a token would exceed that limit, `Scan` returns `false` and stores `bufio.ErrTooLong` ("token too long"). It does **not** panic, does **not** truncate the token and hand you a prefix, and does **not** skip the line and continue. The scanner is finished. ## Why nobody notices `Scan` returning `false` is the loop's normal exit condition. At a clean end of input `Scan` also returns `false`, and in that case `Err()` is `nil`. The only thing distinguishing "the file ended" from "I gave up" is the value of `Err()`, and the idiom that omits it looks completely reasonable: ```go for sc.Scan() { ... } // wrong: indistinguishable from success ``` That is the single most common `bufio.Scanner` defect, and it fails silently by construction. The corrected form is three lines longer: ```go for sc.Scan() { ... } if err := sc.Err(); err != nil { return fmt.Errorf("read %s: %w", path, err) } ``` ## Raising the limit `Buffer(buf []byte, max int)` sets both the initial buffer the scanner uses and the maximum token size it will grow to: ```go sc := bufio.NewScanner(f) sc.Buffer(make([]byte, 0, 64*1024), 4*1024*1024) // start at 64 KiB, cap at 4 MiB ``` Two constraints. It must be called **before the first `Scan`** - calling it afterwards panics with "Buffer called after Scan". And the `max` you choose is a real memory commitment: it is per scanner, so an agent with one scanner per tailed file multiplies it by the number of files. Choosing a huge cap to "be safe" trades a silent truncation for an allocation a hostile or corrupt producer controls. ## Proving it in a test The cheapest way to keep the fix from regressing is a unit test that feeds exactly the failing input: ```go long := strings.Repeat("x", bufio.MaxScanTokenSize+1) + "\n" sc := bufio.NewScanner(strings.NewReader(long)) if sc.Scan() { t.Fatal("expected Scan to fail") } if !errors.Is(sc.Err(), bufio.ErrTooLong) { t.Fatalf("Err = %v", sc.Err()) } ``` Once that test exists, the behaviour you chose - a bigger cap, a skip-and-continue path, or a hard failure - is pinned. ## Two more scanner details worth knowing **`Bytes()` is borrowed.** The slice it returns points into the scanner's own buffer and is only valid until the next `Scan`. Retaining it - appending it to a batch, stashing it in a map - hands you corrupted records later. Copy it, or use `Text()`, which allocates a fresh string. **The split function decides the token.** `bufio.ScanLines` is the default and strips the trailing newline **and** an optional preceding carriage return, so CRLF input yields clean tokens. `bufio.ScanWords`, `bufio.ScanRunes` and `bufio.ScanBytes` are the other supplied ones, and the same size limit applies to whatever they produce. ## When to stop using bufio.Scanner If over-long records are normal rather than exceptional, the scanner is the wrong tool: it is designed around "the whole token fits in memory". Reading with a `bufio.Reader` lets you consume a long line incrementally and decide per record whether to process, truncate or skip it, which is often what a log pipeline actually wants.
- Why not just set the maximum token size to something enormous and forget about it?Because the cap is what stops one line of input from deciding your memory use. A corrupt or hostile producer that emits a single 500 MB line would then allocate 500 MB per scanner, and an agent tailing many files multiplies that. Pick a cap you can afford per concurrent stream and treat exceeding it as a data-quality signal worth reporting.
- How long is the slice returned by Scanner.Bytes valid?Only until the next Scan call - it points into the scanner's own buffer, which the next token overwrites. Batching those slices gives you records that mutate under you. Copy the bytes, or call Text(), which allocates a new string per token at the cost of an allocation.
- What does the default split function do with a CRLF line ending?bufio.ScanLines strips the trailing newline and one optional carriage return before it, so a line written as "a,b\r\n" yields the token "a,b". That is why scanning a Windows-authored file usually needs no special handling, while reading the same file with a raw delimiter search leaves the \r attached.
saying these in an interview costs you the question
- Assumes Scan returning false always means the input ended
- Believes bufio.Scanner has no line-length limit
- Tries to call Buffer inside the scanning loop
- Blames the producer without checking Err
- Retains the slice from Bytes across iterations
- Sets an unbounded maximum token size to make the error go away