A Go ingest job read a truncated line-delimited feed as a complete batch with no error. How do you locate the faulty end-of-stream check?
answer
- write the red test first
- one signal carries two different events
- a false Scan means end or error
- Err stays nil at a clean end
- compare bytes copied against bytes promised
basics
~20 sReproduce it with a deliberately truncated fixture first, then audit where the loop decided it was finished: an unchecked bufio.Scanner.Err, or a break on any error, both turn a failure into a clean end. Then compare bytes actually read against the length the source promised.
solid answer
~50 sStart by writing the failing test: feed the pipeline a fixture cut mid-line, or a reader that fails partway using `testing/iotest.ErrReader`, and assert the job returns an error. That converts an argument into a reproduction. Then audit the two places an ending is decided. First, `bufio.Scanner`: `Scan` returning `false` is ambiguous — it means "end of input **or** error" — and `Err` returns nil at a clean end, so a missing `if err := sc.Err(); err != nil` swallows every mid-stream failure. Second, any hand-written loop shaped `if err != nil { break }`, which files a read failure as a finish. Finally, accept that the byte layer cannot distinguish a finished producer from a dead one: add an independent length check, comparing the count `io.Copy` returns against the size the source declared, or read fixed-size records with `io.ReadFull` so truncation surfaces as `io.ErrUnexpectedEOF`.
code
go · 8 linessc := bufio.NewScanner(r)
for sc.Scan() {
handle(sc.Bytes())
}
if err := sc.Err(); err != nil { // nil when the input simply ended
return fmt.Errorf("scan feed: %w", err)
}
return nilgo deeper
Remember the mechanical check: after a scan loop over a stream, always test the scanner's Err method, because the loop ends the same way whether the input finished or the read failed.
Explain the ambiguity precisely: Scan returning false covers both the end of input and an error, and Err returns nil at a clean end, so the single check after the loop is what separates them.
Show the full diagnosis: reproduce with a truncated fixture, audit every place an ending is inferred, then add an expectation the stream itself cannot fake, such as comparing copied bytes against a declared length.
Own the pipeline-wide position: silent truncation is a contract gap between producer and consumer, so the durable fix is a framing convention plus reported counts, not a patch in one job's read loop.
## The shape of this failure A batch job reads line-delimited records from a stream, processes each one, and reports success. One day the upstream feed is cut off partway — a producer crash, a half-written file, a connection reset — and the job reports a clean run over half the data. Nothing logs an error. Downstream, the missing half is indistinguishable from records that were never produced. This is the most expensive class of stream bug precisely because it is silent. ## Step 1: make it fail on demand Before touching the code, get a red test. Two cheap ways: - a fixture file cut mid-line, checked into the repo alongside the good one; - `testing/iotest.ErrReader(errors.New("connection reset"))` composed after a reader that yields a few good records, so the pipeline sees real data then a real failure. Assert two things: the job returns a non-nil error, and it does not report the partial batch as complete. Until that test is red, every subsequent "fix" is a guess. ## Step 2: audit where the loop decides it is finished There are only a few candidates, and all of them collapse two different events into one. **bufio.Scanner without Err.** `Scan` returns `false` for *both* the end of the input and any error. `Err` is what disambiguates, and it reports the first *non-EOF* error — which means it deliberately returns `nil` when the input simply ran out, and returns the real error otherwise. That design is fine; omitting the check is not: ``` sc := bufio.NewScanner(r) for sc.Scan() { handle(sc.Bytes()) } if err := sc.Err(); err != nil { // nil at a clean end of input return fmt.Errorf("scan feed: %w", err) } ``` Dropping that final block is the single most common cause of this incident. **A raw loop that breaks on any error.** `if err != nil { break }` followed by `return nil` classifies a disk or socket failure as a finish. The correct form tests `io.EOF` specifically and returns every other error. **An ending accepted mid-record.** If the format has structure — a header committing to a body, a multi-line record — an ending is only legal at a record boundary. Code that treats `io.EOF` as "done" everywhere accepts a half record. ## Step 3: accept what io.EOF cannot tell you Even with every error checked, one case remains: the producer's stream was cut in a way that reaches you as a genuine, clean end. A file that was only half written, an object whose upload was abandoned, a sender that exited between records — all of these end the byte stream politely. No amount of error handling detects that, because there is no error. The information simply is not in the stream. The fix is framing plus an independent expectation: - **Compare the byte count.** `io.Copy` returns how many bytes it moved. If the source declared a length — a file size, a manifest entry, a content length recorded by the producer — compare and fail on a mismatch. This is the cheapest check that catches a polite truncation. - **Read fixed-size or length-prefixed records with `io.ReadFull`**, so a short record comes back as `io.ErrUnexpectedEOF` rather than as an ending. - **Terminate the stream explicitly.** A trailer record, a record count, or a checksum turns "the bytes stopped" into a verifiable claim. A count in the trailer is the simplest version and catches both truncation and duplication. - **Bound the input** with `io.LimitReader` where a maximum is known, so the opposite failure — an unbounded stream — is also caught. ## Step 4: make the wrong version impossible Once fixed, keep it fixed. Put the loop behind one function that every ingest path uses, so the `Err` check exists once rather than at every call site. Keep the truncated fixture in the test suite permanently. And report the count: a job that logs "processed 41,233 of 41,233 declared records" gives an operator something to compare, whereas "batch complete" gives them nothing. ## What to say in the interview Name the ambiguity first — an ending and a failure arrive through the same channel — then the two code-level causes, then the structural fix. Candidates who jump straight to "add error handling" miss the half of the problem where there was never an error to handle.
- Everything is error-checked and the feed still truncates silently. What is left?A producer that died between records ends the byte stream cleanly, so there is no error to check — the information is not in the stream. Only an independent expectation catches it: a declared length compared against the bytes actually copied, a record count or checksum in a trailer, or length-prefixed records read with io.ReadFull so a short one fails.
- How would you build the failing test before changing any code?Two fixtures. A file cut mid-line, checked in beside the good one, proves the format-level case. A reader that yields a few good records and then fails — testing/iotest.ErrReader supplies the failure — proves the transport case. Assert the job returns an error and does not mark the partial batch complete.
- Why does bufio.Scanner's Err return nil when the input ends?Because reaching the end is not an error, so surfacing io.EOF would force every caller to filter it. Err reports the first non-EOF error, which makes the correct usage a single check after the loop: nil means the input genuinely ran out, non-nil means the scan stopped for a real reason.
- What would you change so this class of bug cannot recur?Put the read loop behind one shared function so the error check exists once instead of at every ingest site, keep the truncated fixture in the suite permanently, and have the job report processed-versus-declared counts rather than a bare success, so an operator can see a short run without reading logs.
saying these in an interview costs you the question
- Assuming a false result from Scan always means end of input
- Skipping the error check after a scan loop
- Breaking out of a read loop on any non-nil error and returning success
- Believing io.EOF can distinguish a finished producer from a dead one
- Reporting batch success without comparing counts or bytes