skip to content

In Go's encoding/csv, when would you use Reader.Read in a loop instead of Reader.ReadAll?

level: juniorimportance: must knowfreq 55%

answer

  1. one call, or one row at a time
  2. how big is that file, really
  3. the loop ends on a sentinel error
  4. one bad row can cost you every good one

basics

~20 s

ReadAll builds the whole file as one [][]string in memory and returns nil records if any row fails to parse. Read hands back one record per call and ends with io.EOF, so large or partly malformed files stay workable.

solid answer

~50 s

`ReadAll` is the convenience path: it reads to the end, returns every record as a `[][]string`, and reports `err == nil` on a clean run rather than `io.EOF`. Two things push me to `Read` instead. First, memory: the returned records hold the entire file expanded into Go strings and slice headers, so a multi-gigabyte partner export is not an option. Second, error granularity: if `ReadAll` hits a bad record it returns a nil slice and the error, throwing away everything it had already parsed, so one malformed line at row 90,000 costs the whole import. Driving `Reader.Read` in a loop — break on `io.EOF`, check any other error, handle the record — lets me process rows as they arrive and quarantine the bad ones with their line numbers. For a small, trusted file, `ReadAll` is fine and shorter.

go deeper

for a junior

Be able to write the Read loop from memory: call Read, break on io.EOF, check any other error, then use the record. Know that ReadAll exists and holds the entire file in memory.

for a middle

Explain what ReadAll costs relative to the file on disk, and that it discards already-parsed records when it hits an error rather than returning them.

for a senior

Show how you keep an import running over a partner file with a few bad rows: stream with Read, quarantine failures with their line numbers, and bound whatever you batch up in between.

for a principal

Own the policy: does an import abort on the first malformed record or finish and produce a rejects report? That choice decides how re-runnable the job is and what the partner has to fix before the next drop.

## The two ways to pull records out `encoding/csv` gives you a `*csv.Reader` built with `csv.NewReader(r io.Reader)`. A *record* is one row, represented as a `[]string` of already-unquoted fields: a quoted field containing commas or newlines arrives as a single element, so you never re-split anything yourself. There are exactly two ways to get records out: - `func (r *Reader) Read() (record []string, err error)` — one record per call. - `func (r *Reader) ReadAll() (records [][]string, err error)` — everything, at once. ## What ReadAll actually does `ReadAll` loops internally until the input is exhausted and then returns the accumulated slice. Two details matter and both surprise people: 1. **A clean run returns `err == nil`, not `io.EOF`.** `ReadAll` converts end-of-input into success. Code that checks `if err == io.EOF` around `ReadAll` never fires. 2. **An error returns nothing.** If record 4,271 is malformed, `ReadAll` returns a nil `[][]string` together with the error. The 4,270 records it had already parsed are discarded. You cannot recover partial work from it. That second point is why `ReadAll` is a poor fit for files you did not produce. A batch import over partner-supplied exports meets ragged rows, stray quotes and trailing summary lines routinely, and an all-or-nothing call turns each of those into "the whole file failed" with no way to report which row was at fault. ## What Read does `Read` returns one record and advances. End of input is signalled by `io.EOF`, which is the normal, expected termination of the loop, not a failure to log: ```go for { rec, err := r.Read() if err == io.EOF { break } if err != nil { return err } // use rec } ``` Always check the error before touching the record. A parse failure concerns one record and you may decide to log it and keep reading; an error coming from the underlying `io.Reader` ends the stream and there is nothing further to read. ## Memory, concretely A CSV on disk is bytes. A `[][]string` of the same data is Go strings plus a slice header per row plus a string header per field. A 2 GB file with many small fields expands well past 2 GB of live heap, and every one of those objects stays reachable for as long as you hold the outer slice, so the garbage collector cannot help you. Streaming with `Read` keeps only the current record plus whatever you have deliberately accumulated — a batch of a few thousand rows waiting to be written onward, say — and that batch size is a number you control. ## Choosing Use `ReadAll` when the file is small, trusted, and you genuinely need random access to all rows: a checked-in fixture, a lookup table of a few hundred lines, a test. Use `Read` when the file is large, when it comes from outside your system, when you want to stream each row onward as it is parsed, or when you need to keep going past a bad record instead of aborting. A useful middle ground for imports is `Read` in a loop with your own bounded batch: accumulate N records, hand the batch to the next stage, reset the batch. That gives you the throughput of bulk work without the unbounded footprint of `ReadAll`, and it keeps every parse failure attributable to a line number you can put in a rejects report. ## The habit to build Write the `Read` loop by default and reach for `ReadAll` only when you can say out loud why the whole file fits in memory. The loop is four extra lines, and it is the version that still works when the file grows or when a partner ships you a broken row.

  • What does ReadAll return when the file's tenth record is malformed?
    Nothing usable: a nil [][]string plus the error. The nine records it had already parsed are discarded, so there is no partial result to salvage. It returns records only on a clean run to end of input, where the error is nil rather than io.EOF. If you want to keep the good rows and quarantine the bad ones, you have to drive Reader.Read yourself.
  • How should the loop tell end of input apart from a bad record?
    Break on io.EOF — that is the only way Read signals the end of the stream, and it is not a failure. Anything else is about the input: unwrap it to a *csv.ParseError to get the line, and decide whether to abort the import or record the row and keep going. Never log io.EOF as an error.
  • Is ReadAll ever the right call in production code?
    Yes, when the input is bounded and you control it: a small configuration table, a fixture in a test, a lookup file of a few hundred rows you need to index by column. The judgement is whether you can state an upper bound on the file size. If the file comes from a partner or grows with your customers, you cannot, and the loop is the honest choice.

saying these in an interview costs you the question

  • Expects ReadAll to return io.EOF at the end of a clean file
  • Believes ReadAll returns the rows it parsed before an error
  • Treats io.EOF from Read as a failure worth logging
  • Assumes any file size is fine because the OS pages it
  • Re-splits returned fields on commas by hand