skip to content

A Go import sets csv.Reader.ReuseRecord = true and keeps every record; why do all rows end up identical?

level: seniorimportance: nice to knowfreq 28%

answer

  1. the flag says reuse, and it means it
  2. append copies the header, not the array
  3. two views of one backing array
  4. the last record wins — copy before you keep

basics

~20 s

With ReuseRecord set, Read may return a slice sharing the previous call's backing array, so every retained row points at that one array and shows whatever the last record left in it. Copy each record you intend to keep.

solid answer

~50 s

`ReuseRecord` is an opt-in performance knob: it lets `Reader.Read` hand back a `[]string` that shares the backing array of the previous call's record, saving one allocation per row. That is fine when you consume the record fully inside the loop body — parse it into a struct, write it onward — and wrong the moment you retain it. Appending the returned slice to a `[][]string` stores a view of the same array every time, so at the end every row shows the last record read. The fix is to copy before keeping: `slices.Clone(rec)`, or `append([]string(nil), rec...)`. Note this is not a data race — one goroutine, ordinary reuse — so `-race` and `go vet` both stay silent; a unit test over a two-record fixture asserting the two retained rows differ is what pins it. `ReadAll` is unaffected, because it allocates a fresh record per row.

code

go · 14 lines
go
r := csv.NewReader(f)
r.ReuseRecord = true

var rows [][]string
for {
	rec, err := r.Read()
	if err == io.EOF {
		break
	}
	if err != nil {
		return err
	}
	rows = append(rows, slices.Clone(rec)) // without the clone every row aliases one array
}

go deeper

for a junior

Know that a []string returned by a library may point at memory the library reuses, and that keeping it beyond the current iteration needs a copy.

for a middle

Explain why append stores the slice header rather than the elements, and name the copy that fixes it — slices.Clone or append([]string(nil), rec...).

for a senior

Diagnose it from the symptom alone: identical rows, a clean -race run, no vet finding. Reproduce it with a two-record fixture and say why the flag was safe when it was introduced and stopped being safe later.

for a principal

Take a position on knobs whose safety depends on caller behaviour: whether such a micro-optimisation belongs in shared import code at all, and what documentation or test makes the contract survive the next refactor.

## What the flag buys `csv.Reader` has a `ReuseRecord bool` field. When it is true, calls to `Read` may return a slice that shares the backing array of the previous call's returned slice. The point is allocation: a long import calls `Read` millions of times, and each call otherwise allocates a fresh `[]string`. Turning that off is a measurable win in a hot import loop, and it is exactly the kind of knob you reach for after a benchmark with `-benchmem` has shown you where the allocations are. The contract it asks for in exchange is simple and unforgiving: **the record you get is only yours until the next `Read`.** ## How the bug forms ```go r.ReuseRecord = true var rows [][]string for { rec, err := r.Read() if err == io.EOF { break } rows = append(rows, rec) // the bug } ``` `append` copies the slice *header* — pointer, length, capacity — not the elements it points at. Every element of `rows` therefore holds a pointer to the same array. The next `Read` overwrites that array in place. When the loop ends, all N entries of `rows` are N views of one array holding the final record, and the import writes the last row a hundred thousand times. Re-slicing does not help: `rec[:]` produces another header pointing at the same array. Assigning to a new variable, storing it in a struct field, passing it to a goroutine — all the same. Nothing short of copying the elements breaks the aliasing. ## The fix ```go rows = append(rows, slices.Clone(rec)) ``` `slices.Clone` allocates a new backing array and copies the elements, so the retained row is independent. `append([]string(nil), rec...)` is the older spelling of the same thing. If all you need from the record is a couple of fields, take them out by value and let the record go — a `string` you have copied into a struct field is yours regardless of what happens to the record slice. So the rule is: **clone whenever the record outlives the loop iteration.** ## Why nothing warns you This is the part that makes it a senior question. It is not a data race — a single goroutine reads and appends, and the reuse is ordinary sequential mutation — so the race detector reports a clean run no matter how long you exercise it. `go vet` has no check for retaining a reused record. The compiler is happy: the types are all correct. The only signal is the data itself, and only if someone looks at more than one row. It also hides in testing. A fixture with a single record cannot show the problem: with one row there is nothing to overwrite it. A fixture with two rows shows it immediately, which is why the cheap reproduction is a unit test over a two-record input asserting `rows[0][0] != rows[1][0]`. That test is worth writing even after you fix it, because the fix is one call that a future refactor can quietly delete. ## Where it typically surfaces A batch import over partner files runs for months. Somebody profiles it, sees the per-record allocation, sets `ReuseRecord = true`, benchmarks it, ships it. The import loop at that time consumed each record inline, so nothing broke. Six months later a feature needs the rows collected — for a summary, a dedupe pass, a bulk write — and the collecting code is written by someone who never saw the flag three files away. Every imported row is now the last row of the file, and the symptom looks like a corrupt export from the partner rather than a bug in the reader configuration. That is worth remembering when you are the one reviewing the PR: a flag whose safety depends on what the *caller* does with the value is a flag that needs a comment at the point it is set, not only at the point it is used. ## And ReadAll `ReadAll` allocates a fresh record for each row, so the `[][]string` it returns holds independent slices and `ReuseRecord` does not affect it. The flag exists purely for the streaming path. ## The judgement Default to leaving `ReuseRecord` false. Turn it on only when a benchmark says the allocation matters, only when the loop body demonstrably consumes the record before the next iteration, and with a comment saying so. One saved allocation per row is not worth an import that silently writes the same record a million times.

  • Would the race detector or go vet catch this?
    Neither. Nothing is concurrent: one goroutine reads and appends, and the aliasing is ordinary sequential reuse of a backing array, so a -race run is clean however long you exercise it. go vet has no check for retaining a reused record. The cheap reproduction is a unit test over a two-record fixture asserting the two retained rows differ.
  • Does re-slicing the record with rec[:] before appending fix it?
    No. Re-slicing produces another slice header pointing at the same backing array, so the retained value still aliases the record the next Read overwrites. You need a real element copy: slices.Clone(rec), or append([]string(nil), rec...). Anything that does not allocate a new array leaves the bug in place.
  • Does ReuseRecord affect Reader.ReadAll?
    No. ReadAll allocates a fresh record for each row, so the [][]string it returns holds independent slices regardless of the flag. ReuseRecord exists for the streaming path, where you call Read in a loop and are expected to finish with the record before the next call.
  • When is turning the flag on actually worth it?
    When a benchmark with -benchmem shows the per-record allocation is a real cost in a hot import, and the loop body fully consumes each record — parses it into a struct, writes it onward — before the next iteration. Set it with a comment stating that contract, because its safety depends on what the caller does with the value, not on the reader.

saying these in an interview costs you the question

  • Blames the partner's file for the repeated rows
  • Expects the race detector to find it
  • Thinks re-slicing or reassigning counts as a copy
  • Turns ReuseRecord on by default for speed
  • Tests with a one-record fixture and declares it fixed