skip to content

In a Go CSV loop, what does `if err != nil { continue }` hide from the caller?

level: middleimportance: should knowfreq 44%

answer

  1. the error was checked, then routed around
  2. failure arriving as an empty result
  3. green run, missing rows
  4. count it, sample it, bound it
  5. what number proves this happened?

basics

~20 s

It converts a failure into an absence. Bad records vanish, the loop finishes normally, and the function returns success, so the caller cannot tell a clean run from one that skipped a quarter of the input.

solid answer

~50 s

Skipping on error is the second way to swallow one: the error is never assigned to `_`, it is simply routed around by control flow. The loop keeps going, the function returns success, and the only trace of the failure is data that is not in the output. The same shape appears as `if err != nil { return nil }` in a helper, which leaves the caller unable to distinguish "no rows" from "the read failed". Skipping bad records can be a legitimate product decision — but then it must be counted, bounded and reported: keep a skip counter and a sample of the first few offending row numbers, return an error when the count crosses the threshold you agreed, and put the count in the run's summary. The anti-pattern is not the skip, it is the skip nobody can see.

code

go · 21 lines
go
var skipped int
var firstBad []string // sample of offending rows

for row := 1; ; row++ {
	rec, err := r.Read()
	if err == io.EOF {
		break
	}
	if err != nil {
		skipped++
		if len(firstBad) < 5 {
			firstBad = append(firstBad, fmt.Sprintf("row %d: %v", row, err))
		}
		continue
	}
	totals.add(rec)
}

if skipped > maxSkipped {
	return fmt.Errorf("skipped %d rows (limit %d): %v", skipped, maxSkipped, firstBad)
}

go deeper

for a junior

Recognise the shape: an error that is checked and then skipped over with continue, or replaced by an empty return value, is still a swallowed error. The run succeeds and the data is quietly missing.

for a middle

Explain why the zero value is a dangerous substitute for a failure, and show the counted-and-bounded skip — counter, sample with row numbers, threshold — that turns an invisible drop into a reported outcome.

for a senior

Bring the operational view: what number the run publishes, what threshold fails it, and why an incomplete but valid output file is more dangerous to a downstream consumer than a run that failed loudly.

for a principal

Own the policy question of how much bad input a pipeline tolerates before it stops, who agrees that bound with the data's consumers, and how it gets revisited when the input's quality changes.

## Two ways to swallow an error The obvious one is lexical: `_ = doThing()`, or a bare call whose error return is dropped. It is visible on the line where it happens. The second is structural, and it is much harder to spot in review because the error *is* checked — it is checked and then discarded by the control flow: ```go for { rec, err := r.Read() if err == io.EOF { break } if err != nil { continue // the swallow } // ... roll rec into the totals ... } return nil ``` Every `if err != nil` box is ticked. The function still reports success on a run where a quarter of the input never made it into the output. ## The same shape in a helper ```go func loadTotals(path string) []Total { data, err := os.ReadFile(path) if err != nil { return nil // failure rendered as "nothing here" } ... } ``` The caller receives an empty result and has no way to distinguish *there were no totals* from *the file could not be read*. Turning a failure into a zero value is the single most common way a Go program loses information about itself, because the zero value is always a plausible answer. ## Why skipping is not automatically wrong A batch job that must not die on one malformed row is a reasonable requirement. Real exports contain garbage. "Fail the whole nightly run because record 8,400,112 has an unparseable date" is often the worse policy. So the fix is not "never skip". The fix is to make the skip a **reported outcome** instead of an invisible one. Three things turn it around: 1. **Count it.** A `skipped` counter costs nothing and is the number the postmortem will ask for first. 2. **Sample it.** Keep the first few failures with their row numbers, so somebody can look at an actual bad record rather than guessing. 3. **Bound it.** Agree a threshold — an absolute count, or a fraction of rows read — above which the run fails instead of succeeding. A job that silently drops 25% of its input is a different event from one that drops three rows, and only the threshold encodes that difference. Then return the outcome. Either return an error when the bound is exceeded, or return a summary value the caller records; what you must not do is return success with no number attached. ## Distinguishing a terminator from a failure One genuine subtlety in this loop shape: `io.EOF` from a reader is the loop's normal ending, not a failure, and treating it as one is its own bug. The discipline is that *exactly one* error value means "stop, we are done", and every other error is either handled or counted. A loop that lumps all non-EOF errors into `continue` has not made that distinction — it has erased it. ## What the caller and the operator actually see Think about the three people downstream of this loop. - **The caller** gets `nil`. It has no reason to retry, alert, or refuse to publish. - **The scheduler** gets exit code 0. The run is green in every dashboard. - **The consumer of the output file** gets a file that is structurally valid and quietly incomplete, which is worse than a missing file — a missing file trips an alert, an incomplete one gets used. That asymmetry is the whole argument. A loud failure costs one night of a batch run. A silent one costs a quarter of a quarter's data and is discovered by somebody reconciling numbers weeks later. ## In review When you see `continue` or `return nil` immediately under an `if err != nil`, ask one question: *what number tells us this happened?* If the answer is "nothing", the error has been swallowed no matter how carefully it was checked.

  • Why is returning a zero value on error especially dangerous compared with returning an explicit error?
    Because the zero value is always a plausible answer. An empty slice reads as "there was nothing", `0` reads as "the total is zero", and the caller has no way to tell either apart from a failed read. The failure does not just get lost — it gets replaced by a value the rest of the program will confidently act on.
  • How do you decide the threshold above which a batch job should fail rather than skip?
    From what the output is used for. If a consumer reconciles totals, any drop matters and the bound is near zero; if the job feeds a trend chart, a fraction of a percent may be tolerable. Express it as both an absolute count and a share of rows read, agree it with the consumer, and record the actual number on every run so the threshold can be revisited with evidence.
  • Is treating `io.EOF` as a failure in a read loop the same mistake?
    It is the mirror image. Here a normal terminator is reported as a failure, which produces noisy alerts and, worse, trains people to ignore them. A well-written loop names exactly one value as "we are done" and treats every other error as something to handle or count — collapsing that distinction in either direction loses information.

saying these in an interview costs you the question

  • Thinks checking err and then continuing counts as handling it
  • Returns an empty slice on read failure and calls it graceful degradation
  • Skips bad records with no counter and no bound
  • Cannot say how many rows a run dropped
  • Treats an incomplete output file as safer than a failed run