skip to content

A Go nightly job exits 0 but its output file is short a quarter of its rows. How do you find the swallowed error?

level: seniorimportance: should knowfreq 38%

answer

  1. exit zero means nobody read a return value
  2. count the rows, do not trust the log
  3. re-open the artefact you just wrote
  4. which side of the gap lost them
  5. a truncated last record means never flushed

basics

~20 s

Establish the numbers first: records read, records handed to the writer, and rows actually present when you re-open and count the file. The gap says which side lost them, and a zero exit code means a returned error nobody read.

solid answer

~50 s

Start with counts, not code. Record rows read from the inputs, records handed to the writer, and rows present in the output file when you re-open and count it afterwards. If read and handed-to-writer disagree, records were skipped on the read side by control flow. If handed-to-writer and on-disk disagree, the loss is on the write path, and a truncated final record points straight at a buffered writer that was never flushed. Because the process exited zero, the failure was reported through a return value nobody read: audit every `Write`, `Flush`, `Close` and `Sync` on that path, including the ones hidden inside deferred calls, and remember `csv.Writer.Flush()` returns nothing so its error is only visible through `Error()`. Then reproduce deliberately against a nearly full volume, and make the job verify its own output before it can exit zero.

code

go · 22 lines
go
func verify(path string, want int) error {
	f, err := os.Open(path)
	if err != nil {
		return err
	}
	defer f.Close() // read path: nothing buffered, nothing to lose

	got := 0
	r := csv.NewReader(f)
	for {
		if _, err := r.Read(); err == io.EOF {
			break
		} else if err != nil {
			return err
		}
		got++
	}
	if got != want {
		return fmt.Errorf("wrote %d rows, file holds %d", want, got)
	}
	return nil
}

go deeper

for a junior

Take away the diagnostic order: get the numbers first — rows read, rows written, rows actually in the file when you open it again — before opening the source. A green run is not evidence the output is complete.

for a middle

Be able to map each gap to a cause: a read-side skip, an unflushed writer, a discarded Close error. Know that a truncated final record points at a buffer that was never flushed.

for a senior

Show the whole investigation: independent measurement, the audit of discarded returns on the write path, a deliberate reproduction against a full volume, and a fix that makes the job verify its own output before it can exit zero.

for a principal

Frame the finding as a class of failure, not an incident: green runs were accepted as proof of complete data for months. Decide what every batch job in the estate must assert about its own output before it is allowed to report success.

## Why this incident is always an unread return value A process that exits 0 is asserting it succeeded. If the output is incomplete, then either nothing detected the problem, or something detected it and the detection went nowhere. In Go the second is overwhelmingly more likely, because every I/O failure on the write path arrives as a returned error — and a returned error that nobody binds is gone without trace. That single observation narrows the search enormously before you have read any code. ## Step 1: get four numbers Do not start by reading the write path. Start by finding out where the rows disappeared. 1. **Rows read** from the input exports. 2. **Records handed to the writer** by the job. 3. **Rows in the output file**, obtained by re-opening the file the job just wrote and counting them with a reader. 4. **Bytes on disk**, from a stat of the same file. Re-reading the artefact is the key move: it is the only number that is independent of the job's own belief about what it did. Everything else is the program marking its own homework. ## Step 2: read the gap - **Read count > handed-to-writer count.** Records were dropped on the read side. Look for an `if err != nil { continue }`, a filter that treats a parse failure as a non-match, or a helper returning a zero value on error. - **Handed-to-writer count > rows on disk.** The loss is on the write path. - **The last row on disk is truncated mid-record.** That is the signature of a buffer that was never flushed: the file ends wherever the last full buffer landed, not at a record boundary. - **Counts match but a whole input file is missing.** Something failed to open and the error was skipped at a level above the row loop. ## Step 3: audit the write path for discarded returns With the gap on the write side, go looking for the errors that had nowhere to go: - `defer f.Close()` — a deferred call's return value is discarded, and `Close` is where a delayed write failure such as a full volume is reported. - `defer w.Flush()` — same discard, and if the flush fails the bytes are simply gone. - Bare `w.Write(...)` or `fmt.Fprintln(...)` calls whose results were never bound. - `f.Sync()` where the result was dropped. - `csv.Writer.Flush()`, which returns **nothing at all** — the error is only available from `csv.Writer.Error()`, so there is no discarded return value for a reviewer to notice. On a CSV write path this is the first line to look for. ## Step 4: reproduce on purpose A finding you cannot reproduce is a guess. Point the job at a small filesystem or a constrained volume so the writes actually fail with no space left, run it, and confirm the current code exits 0 with a short file. That both proves the diagnosis and gives you the regression test for the fix. Checking whether the volume was near capacity on the night in question, and whether the input volume grew, usually closes the loop on *why now* after months of clean runs. ## Step 5: make the next one impossible to miss Fix the specific discards, then remove the class of failure: - Flush and close explicitly on the success path and return those errors, keeping any deferred close purely as a safety net for early returns. - Have the job **verify its own output**: re-open the file it just wrote, count the rows, compare with what it read, and exit non-zero on a mismatch. This is the same check you just performed by hand, promoted into the job. - Emit the counts — read, written, skipped — in the run summary, so a drift of 25% is visible in the record of the run rather than discovered by someone reconciling totals a month later. ## What to say in the postmortem The honest finding is rarely "the disk filled". Disks fill. The finding is that the job had no way to tell anyone, because the only report of the failure was a value returned by a call whose result was thrown away — and that a green run was accepted as evidence of a complete file for months. The corrective action that matters is the self-verification, because it is the only one that would have caught a cause nobody predicted.

  • Why is re-reading the written file more useful than adding more logging to the job?
    Because logging reports what the program believes it did, and the program's belief is exactly what is wrong. Re-opening the artefact and counting rows is an independent measurement: it is the only number that can disagree with the job. It also converts directly into the permanent fix, since the same check can run at the end of every job.
  • What in the output file itself tells you a writer was never flushed?
    The end of the file. An unflushed buffer stops at whatever the last write-through boundary was, not at a record boundary, so the final line is usually truncated mid-record and the file size lands suspiciously close to a multiple of the buffer size. A file that ends on a clean record boundary points elsewhere — typically to records dropped on the read side.
  • How would you reproduce a full-disk failure to confirm the diagnosis?
    Run the job against a deliberately small or nearly full volume so the writes fail with no space left. Confirm that the current build still exits 0 with a short file — that reproduces the incident — then keep the setup as the regression test for the fix, which must now exit non-zero and name the failure.
  • Which single change would have caught this even if the cause had been something nobody predicted?
    Self-verification at the end of the run: re-open the output, count the rows, compare with the number read, and exit non-zero on any mismatch. Fixing the specific discarded errors closes one hole; verifying the artefact closes the class, because it does not depend on having anticipated the failure mode.

saying these in an interview costs you the question

  • Starts by reading code instead of establishing counts
  • Trusts the job's own log lines over the file on disk
  • Blames the input data without comparing rows read to rows written
  • Concludes the fix is more logging
  • Treats a zero exit code as evidence the file is complete