A nightly Go ETL job exits 0 but some output files are short — how do you find the dropped cleanup error?
answer
- exit 0 means someone had the error
- short output means a missing tail
- three layers: flush, sync, close
- reproduce with a quota, not with luck
- publish by rename, never in place
basics
~20 sSuspect a cleanup error the job threw away. Audit every write path for an unchecked Flush, Sync or Close, reproduce with a disk quota so the failure happens on demand, then make the job check those calls and publish output only after a successful close.
solid answer
~50 sA zero exit status with truncated output is the signature of a failure that was reported and discarded, not one that never happened. I start on the write path: every layer wrapping the file has to be finished and checked — a `bufio.Writer` that was never flushed loses its last buffer, and `defer f.Close()` drops the error the filesystem raises at writeback when a disk or quota filled up mid-run. I reproduce it deliberately by pointing the job at a small filesystem or a tight quota, which turns an intermittent night into a repeatable test. Then I sweep the codebase for deferred `Close` calls whose errors are discarded on write paths and fix them with a named result. The structural fix is to stop publishing partial files at all: write to a temporary name, flush, `Sync`, close with the error checked, and only then rename into place, so a failed run leaves nothing for the next reader to reconcile.
code
go · 26 linesfunc writeAtomic(dir, name string, data []byte) error {
f, err := os.CreateTemp(dir, name+".tmp")
if err != nil {
return err
}
tmp := f.Name()
err = func() error {
w := bufio.NewWriter(f)
if _, werr := w.Write(data); werr != nil {
return werr
}
if ferr := w.Flush(); ferr != nil { // buffer to descriptor
return ferr
}
return f.Sync() // descriptor to stable storage
}()
if cerr := f.Close(); cerr != nil && err == nil {
err = cerr
}
if err != nil {
_ = os.Remove(tmp)
return err
}
return os.Rename(tmp, filepath.Join(dir, name))
}go deeper
Take away the shape of the bug: a job can finish with status 0 and still have lost data, because an error was returned by a cleanup call and thrown away.
Be able to order the layers — flush the wrapper, sync the file, close the file — and say which failure each one hides when it is skipped.
Show the whole arc: reproduce with a quota, sweep the write paths, fix with a named result, then change the publish shape to write-then-rename so partial output is never visible.
Own the prevention: reconciliation inside the run, disk pressure as a tested condition, and one cleanup convention that a build-time check can enforce without noise.
### Read the symptom precisely Two facts are doing the work: the files are **short**, not corrupt, and the process exited **0**. Short means the tail is missing, which points at a buffer that was never drained or at a write that was rejected near the end of the file. Exit 0 means some code had an error value in hand and did not return it. Failures that were never detected produce different symptoms — a panic, a partial line, a wrong value — so the first hypothesis should be a discarded cleanup error rather than a bug in the transform. ### The three calls to audit, innermost first On a write path there are usually three chances to lose data, and they are lost in this order: 1. **The wrapper's flush.** `bufio.Writer` has no `Close`; if the code deferred only `f.Close()`, up to one buffer of the tail was never handed to the file. A `compress/gzip.Writer` is worse: skipping its `Close` omits the trailer and the file will not decompress at all. 2. **The file's close.** Delayed allocation means a full disk or exhausted quota can be reported at writeback rather than at `Write`, and `Close` is where that becomes visible. `defer f.Close()` discards it. 3. **Durability.** `Close` is not an fsync. If the machine or container died mid-run, `Sync` is the call whose absence explains a file the operating system had accepted but not stored. ### Make the failure repeatable An intermittent nightly failure is unfixable until you can cause it. Point the job at a small filesystem or apply a quota sized just under the expected output, then run one shard. If the job still exits 0 with short output, you have reproduced the defect in seconds and you have a regression test for the fix. This is far more reliable than reasoning about which of last week's runs hit the disk pressure. ### Sweep rather than guess Search the repository for deferred close calls and triage each one by whether the handle was written to. A lint pass that reports unchecked errors on deferred calls makes this systematic and can then be run in the build to stop the defect coming back; the reason the explicit `_ = f.Close()` form is worth adopting on read paths is precisely that it lets such a pass be enabled without drowning in false positives. ### Fix the shape, not just the call Checking the error is necessary but leaves a worse artefact behind: a half-written file under the real name, which the next stage will happily read. The durable pattern is write-then-rename. Create a temporary file in the same directory, write, flush the wrapper, `Sync`, `Close` with the error checked, and only then `os.Rename` onto the final path. Rename within a directory is atomic, so a reader sees either the previous file or the complete new one, never a truncated one. On Linux, fsyncing the parent directory after the rename makes the new name itself survive a crash. If any step fails, remove the temporary file and return the error. ### Make the job's exit status honest The last link is the one that let this reach production: whatever wraps the per-file work has to propagate the error to a non-zero exit. A worker loop that logs a failure and continues, or a `main` that ignores what its run function returned, converts a checked error back into the silent failure you started with. The rule that keeps this honest is that the run is only recorded as successful once every output file has been closed successfully. ### Prevention worth naming in the postmortem - Reconciliation as a job step: the number of records written is compared against the number read, in the same run, and the mismatch fails the job rather than waiting for a human days later. - Disk pressure treated as an expected condition in tests, not an operational surprise. - One documented cleanup shape for write paths, so a reviewer can recognise a missing check without reasoning from first principles every time. ### What to say in the interview Name the discarded value as the prime suspect, order the three layers correctly, reproduce with a quota, fix with a named result, and close with the write-then-rename shape and an honest exit status. That progression — symptom, mechanism, reproduction, fix, prevention — is what the question is testing.
- How would you reproduce this on demand rather than waiting for another bad night?Run one shard against a deliberately small filesystem or a quota set just below the expected output size. That forces the late allocation failure every time. If the job still exits 0 with a short file, the defect is confirmed in seconds, and the same setup becomes the regression test that keeps the fixed version honest.
- Once every cleanup error is checked, why still write to a temporary file and rename?Checking makes the run fail loudly, but it does not remove the truncated file already sitting under the final name for the next stage to read. Rename within a directory is atomic, so readers see either the old complete file or the new one. On failure you delete the temporary file and nothing partial is ever published.
- What would make this reach production again even after the write path is fixed?A caller that swallows the error. A per-file loop that logs and continues, or a top-level function whose returned error is not turned into a non-zero exit, restores the original silent failure. The run should only be recorded as successful after every output file has been closed successfully.
saying these in an interview costs you the question
- assuming short output means the input was short
- blaming the transform before auditing the write path
- adding a retry instead of checking the cleanup errors
- treating exit 0 as proof the run succeeded
- fixing the close check but still writing in place