A Go ETL job exports float64 readings with %.3f and re-ingests them; totals drift. How do you diagnose it?
answer
- three decimals is a grid, not a value
- many inputs, one output string
- the residuals do not cancel
- round once, at the source
- prove it with a re-parse over the real corpus
basics
~20 sFixed-precision formatting is lossy: each re-ingested row can differ by half a unit in the last printed place, and the errors accumulate. Prove it with a round-trip test, then emit shortest-unique text or round once at the source.
solid answer
~50 sStart by proving the export is the lossy step rather than the arithmetic. Take the real corpus, format each value the way the job does, re-parse it, and compare with the original: `%.3f` collapses every value within half a milli-unit onto one string, so the differences are non-zero and biased by whatever the data looks like. Summing the absolute differences should reproduce the observed drift to within the noise; if it does not, the loss is somewhere else. The fix has two shapes. If the exported text is meant to reconstruct the value, format with `strconv.FormatFloat(x, 'g', -1, 64)` or `%v`, which is guaranteed to reparse identically. If a fixed scale really is the contract, quantise **once at the source** so the stored value already equals the printed value, and never recompute a total from a rounded export — recompute it from the source rows. For money, stop carrying floats at all and use integer minor units.
code
go · 16 linesfunc checkRoundTrip(t *testing.T, f float64) {
s := strconv.FormatFloat(f, 'g', -1, 64)
back, err := strconv.ParseFloat(s, 64)
if err != nil {
t.Fatalf("ParseFloat(%q): %v", s, err)
}
if math.IsNaN(f) {
if !math.IsNaN(back) {
t.Fatalf("NaN lost through %q", s)
}
return
}
if back != f {
t.Fatalf("%v formatted as %q parsed back as %v", f, s, back)
}
}go deeper
Understand that printing with a fixed number of decimals throws information away, so a value that goes out and comes back is not always the value you started with.
Explain the mechanism: a fixed precision maps a whole interval of float64 values onto one string, and re-parsing returns the grid point. Know that formatting with precision -1 avoids it entirely.
Diagnose it rather than assert it: re-parse the real corpus, sum the residuals, check the magnitude against the observed drift, and rule out duplication with the per-row error bound. Then pick between a lossless format and rounding once at the source.
Own the standing rule: which pipeline stage holds the authoritative value, that aggregates are never recomputed from a rounded export, and that a change to output precision is treated as a contract change with a reconciliation plan.
## What actually happened `%.3f` prints three digits after the decimal point. That is a projection of the `float64` value onto a grid of thousandths, and it is not injective: every value within half a thousandth of a grid point produces the same string. Re-parsing gives you the grid point, not the value you had. Per row the error is at most 0.0005; over millions of rows, unless the residuals happen to be symmetric around zero, they accumulate into a visible total. The accumulation is the part people find surprising, and it is worth being explicit about why the errors do not cancel. They cancel only if the fractional parts of your data are uniformly distributed. Real measurement feeds are not: sensors quantise to their own resolution, unit conversions multiply by constants like 2.54 or 0.001, and prices cluster on values ending in 9. Any of those makes the rounding residual biased, and a biased residual times ten million rows is a number the finance or ops report shows. ## Confirming the cause rather than assuming it The postmortem should not rest on "floats are inexact". Run the export function over the actual corpus and measure: 1. **A round-trip property test.** For each value, format it exactly as the job does, parse it back, and record the difference. Assert the *shortest-unique* formatting round-trips exactly — it will — so you have a control showing the pipeline machinery is sound and only the chosen precision is lossy. 2. **Sum the residuals, signed and absolute.** The signed sum should match the observed drift in magnitude and direction. If the observed drift is much larger, something else is also wrong: a double-counted batch, a unit mismatch, a re-ingest that reads its own output twice. 3. **Check the special values.** Any `NaN` or infinity in the feed changes the story completely: `NaN` propagates through a sum and makes the whole total a `NaN`, and `encoding/json` refuses to marshal either, so their presence usually shows up as a failed record rather than as drift. Branch on `math.IsNaN` in the round-trip test, because `NaN != NaN` would otherwise report a false failure on every one. 4. **Compare against a exact-arithmetic control** for a sample: sum the same rows with `math/big`'s rational type, and you have a reference total that no float rounding touched. ## The two honest fixes **If the text exists so the value can be reconstructed** — an intermediate file, a fixture, a queue payload, anything a program re-reads — then the format must be lossless. `strconv.FormatFloat(x, 'g', -1, 64)` emits the fewest digits that uniquely identify the value, and re-parsing it returns the identical bits. `%v` on a float does the same thing. The cost is uglier text and a variable number of digits; that is the price of a value you can reconstruct. **If a fixed scale is genuinely the contract** — the warehouse column is defined to three decimals, or the readings are only meaningful to a milli-unit — then round **once, at the source**, and store the rounded value as the authoritative one. After that, printing at three decimals is not lossy, because the value on the grid *is* the value. The disease is having two authorities: a full-precision value in memory and a rounded value in the export, with different totals, each defensible. And the standing rule regardless: **never recompute an aggregate from a rounded export**. Recompute it from the source rows and carry the total as its own quantity. A total derived from rounded parts is a different number than a rounded total, and reconciliation between two systems that disagree about which one they computed is where the days go. ## Money is a separate answer If any of these columns are monetary, the fix is not a better format. Carry money as `int64` minor units — cents, or a scaled integer with a stated exponent — so that addition is exact and equality means what it says, and format for display only at the very edge. Floats have no business in a ledger that has to reconcile with another system's ledger to the cent. ## What goes in the postmortem Name the lossy step precisely ("the export formatted with three decimals, the loader treated the export as authoritative"), give the measured residual per row and the extrapolated total, state which of the two fixes was chosen and why, and add the round-trip property test to the suite so a future change of format is caught by CI rather than by a reconciliation report. Add a check at the ingestion edge for `NaN` and infinities while you are there — they are cheap to detect and expensive to discover downstream.
- Why do the per-row rounding errors not cancel out over millions of rows?They cancel only if the fractional parts are symmetric around the grid. Real feeds are not: sensors quantise at their own resolution, unit conversions multiply by fixed constants, and prices cluster on particular endings. A biased residual multiplied by the row count is exactly the drift the report shows.
- How do you tell rounding drift apart from a double-counted batch?Bound it. The rounding residual per row is at most half a unit in the last printed place, so the maximum possible drift is that bound times the row count. If the observed gap exceeds it, rounding cannot be the whole cause and you are looking at duplicated or missing rows.
- The warehouse column is defined to three decimals and cannot change. What is the fix then?Round once at the source and treat the rounded value as authoritative, so the printed value and the stored value are the same number. Then recompute totals from those authoritative rows rather than from a full-precision copy, so the two systems agree on which quantity they are summing.
- What do you add to the test suite so this cannot come back?A property test that formats and re-parses the real export function over a corpus of production-shaped values, including special values, and asserts exact recovery for lossless formats or a bounded residual for the fixed-scale one. It turns a silent format change into a CI failure.
saying these in an interview costs you the question
- Blames float64 in general instead of naming the lossy step
- Increases the precision to %.6f and calls it fixed
- Assumes rounding errors always cancel over many rows
- Recomputes totals from the rounded export
- Rounds on the way out while keeping a full-precision authority
- Writes the round-trip test with == and gets false failures on NaN