skip to content

When does math.fsum give a different total than the built-in sum() over the same floats?

level: middleimportance: should knowfreq 32%

answer

  1. Each addition rounds, and roundings accumulate
  2. The built-in got smarter in 3.12
  3. One correction float is not always enough
  4. Exact partial sums, rounded once at the end

basics

~20 s

math.fsum tracks exact partial sums and rounds once at the end, so its total is always correctly rounded. The built-in sum() has compensated since Python 3.12, but its single correction term can still drift across widely separated magnitudes.

solid answer

~40 s

Each float addition rounds, so a left-to-right summation accumulates error, and small values are absorbed once the running total is large. `math.fsum` avoids that by keeping a list of exact non-overlapping partial sums and rounding only once at the end, which makes its result the correctly rounded exact total and independent of input order. The built-in `sum()` is no longer naive: since **Python 3.12** it uses Neumaier compensated summation for floats, so `sum([0.1] * 10)` is exactly `1.0` on 3.14. One correction term is not always enough, though — `sum([1e40, 1e20, 1.0, -1e40, -1e20])` returns `0.0` where `math.fsum` returns `1.0`. I use `fsum` when magnitudes span many orders or large terms cancel, and accept its roughly 2-3x cost; a hand-written `total += v` loop gets no compensation at all.

code

python · 6 lines
python
import math

values = [1e40, 1e20, 1.0, -1e40, -1e20]

print(sum(values))        # 0.0 - the 1.0 is absorbed
print(math.fsum(values))  # 1.0 - exact total, rounded once

go deeper

for a junior

Recall that a float addition rounds and that repeated additions accumulate that rounding, so a long total can end a few ulps off. Know that math.fsum exists as the accurate alternative to sum() for floats.

for a middle

Explain the mechanics: absorption of small values into a large running total, Neumaier's single correction term inside sum() since 3.12, and Shewchuk exact partial sums inside math.fsum. Be ready to name a list where the two still disagree.

for a senior

Show the judgement: quantify the gap with math.isclose instead of calling a total wrong, decide when the 2-3x cost of fsum is warranted, and recognise heavy cancellation or wide magnitude spread as the trigger. Also handle the OverflowError path.

for a principal

Own the policy question of which quantities may live in binary floats at all. Set the boundary between accurate-enough summation, exact decimal arithmetic for money, and reproducibility requirements across differently ordered or parallel reductions in a pipeline.

### The mechanism: every addition rounds A Python `float` is an IEEE 754 binary64 value with 53 bits of significand. Every arithmetic operation on two floats computes the exact mathematical result and then rounds it to the nearest representable float. One rounding is at most half an ulp (unit in the last place) of error — tiny. The problem is that a summation performs one rounding **per addition**, and those errors do not cancel. Add `n` values left to right and the worst-case relative error grows with `n`; add values of very different magnitudes and the small ones are rounded straight out of the running total, because the exact sum needs more than 53 bits to hold. This is the "absorption" failure: once the accumulator reaches `1e16`, adding `1.0` to it cannot change it, so the `1.0` is silently discarded. The error is a property of the accumulation order, not of the inputs — which is why naive float summation is not even commutative. ### What the built-in `sum()` actually does today The interview folklore is that `sum()` drifts and `math.fsum()` fixes it. That was true through Python 3.11, where `sum([0.1] * 10)` returned `0.9999999999999999`. **In Python 3.12 the built-in `sum()` was changed to use Neumaier compensated summation for floats** (and for mixed ints and floats). It keeps a second float `c` that carries the low-order bits lost by each addition, and adds `c` back at the very end. On 3.12 and later — including 3.14 — `sum([0.1] * 10)` returns exactly `1.0`, and the result is far less dependent on ordering. So the honest answer to "does `sum()` drift?" starts with *when*: on any currently supported interpreter it compensates, and for ordinary data — a column of prices, a list of durations, a million samples of the same rough magnitude — it now agrees with `math.fsum` bit for bit. Note what is **not** compensated: a hand-written accumulator. `total = 0.0` followed by `total += v` in a `for` loop, or `functools.reduce(operator.add, values)`, or `itertools.accumulate(values)`, all still perform plain left-to-right addition and still produce `0.9999999999999999` on 3.14. The compensation lives inside `sum()`, not in the `+` operator. ### What `math.fsum` does differently `math.fsum` implements Shewchuk's exact-summation algorithm. Instead of one running total plus one correction term, it maintains a **list of non-overlapping partial sums** whose exact mathematical sum equals the exact sum of everything consumed so far. Each incoming value is merged into that list using exact two-sum arithmetic, so nothing is ever thrown away mid-stream. Only at the very end are the partials collapsed into a single `float`, with **one** rounding. The guarantee is therefore stronger than "more accurate": the result is the *correctly rounded* value of the exact sum, and it is independent of the input order. Neumaier's single correction term is not always enough, because the discarded bits can themselves be too large to represent alongside `c`. Three or more widely separated magnitudes that cancel is the standard counterexample, and it still separates the two functions on 3.14: ```python import math values = [1e40, 1e20, 1.0, -1e40, -1e20] sum(values) # 0.0 -- the 1.0 is absorbed and never recovered math.fsum(values) # 1.0 -- the exact total, rounded once ``` ### The cost, and when to pay it `math.fsum` is roughly two to three times slower than `sum()` on a long list of floats, and it allocates a small partials list. It also has a sharper failure mode: an intermediate partial can overflow, and `math.fsum([1e308, 1e308, -1e308, -1e308])` raises `OverflowError`, where `sum()` quietly returns `inf`. A `nan` anywhere propagates through both. Reach for `math.fsum` when the data genuinely spans many orders of magnitude, when there is heavy cancellation between large positives and large negatives, when a residual or a difference of two sums is the quantity of interest, or when a total must be reproducible regardless of the order rows arrive in. Reach for `sum()` everywhere else — it is the default for a reason, and since 3.12 it is a good one. ### What neither of them is for Accuracy is not exactness. If the requirement is that a total of money be *exact* to the cent, no binary summation is the answer, however well rounded — that is a job for exact decimal arithmetic. And if you are averaging, `statistics.mean` already sidesteps the whole question by accumulating in exact rational arithmetic (`statistics.mean([0.1] * 10)` is exactly `0.1`), while `statistics.fmean` is the fast float path. The senior version of this answer is to stop asserting a total is "wrong" and instead say by how much: compute both totals, take the difference, and compare it against a tolerance with `math.isclose`. Most of the time the gap is a few ulps and irrelevant; when it is not, the data has told you it needs `math.fsum` — or that it should never have been in floats at all.

  • Can math.fsum fail on input that the built-in sum() handles?
    Yes. Its exact partial sums can overflow mid-stream even when the final total is representable, and it raises `OverflowError: intermediate overflow in fsum` in that case, where `sum()` silently returns `inf`. A `nan` in the input propagates through both. So `fsum` trades a quiet wrong answer for a loud failure, which is usually what you want, but it does mean the call needs handling in code that ingests unvalidated float data.
  • How would you accurately total a stream of floats too large to hold in memory?
    `math.fsum` accepts any iterable, not just a list, so you can hand it a generator and it consumes the stream in one pass. Its memory is the partials list, which grows only with the number of distinct exponent ranges seen, not with the number of values, so it stays small in practice. If you also need a running total exposed mid-stream, keep your own Neumaier correction term rather than re-summing.
  • Why does statistics.mean([0.1] * 10) return exactly 0.1?
    `statistics.mean` converts the values to exact rationals and does the accumulation and the division in exact arithmetic, converting back to `float` only once at the end. That gives a correctly rounded mean rather than the total-then-divide result, at a real speed cost. `statistics.fmean` is the float fast path when you want the speed and the ordinary float accuracy instead.

A cashier who rounds the running subtotal to the nearest cent after every item ends the day off by a few cents; one who keeps the exact ledger and rounds only the final line is off by at most half a cent, no matter how many items there were.

saying these in an interview costs you the question

  • Claims the built-in sum() over floats is exact
  • Says math.fsum is simply a faster sum()
  • Insists sum([0.1] * 10) is 0.9999999999999999 on 3.14
  • Thinks a manual += loop matches sum()'s accuracy
  • Believes math.fsum can never raise OverflowError
  • Reaches for fsum to make currency totals exact

context