skip to content

How does itertools.accumulate with a custom binary function differ from functools.reduce?

level: middleimportance: should knowfreq 34%

answer

  1. Every step, not just the last
  2. Two-argument function, running value first
  3. Default operation is plain addition
  4. Seeding changes the output length
  5. Sequential float folding drifts

basics

~10 s

itertools.accumulate folds a binary function over the input and yields every intermediate result, one per item, lazily. functools.reduce performs the same left fold but returns only the final value. accumulate defaults to addition.

solid answer

~40 s

`itertools.accumulate(iterable, func)` calls `func(running, item)` for each item and yields the running value every time, so a five-item input produces five outputs: a running total by default, a running product with `operator.mul`, a running maximum with the builtin `max`. `functools.reduce` performs the identical left fold but returns just the last value, and takes its optional seed positionally. accumulate's seed is the keyword-only `initial=`, added in Python 3.8; it is yielded first, so the output then has one more element than the input, and it is the only way to get output at all from an empty iterable — without it, empty input yields nothing rather than a zero. Because the fold is sequential, a running float total accumulates rounding error; use `math.fsum` when only an accurate total matters.

code

python · 9 lines
python
import itertools, operator

prices = [3, 5, 2, 8]
print(list(itertools.accumulate(prices)))
print(list(itertools.accumulate(prices, operator.mul)))
print(list(itertools.accumulate(prices, max)))
print(list(itertools.accumulate(prices, operator.add, initial=100)))
print(list(itertools.accumulate([])))
print(list(itertools.accumulate([], initial=0)))

go deeper

for a junior

Recall the shape of the output: itertools.accumulate over [3, 5, 2] gives 3, 8, 10 — a running total, produced lazily, one value per input item.

for a middle

Explain the second argument: any two-argument callable, applied as func(running, item), turning accumulate into a running product, maximum or anything else — and know how initial= seeds the fold and lengthens the output.

for a senior

Show where the running series itself is the deliverable — cumulative columns, high-water marks, drift detection — and flag the numeric trap that a sequential float fold is not the same number as an exactly-rounded total.

for a principal

Own the readability line: a fold expressed as accumulate plus a named operator is elegant, but one whose function needs a paragraph of explanation belongs in an explicit loop, and money folds belong in decimal rather than either.

### The shape of the operation Both `itertools.accumulate` and `functools.reduce` are left folds: they thread a running value through a sequence, combining it with each item using a two-argument function. The only difference is how much of that thread you get to see. `reduce` hands back the last value and throws the rest away. `accumulate` yields the running value at every step, lazily, as an iterator. So `functools.reduce(operator.add, [3, 5, 2])` is `10`, while `list(itertools.accumulate([3, 5, 2]))` is `[3, 8, 10]`. When the *series* is the deliverable — a cumulative total per row, a high-water mark over time, a running balance — accumulate is the direct expression of it, and reduce is the wrong tool. ### The function argument The second positional argument is any callable of two arguments, invoked as `func(running_value, next_item)`. Omit it and addition is used. Common instantiations, all one-liners: ```python import itertools, operator itertools.accumulate(xs, operator.mul) # running product itertools.accumulate(xs, max) # running maximum itertools.accumulate(xs, min) # running minimum ``` Argument order matters as soon as the operation is not commutative. `func(running, item)` means subtraction folds as `running - item`, and a lambda written the other way round produces a completely different, entirely plausible-looking sequence. That is a silent bug, not an exception, so it is worth stating the order explicitly when you answer. The first output is always the first item itself — the function is not called for it, because there is nothing yet to combine it with. That is why `list(itertools.accumulate([3, 5, 2], max))` starts with `3`. ### The initial keyword `initial=` is keyword-only and was added in Python 3.8. It supplies a starting running value, which is yielded first, so the result has `len(input) + 1` elements. Two things make it more than cosmetic. First, it lets a fold start from a meaningful zero — an opening balance, an identity element of `1` for products. Second, it is the only way an empty input produces any output at all: bare `accumulate([])` yields nothing, so downstream code that expects at least one value breaks precisely on the empty edge case, while `accumulate([], initial=0)` yields a single `0`. ### Laziness, and what it buys accumulate returns an iterator. Nothing is computed until you pull, and you may stop pulling early — a running total over a very long stream can be abandoned as soon as it crosses a threshold, without folding the tail. It is also one-shot: iterate it twice and the second pass yields nothing, and there is no `len()` on it. ### The numeric trap A fold over floats is sequential addition, and sequential addition of floats loses low-order bits at every step. Consider a geocoding batch where each record contributes a small distance in kilometres and the pipeline wants both the running cumulative distance and a final total. The running series must be computed sequentially — that is what a cumulative series *is* — so it carries the drift by construction: the classic demonstration is that folding ten copies of `0.1` lands on `0.9999999999999999`, not `1.0`. The final number is where you have a choice: `math.fsum` performs an exactly-rounded summation and returns `1.0`, so a report that shows a cumulative column from accumulate and a grand total from `math.fsum` can display a last row that disagrees with the total by a few ulps. Decide deliberately which number is authoritative, and round for presentation rather than hoping the two agree. For money, neither float path is the answer — fold `decimal.Decimal` values instead. ### accumulate versus a plain loop accumulate threads exactly one value and feeds it straight back in. The moment you need two running values, a side effect, or a state that is not simply the previous output, the fold stops fitting and an explicit `for` loop is both clearer and shorter than the tuple-carrying lambda that would be required. The readable dividing line is roughly: if the folding function is a named operator or a builtin, use accumulate; if it needs a paragraph of explanation, use a loop. ### What an interviewer is checking That you know accumulate yields *every* step while reduce yields one; that you can name the call order `func(running, item)`; that you know `initial=` is keyword-only and changes the output length; and, at senior level, that a sequential float fold is not the same number as an accurate sum.

  • What does itertools.accumulate yield for an empty input, with and without initial=?
    Without `initial=` it yields nothing at all — not a zero — because there is no first item to become the running value. With `initial=0` it yields exactly `0` and stops. That asymmetry is why the seeded form is the safe one whenever downstream code expects at least one value, such as code that reads the last element.
  • In which order does itertools.accumulate call your function, and why does it matter?
    As `func(running_value, next_item)`, folding left to right. For commutative operations such as addition or `max` the order is invisible, but for subtraction, division or string concatenation, swapping the parameters yields a different and entirely plausible-looking sequence with no exception raised — a lambda written backwards is a silent bug.
  • When would you write a plain for loop instead of itertools.accumulate?
    When more than one running value is needed, when the step has side effects, or when the state is not simply the previous output — accumulate threads exactly one value. A loop is also the better choice once the folding function grows past a named operator or builtin, because a multi-line lambda folding a tuple of state is harder to read than the loop it replaces.

reduce is the closing balance printed at the bottom of a statement; accumulate is the balance column printed on every line.

saying these in an interview costs you the question

  • Says accumulate returns a list rather than an iterator
  • Expects only the final value, like functools.reduce
  • Thinks the function is called as func(item, running)
  • Passes initial as a positional argument
  • Expects accumulate on empty input to yield zero
  • Trusts a running float total where math.fsum is needed

context