skip to content

Why does sum() over a generator of booleans work as a match counter in Python?

level: middleimportance: should knowfreq 38%

answer

  1. What type is True, really?
  2. Booleans take part in arithmetic
  3. One per match, added up
  4. The empty total is the identity element
  5. No single element can settle a count

basics

~20 s

bool is a subclass of int, with True equal to 1 and False to 0, so adding booleans totals the matches. sum(pred(x) for x in items) counts without building a list, and returns an int.

solid answer

~40 s

`bool` is a subclass of `int`: `True` behaves as `1` and `False` as `0` in arithmetic, so `sum()` over booleans totals the number of true ones and returns an `int`, not a `bool`. That gives the lazy counting idiom `sum(pred(x) for x in items)`, or equivalently `sum(1 for x in items if pred(x))`, which counts matches without allocating the `len([x for x in items if pred(x)])` list. The empty case is `0`, because `sum()` starts from a zero accumulator. The important difference from `any()` and `all()` is that counting cannot short-circuit — the total is only known after every element has been seen — so reach for `sum()` when you need *how many* and for `any()`/`all()` when you only need *whether*, since the latter can stop at the first decisive element.

code

pycon · 9 lines
pycon
>>> flags = [True, False, True, True]
>>> sum(flags)
3
>>> isinstance(True, int)
True
>>> sum(1 for w in ["ok", "", "fail"] if w)
2
>>> sum(x > 2 for x in [])
0

go deeper

for a junior

Recall that True counts as 1 and False as 0, so sum() over booleans gives the number of true ones, and an empty input gives 0.

for a middle

Explain the subclass relationship between bool and int, contrast the two spellings of the idiom, and say why the version with a literal 1 is safer when the predicate is not guaranteed to return a bool.

for a senior

Show the judgement: counting consumes the whole source, so use any() for existence, keep the count for thresholds, and collect several counters in one explicit pass when the source can only be read once.

for a principal

Own the readability tradeoff — terse arithmetic over booleans is idiomatic but couples correctness to a predicate's return type, so decide as a codebase which spelling is the house idiom rather than mixing all three.

The idiom rests on one fact about the object model: `bool` is a subclass of `int`. `True` is an `int` whose value is 1 and `False` is an `int` whose value is 0, so booleans participate in arithmetic like any other integer — `True + True` is `2`, and `sum([True, False, True])` is `2`. `sum()` needs nothing special to count; it is doing ordinary integer addition over values that happen to be 0 and 1. ### The two spellings ```python sum(pred(x) for x in items) # add the predicate's own True/False sum(1 for x in items if pred(x)) # add 1 per surviving element ``` Both are lazy and both allocate nothing beyond the generator object and the running total. The first requires the predicate to return an actual `bool` (or an int); if it returns a truthy *object* — a non-empty string, a list — the addition raises `TypeError` or, worse, adds something meaningful, so it is only safe when the predicate genuinely returns a boolean. The second is immune to that because it uses the standard truth test in the `if` clause and adds a literal `1`, which is why it is the safer default when the predicate is someone else's function. Both start from `0`, so an empty input counts to `0`. The eager alternative, `len([x for x in items if pred(x)])`, is correct and readable but builds a list solely to measure it — an allocation proportional to the number of matches, thrown away on the next line. On a small in-memory container that is fine; over a large or streamed source it is the same eagerness defect that spoils a reduction elsewhere. ### Why counting cannot short-circuit `any()` and `all()` can stop early because one element can settle the answer. A count has no such element: the total is only known once every candidate has been examined, so `sum()` necessarily consumes the whole iterable. This has two practical consequences. First, `sum()` over an unbounded source never returns — the lazy argument saves memory, not time. Second, when you only need to know *whether* something matched, counting everything and comparing to zero is strictly more work than asking `any()`; `if sum(...) > 0` is a common and needless pessimisation of `if any(...)`. There is one place the count genuinely beats the existence check even for a yes/no question: thresholds. "Are more than three items failing?" cannot be answered by `any()`, and short-circuiting can be recovered by hand — count in a loop and break once the threshold is crossed, or slice the source so the count is bounded. ### Related counting tools and their limits `sum()` counts one predicate. When you need several counts over the same pass, adding several generator expressions means several passes; a single explicit loop with several counters, or a `collections.Counter` keyed by the classification, is clearer and reads the source once — which matters when the source is a one-shot iterator that cannot be walked twice at all. `sum()` also takes a starting value as its second argument, so a running total can be seeded, and it refuses to concatenate strings on purpose — `str.join` is the tool there. Note also that `sum()` over booleans returns an `int`: `sum(flags) == 3` is a count, and if some caller expects a `bool` back, `bool(sum(flags))` (or better, `any(flags)`) is the explicit conversion. ### In review The idiom is idiomatic Python and worth recognising instantly when reading, but it deserves a moment's thought when writing. `sum(1 for x in items if pred(x))` states "count the matches" clearly. `sum(pred(x) for x in items)` is terser but couples correctness to the predicate's return type. And when the surrounding code only branches on whether the count is non-zero, the honest expression is `any()`, which says what is meant and stops early. The behaviour described here is stable across the Python 3 line and unchanged on 3.14. ### One more corner worth knowing Because `bool` subclasses `int`, booleans also work as indices and as multipliers: `["no", "yes"][flag]` and `count += flag` are both legal, and `sum()` is only the most useful member of that family. The relationship is one-directional — every `bool` is an `int`, but `isinstance(1, bool)` is `False`, so a function that must distinguish a genuine boolean from the integer 1 has to test `isinstance(value, bool)` before treating it as a number. That distinction rarely matters when counting and matters a great deal when serialising.

  • When should you use any() instead of comparing a count to zero?
    Whenever the question is only *whether* something matched. `any(pred(x) for x in items)` stops at the first match, while `sum(...) > 0` must examine every element and never returns at all over an unbounded source. The count is the right tool only when a threshold or the actual number matters.
  • What breaks if the predicate returns a truthy object rather than a bool?
    `sum(pred(x) for x in items)` then adds the objects themselves — a `TypeError` for strings, or a meaningless numeric total for numbers. `sum(1 for x in items if pred(x))` is immune because the `if` clause applies the standard truth test and the value added is always the literal 1.
  • How would you produce several counts in one pass?
    Do not write several generator expressions — each one walks the source again, which is impossible for a one-shot iterator. Use an explicit loop with a counter per category, or classify each item and feed the labels to a `collections.Counter`, so the source is consumed exactly once.

saying these in an interview costs you the question

  • Thinks bool is unrelated to int in Python
  • Says sum() over booleans returns True or False
  • Uses len() on a list built only to be measured
  • Expects sum() to short-circuit like any()
  • Writes sum(...) > 0 where any() is meant
  • Repeats the generator over a one-shot iterator

context