skip to content

Why does any() lose its short-circuit benefit when passed a list comprehension?

level: middleimportance: must knowfreq 55%

answer

  1. Think about when the argument is evaluated
  2. The builtin only sees what it was handed
  3. Brackets build; parentheses yield
  4. Short-circuit after the work is already paid
  5. Laziness has to reach the predicate

basics

~20 s

Arguments are evaluated before the call. A list comprehension runs the predicate on every element and allocates the whole list first, so any() short-circuits over a result that has already cost full time and memory.

solid answer

~50 s

`any()` and `all()` still stop at the first decisive element, but by then the damage is done: Python evaluates the argument expression fully before the call happens, so `any([check(x) for x in items])` builds a complete list — one `check()` call per item plus an allocation proportional to the input — and only then hands it over. Swapping the brackets for a generator expression, `any(check(x) for x in items)`, makes the argument lazy: the builtin pulls one value at a time and stops at the first truthy one, so the predicate may run once instead of a million times and nothing is allocated. It also handles unbounded sources, which the list form cannot. The list form is harmless when the input is small and already in memory and the predicate is cheap; it is a real defect when the predicate does I/O or the input is large.

code

python · 6 lines
python
def expensive(n):
    print("evaluated", n)
    return n % 7 == 0

print(any(expensive(n) for n in range(1, 30)))
print(any([expensive(n) for n in range(1, 30)]))

go deeper

for a junior

Remember that a list comprehension inside the call is finished before the call starts, and that swapping to a generator expression is usually the right default for any() and all().

for a middle

Explain argument evaluation order precisely: the comprehension runs the predicate once per element and allocates the list, so the builtin's short-circuit saves nothing. Say what the generator form saves in both time and memory.

for a senior

Be ready to size the impact — predicate cost, input size and where matches typically fall — and to point out the case the eager form cannot handle at all, such as a stream of unknown length.

for a principal

Frame it as a default worth encoding: the lazy form is never slower, so it belongs in the codebase's idiom guide and lint configuration rather than being rediscovered each time a predicate quietly grows an I/O call.

Python evaluates a function's arguments before the function runs. That single rule explains the whole difference between the two shapes: ```python any([check(x) for x in items]) # list built first, then any() runs any(check(x) for x in items) # generator handed over, any() pulls lazily ``` In the first line the list comprehension is a complete computation: it calls `check()` on every element of `items` and builds a `list` of that many booleans before `any()` is even entered. `any()` then does short-circuit — it looks at the first `True` in the list and returns — but the short-circuit saves nothing, because all the work it might have skipped has already happened. The cost is `n` predicate calls and `O(n)` memory, whatever the data looks like. In the second line the argument is a generator expression. Creating it does almost no work: it produces a generator object. `any()` then drives it with the iteration protocol, pulling one value at a time, and stops pulling at the first truthy result. If the first element already matches, `check()` ran exactly once and no list was created. The worst case (no match) is identical to the list version minus the allocation; the best and average cases can be dramatically cheaper. ### How big the difference is Three properties of the workload decide whether the difference is a curiosity or a bug: 1. **Predicate cost.** If `check()` is a comparison, `n` extra calls are noise. If it hashes a file, parses a payload, or crosses a network boundary, the list form multiplies that cost by the length of the input. 2. **Input size.** A three-element list is unmeasurable. A million-row cursor is not, and the list form also materialises a million-element result. 3. **Match position.** The lazy form's advantage is the distance from the start to the first decisive element. When matches are typically early, the lazy form is close to `O(1)`; when they are rare, the two converge. There is one thing the list form cannot do at all: a lazy source of unknown or unbounded length. `any(line.startswith("ERROR") for line in stream)` answers as soon as the first matching line arrives; the list version would first read the stream to its end, and over an endless source it never returns. ### The same trap in other shapes The defect is about eagerness, not about square brackets specifically. `any(tuple(...))`, `any(set(...))`, `any(list(map(check, items)))` and `any(sorted(...))` are all eager for the same reason — each builds a complete container before the call. Conversely `any(map(check, items))` and `any(filter(...))` are lazy in Python 3 and short-circuit exactly like a generator expression does. A subtlety worth knowing about generator expressions: only the outermost iterable is evaluated when the generator object is created; the rest of the pipeline runs on demand. So `any(check(x) for x in build_items())` calls `build_items()` immediately, and any laziness you wanted from `build_items` itself has to come from `build_items` being a generator too. ### When the list form is fine, and how to argue it in review Use the list form deliberately when you need the list for something else anyway — if you are about to report *which* items failed, you are not doing a short-circuiting existence check, you are doing a filter, and materialising is honest. Use it when the input is a small, already-materialised container and the predicate is a cheap attribute access; the readability of the two is identical and so is the runtime. Treat it as a defect when the predicate has any real cost or the input can grow. The review heuristic is simple: if the predicate calls anything you would put on a latency budget, the brackets are wrong. The fix costs two characters and never makes the code slower, which is why the lazy form is the sensible default even where the difference does not currently matter — inputs grow, predicates acquire I/O, and nobody revisits the brackets. All of this applies identically to `all()`, which stops at the first falsy element, and to `sum()`, `min()` and `max()` in the sense that they too accept a generator expression instead of a built container — although those must consume everything and so save only the allocation, not the predicate calls.

  • Does the same argument apply to map() and filter() in Python 3?
    Yes, and in their favour: `map()` and `filter()` return lazy iterators in Python 3, so `any(map(check, items))` short-circuits exactly like a generator expression. Only wrapping them in `list()`, `tuple()`, `set()` or `sorted()` reintroduces the eager evaluation.
  • When would you deliberately build the list anyway?
    When you need the elements afterwards — reporting which items failed, counting them, or logging them — because that is a filter, not an existence check. Also when the container is small and already in memory and the predicate is a cheap attribute read, where the two forms have indistinguishable cost.
  • Is a generator expression fully lazy?
    Almost. The outermost iterable is evaluated immediately when the generator object is created; everything else runs on demand as values are pulled. So `any(check(x) for x in load_all())` still calls `load_all()` straight away, and that call has to be lazy itself if you want the laziness to reach that far.

Asking "is anyone in the building?" is answered the moment you find one person; a list comprehension is the guard who first walks every floor writing down each room's occupancy, then hands you the report.

saying these in an interview costs you the question

  • Says any() short-circuits the list comprehension itself
  • Thinks the brackets are purely a style choice
  • Believes Python evaluates arguments lazily by default
  • Claims a generator expression is always slower
  • Cannot say what changes about memory use
  • Uses list() around map() and still expects short-circuiting

context