skip to content

Generator Expression Laziness

A generator expression looks like a comprehension but yields lazily and can be consumed only once. The memory-versus-reuse tradeoff is standard, as is dropping the parentheses inside sum() or any().

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

In Python, how does `(x*x for x in nums)` differ from `[x*x for x in nums]`?

level: juniorimportance: must knowfreq 75%

answer

  1. Brackets or parentheses changes the object
  2. One builds now, one builds on demand
  3. Memory grows with input, or stays flat
  4. No len(), and only one pass

basics

~10 s

The square-bracket form builds the whole list in memory immediately. The parenthesised form builds a generator object that computes each value only when asked, uses near-constant memory, and can be iterated only once.

solid answer

~40 s

`[x*x for x in nums]` is a list comprehension: it runs the loop to completion at once and returns a `list` holding every result, so memory grows with the input. `(x*x for x in nums)` is a generator expression: evaluating it returns a generator object and runs no loop body yet. Each `next()` resumes it, produces one value and suspends again, so it holds one item at a time whatever the input size. That object is one-shot: once exhausted it stays exhausted, and it supports neither `len()` nor indexing. Choose the list when you need the values more than once, want random access, or want the source snapshotted before something mutates it. Choose the generator expression for a single streaming pass over something large, or when the consumer may stop early.

code

python · 9 lines
python
nums = [1, 2, 3, 4]
squares_list = [x * x for x in nums]
squares_gen = (x * x for x in nums)

print(type(squares_list))   # <class 'list'>
print(type(squares_gen))    # <class 'generator'>
print(len(squares_list))    # 4
print(next(squares_gen))    # 1
print(list(squares_gen))    # [4, 9, 16] - the 1 is already gone

go deeper

for a junior

Be ready to point at the brackets versus the parentheses and say which one holds every value in memory. Knowing that list() materializes the lazy form, and that a generator object has no length, is enough at this level.

for a middle

Explain the mechanics: evaluating the parenthesised form returns a generator object, and each next() runs the loop body until one value is ready, then suspends. Say why memory stays flat and why a second pass sees nothing.

for a senior

Show judgement about when materializing is the right call - repeated passes, indexing, or pinning values before the source mutates or closes - and when streaming is what keeps a job from being killed for memory.

for a principal

Own the guidance. A blanket 'always use generators' rule trades memory for silent single-pass bugs and sources that close before consumption. Frame it as a per-boundary decision, with materialization at API edges where laziness would leak.

### Two syntaxes, two different kinds of object The two forms look almost identical, and that is the trap. `[x*x for x in nums]` is a **list comprehension**. Python runs the entire loop straight away and hands back a `list` containing every computed value. `(x*x for x in nums)` is a **generator expression**. Evaluating it constructs and returns a *generator object* - `type()` reports `<class 'generator'>` - and the loop body has not run even once at that point. Nothing is computed until something pulls values out of it. The difference is not cosmetic; it changes what you are holding. A list of a million squares is a million pointers plus a million integer objects on the heap. A generator over the same source is a single small object holding a suspended execution state: `sys.getsizeof()` reports a couple of hundred bytes for it whether the source has ten items or ten million, because the values do not live inside it. What the generator does hold is a reference to the underlying iterable, so it keeps that alive - a generator expression over a huge list does not free the list. ### How the lazy one actually runs A generator expression compiles to an implicit function whose body is the comprehension loop, with the produced value handed out one at a time. When a `for` loop, `list()`, `sum()` or any other consumer calls `__next__()` on the object, that hidden function runs until it has one value ready, hands it over, and *suspends with its local state intact* - the loop counter, the partially consumed source iterator, everything. The next `__next__()` resumes exactly there. When the source runs out, the generator raises `StopIteration`, which `for` loops and the builtins catch as the normal end signal. That suspend-and-resume model buys three properties. Memory stays flat, because only one value is alive at a time. Work is skipped if the consumer stops early: `any(...)` or `next(...)` over a generator expression leaves the rest of the source untouched. And the values can be produced while the source is still arriving - reading a file, walking a large query result - instead of waiting for the whole thing. ### What laziness costs you A generator object is a one-shot iterator. Consume it fully and it is finished forever; a second `for` loop over the same object ends immediately with no error and no values, which is a silent wrong answer rather than a crash. It also has no length and no subscription: `len(g)` raises `TypeError`, and `g[0]` raises `TypeError` as well. You cannot sort it in place, you cannot look at it in a debugger without consuming it, and `copy.copy()` on it fails. If you need the data twice, materialize it once with `list()` and reuse that list, or write a small function that returns a fresh generator expression on each call. Laziness also moves *when* code runs. Exceptions raised by the loop body surface at the consumption site, not at the line that built the expression, and side effects inside the expression happen wherever the values are pulled. If the source is a file handle, a cursor or a lock-protected structure, that source has to still be valid at consumption time - a generator expression returned out of a `with` block is a classic bug, because the file is closed by the time anyone iterates. ### Choosing between them Reach for the list comprehension when the result is small and needed repeatedly, when you want indexing, sorting or a length, or when you deliberately want to snapshot values before the source changes underneath you. Reach for the generator expression when the sequence is large or unbounded, when it feeds straight into a single consuming call such as `sum()`, `max()`, `min()` or `str.join()`, or when the consumer can bail out early. For small inputs the difference is negligible and readability wins - and there are cases, `str.join()` among them, where the consumer materializes the iterable internally anyway, so the list form is no worse. One performance footnote for the modern interpreter: since Python 3.12 (PEP 709) list, dict and set comprehensions are *inlined* into the enclosing function's frame, removing the per-comprehension function-call overhead they used to pay. Generator expressions were deliberately left out of that change, because they genuinely need their own suspendable frame to be resumable. So on 3.12 and later the eager forms got measurably faster while the lazy one did not - a small extra reason not to reach for a generator expression on a ten-element list.

  • What does `sys.getsizeof()` report for a generator expression over a million items?
    A couple of hundred bytes, and the same figure for ten items - the generator holds a suspended frame, not the values. The equivalent list grows linearly with the source and, for integers, also drags in a separate object per element. Note the generator still keeps a reference to the underlying iterable, so it does not free a large source list.
  • You need the values twice. What is the cheapest correct fix?
    Materialize once with `list()` and iterate the list twice, accepting the memory. If the data is too large for that, either write a small function that returns a fresh generator expression on each call and pay for two passes over the source, or restructure to compute both results in a single pass with two accumulators.
  • Can you index or slice a generator expression?
    No. A generator object supports only iteration: `g[0]` and `g[1:3]` both raise `TypeError`, as does `len(g)`. To get positional access you materialize it with `list()` first, or pull items with `next()` in order. That absence of random access is the main practical reason people go back to a list.

A list comprehension is a printed report: every page already exists and you can flip back and forth. A generator expression is the printer - it produces the next page only when you ask, and the pages it has already handed over are gone.

saying these in an interview costs you the question

  • Calls a generator expression just a faster list comprehension
  • Claims len() works on a generator object
  • Thinks the values are computed when the expression is created
  • Reuses one generator object for two passes and expects data
  • Uses a generator where the result is needed repeatedly

context

open as a page

Why does a Python generator expression produce nothing the second time you iterate it?

level: middleimportance: should knowfreq 60%

basics

~20 s

A generator expression is a one-shot iterator. The first full pass drives it to completion, after which it is exhausted and every further loop ends immediately with no items and no error. Rebuild it, or materialize it once with list().

open as a page

Why does `any(check(r) for r in rows)` stop early while `any([check(r) for r in rows])` does not?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Arguments are evaluated before the call runs. The bracketed form calls check on every row to build the list first, then any scans it. The bare generator expression is consumed only until the first true value, so the remaining calls never happen.

open as a page