skip to content

Streaming vs Materializing

Choosing between building the whole list and yielding items one at a time - the decision that says whether a job fits in memory. Interviewers hand you a huge file and watch for a generator.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

What is the difference between a generator expression and a list comprehension in Python?

level: juniorimportance: must knowfreq 78%

answer

  1. Square brackets versus round parentheses
  2. A container against a recipe
  3. Peak memory versus per-item overhead
  4. One value alive at a time
  5. No length, no index, one pass

basics

~20 s

A list comprehension builds the whole list in memory at once. A generator expression builds nothing up front: it produces items one at a time on demand, holds only the current value, can be consumed once, and supports neither len() nor indexing.

solid answer

~50 s

The syntax differs by one pair of brackets: `[f(x) for x in data]` is a list comprehension, `(f(x) for x in data)` is a generator expression. The comprehension runs the loop immediately and returns a real `list` holding every result, so peak memory scales with the number of items. The generator expression returns a lazy iterator object; nothing is computed until something iterates it, and only one result exists at a time, so memory is roughly constant regardless of input size. That laziness costs a resume per item, so for small inputs the list is usually the faster of the two. The generator is one-shot: once exhausted, iterating it again yields nothing. You also lose `len()`, indexing, slicing and repeated passes, which is exactly why a list still wins when you need any of those.

code

python · 10 lines
python
import sys

squares_list = [n * n for n in range(1000)]
squares_gen = (n * n for n in range(1000))

print(type(squares_gen).__name__)
print(len(squares_list))
print(sys.getsizeof(squares_list) > sys.getsizeof(squares_gen))
print(sum(squares_gen))
print(sum(squares_gen))

go deeper

for a junior

Be ready to state the bracket difference and what each expression returns: one gives a real list you can index and measure, the other gives a lazy one-shot iterator. Knowing that len() fails on the lazy form is the usual checkpoint.

for a middle

Explain the mechanics: nothing runs until iteration, one value is alive at a time, the resume-per-item cost, and how a chain of generator expressions collapses the moment an eager consumer such as sorted() or list() drains it.

for a senior

Show judgement about where the boundary sits in real code — which consumers keep a pipeline streaming, why a silently exhausted generator is a nastier bug than an exception, and how you would measure peak memory rather than assert it.

for a principal

Own the guidance your codebase gives: where laziness is worth its debugging cost, where a materialized list is the honest simpler choice, and how to stop lazy pipelines from leaking open resources across module boundaries.

## Two syntaxes, two very different objects `[n * n for n in data]` is a **list comprehension**. It runs the loop to completion at the moment the expression is evaluated and hands back a `list` object containing every result. `(n * n for n in data)` is a **generator expression**. It runs nothing: it hands back a generator object, a lazy iterator that computes the next value only when something asks for one. The difference is not a style preference. It is the difference between a container and a recipe for producing values. ## Peak memory A list comprehension over ten million rows holds ten million objects alive simultaneously, plus the list's own pointer array. A generator expression over the same source holds exactly one result at a time; the rest either has not been produced yet or has already been released. This is why the streaming form is the answer when the input does not comfortably fit in RAM, and why `sum(len(row) for row in rows)` is preferable to `sum([len(row) for row in rows])` on a large source: the bracketed form materializes a throwaway list purely to hand it to a function that only ever iterates it. A generator expression is not free, though. Every item costs a resume of the generator frame, which is more per-item work than the tight C-level loop inside a list comprehension. On a few thousand items the list is normally the faster of the two; the generator wins when the collection would be large enough that allocation, cache pressure or paging dominates, or when the consumer stops early. ## One-shot, and what you give up A generator is an iterator over itself. Once it is exhausted it stays exhausted: a second `sum()` over the same generator object returns `0`, silently, with no error. That silence is the single most common bug in this area. You also give up everything a sequence offers: - `len()` raises `TypeError` — the length is not known without running the whole thing. - Indexing and slicing raise `TypeError`; the closest lazy equivalent is `itertools.islice`, which consumes forward and cannot go backwards. - `reversed()` does not work. - Anything needing two passes (a mean and then a variance, a maximum and then a filter against it) needs either two independent generators, one restructured single pass, or a materialized list. ## Laziness is contagious in one direction Chaining generator expressions builds a pipeline that still streams: `stripped = (line.rstrip() for line in handle)` followed by `wanted = (line for line in stripped if line)` reads one line through both stages before touching the next. But a single eager consumer collapses the whole pipeline back into memory: `list()`, `sorted()`, `len()`, `tuple()` and `str.join` all drain their input completely. Streaming end to end means the *consumer* must be streaming too — `sum()`, `min()`, `max()`, `any()`, `all()`, a `for` loop, or a writer that emits as it goes. ## Two gotchas worth knowing The **outermost iterable is evaluated eagerly** when the generator expression is created; everything else is deferred. So `gen = (x * 10 for x in data)` captures the object `data` refers to at that instant. Rebinding the name `data` afterwards does not change what the generator iterates — but *mutating* that same object does, because the generator has not walked it yet. That mix of eager and lazy surprises people. Second, a generator expression that is the sole argument to a call needs no extra parentheses: `sum(n * n for n in data)` is legal, while `sum(n * n for n in data, 0)` is not. ## Version notes Both forms have their own scope in Python 3, so the loop variable never leaks into the enclosing function. CPython 3.12 inlined list, set and dict comprehensions into the enclosing function (PEP 709), removing a function call per comprehension and making them measurably faster; generator expressions were not part of that change and still create their own frame. None of the semantics above changed through 3.14. ## The decision rule Stream when the data is large, the consumer is a single-pass aggregate, or you may stop early. Materialize when the data is small and bounded, when you need `len()`, indexing or several passes, or when the source must be released before the results are used.

  • Why is sum(n * n for n in rows) usually preferable to sum([n * n for n in rows])?
    `sum()` only iterates its argument, so the bracketed version builds an entire throwaway list just to feed a single pass. The unbracketed generator expression keeps one value alive at a time, so peak memory stays flat no matter how many rows there are. On a small collection the list version can be marginally faster, but it buys nothing and scales badly.
  • What happens if you iterate the same generator expression object twice?
    The second iteration yields nothing. A generator is its own iterator and does not rewind, so once exhausted it simply raises StopIteration immediately — meaning `sum()` returns 0 and `list()` returns an empty list, with no exception to warn you. If two passes are genuinely needed, either build the generator twice from the source, materialize a list, or restructure into one pass that computes both results.
  • Does the loop variable of a comprehension leak into the enclosing scope in Python 3?
    No. Both list comprehensions and generator expressions have their own scope, so the loop name is not visible after the expression, and it does not clobber an existing variable of the same name. That was true of Python 2 list comprehensions, which did leak, and it is a classic version-trivia follow-up.

A list comprehension is a printed report: every page exists before you read a word. A generator expression is a ticket printer: it prints the next ticket only when someone steps up, and the used ones are gone.

saying these in an interview costs you the question

  • Calls a generator expression a faster list comprehension
  • Expects len() to work on a generator expression
  • Thinks reusing an exhausted generator restarts it
  • Believes generators are always faster than lists
  • Assumes a generator expression computes nothing at creation
  • Wraps a generator in list() then claims it still streams

context

open as a page

How do you stream a 40 GB log file in Python without loading it into memory?

level: middleimportance: must knowfreq 70%

basics

~20 s

Iterate the open file object directly with a for loop: it is its own iterator and hands back one buffered line at a time. Never call read() or readlines(), and keep every downstream stage lazy so nothing collects the whole file.

open as a page

When does materializing a list beat streaming in a Python data pipeline?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Materialize when you need more than one pass, a length, indexing or a global ordering, when the source must be released before the results are used, or when the data is small enough that a list is simpler. Streaming only buys bounded peak memory for one forward pass.

open as a page

What does itertools.tee cost in memory when one of its branches runs far ahead?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

itertools.tee buffers every item the fastest branch has consumed but the slowest has not. If one branch is drained before the other starts, that buffer holds the entire stream, so tee costs as much as building a list and adds indirection.

open as a page