Why does a Python generator expression produce nothing the second time you iterate it?
answer
- The second loop is silent, not loud
- It is an iterator, not a collection
- Iterating it again asks the same exhausted object
- Materialize once, or rebuild per pass
basics
~20 sA generator expression is a one-shot iterator. The first full pass drives it to completion, after which it is exhausted and every further loop ends immediately with no items and no error. Rebuild it, or materialize it once with list().
solid answer
~40 sIterating a generator object repeatedly calls `__next__()` until the underlying loop finishes and `StopIteration` is raised; at that point the suspended frame is done and discarded. Iterating the same object again asks for the next value, is told immediately that there is none, and the `for` body never runs. Nothing raises - you get an empty pass and a wrong number. The classic form is a nightly report generator that binds `rows = (parse(line) for line in lines)`, computes a row count, then computes a total over the same name and reports zero. Fixes: materialize once with `list(rows)` and iterate the list twice; wrap the expression in a function so each call returns a fresh generator; or restructure into one pass that accumulates both figures.
code
python · 4 linesrows = (n for n in range(5))
print(sum(rows)) # 10
print(sum(rows)) # 0 - already exhausted
print(list(rows)) # []go deeper
Remember the rule and the symptom: a generator expression can be looped over once, and the second loop quietly does nothing. If you need the values twice, call list() on it and use the list.
Explain the mechanics - the object is its own iterator with no rewind, the frame finishes and StopIteration is raised from then on - and name all three fixes: materialize, rebuild per pass, or accumulate in one pass.
Demonstrate you can spot it in review. Helpers that iterate an argument twice, aggregates computed over a shared name, and generators handed to worker threads are where this bites, and the failure is a plausible wrong number rather than a crash.
Own the convention. Decide where in your pipelines laziness ends - which layers may pass iterators and which must hand over materialized sequences - so single-pass semantics never cross a boundary where a caller will assume it can re-read the data.
### One object, one pass A generator expression evaluates to a generator object, and a generator object is an *iterator*, not a re-iterable collection. The distinction matters exactly here. A `list` is an iterable: every `for` loop over it asks it for a brand-new iterator and starts from the front. A generator object is its own iterator - asking it for an iterator hands back the very same object, positioned wherever the last consumer left it. There is no rewind operation anywhere in the protocol. So the first full consumption runs the hidden loop to completion. When the source is exhausted the generator raises `StopIteration`, its frame is finished and released, and the object is permanently marked as done. Every later `next()` on it raises `StopIteration` again straight away. A `for` loop treats that as 'the sequence has ended', so it executes zero iterations, silently. `list()` on it returns `[]`. `sum()` returns `0`. `max()` raises `ValueError` for an empty sequence unless you passed `default`, which is often the only reason anyone notices. ### The bug this actually produces Consider a nightly report generator that streams parsed rows: ```python rows = (parse(line) for line in lines) count = sum(1 for _ in rows) total = sum(r.amount for r in rows) # always 0 ``` The count is right, the total is zero, and nothing failed. The report goes out with a real row count and no money in it. Because the failure is silent and only shows on the *second* aggregate, it survives code review easily - the two lines look symmetric, and each one is correct in isolation. The same trap appears whenever a generator expression is passed around: a helper that takes an iterable and loops over it twice works perfectly with a list argument and silently returns nonsense with a generator argument. That is why functions that iterate their input more than once should either document that they need a sequence or defensively materialize with `list()` at the top. ### The three fixes, and when each is right **Materialize once.** `rows = list(parse(line) for line in lines)` - or just use a list comprehension in the first place. Now both aggregates work and you can index and re-sort. The cost is memory proportional to the data, which is precisely what the generator expression was avoiding, so this is the right answer when the data comfortably fits. **Rebuild the expression.** Wrap it in a function or a small `lambda` and call it once per pass: each call evaluates the expression again and returns a fresh generator object over a fresh iteration of the source. Memory stays flat, but you pay for two passes over the source - fine over an in-memory list, wrong over a network stream or a file handle you have already consumed, and outright incorrect if the source itself is a one-shot iterator. **Single pass with two accumulators.** Loop once and update both figures inside the loop body. This is the only fix that works when the data is too large to hold *and* the source can only be read once. It is more code, and it is the answer an interviewer is usually fishing for in a data-pipeline context. ### Related traps worth naming Re-binding does not help. `a = (x for x in xs); b = a` gives two names for one generator object, so consuming through `a` empties it for `b`. Copying does not help either - a generator object cannot be copied or pickled. Partial consumption is sticky. `next(g)` inside a guard clause, or a `for` loop with a `break`, leaves the generator positioned mid-stream; the next consumer picks up from there rather than the beginning. That makes 'peek at the first item then process everything' a subtle off-by-one source unless you deliberately re-attach the peeked item. Sharing one generator object between threads is worse than a stale read. The object holds a single suspended frame, so concurrent consumers interleave items unpredictably - each item goes to exactly one thread, which is occasionally what people want and usually not - and CPython raises `ValueError` if a generator is re-entered while it is already running. If several workers must process the same values, materialize the data or give each worker its own expression; if they should split the work, put the items in a `queue.Queue` where the hand-off is explicit and thread-safe. Finally, exhaustion is not the same as emptiness. `list(g)` returning `[]` tells you nothing about whether the source was empty or whether someone upstream already drained it - which is exactly why the bug reads as 'no data today' rather than as an error.
- How would you get both a count and a total from a single pass?Loop once and keep two accumulators: `for r in rows: count += 1; total += r.amount`. That is the only shape that works when the data will not fit in memory and the source can be read once. Reserve the two-aggregate style for sequences you have already materialized.
- Is it safe for two threads to pull from the same generator object?No. The object wraps one suspended frame, so items are split unpredictably between the consumers rather than delivered to both, and CPython raises `ValueError` if the generator is re-entered while already running. Give each worker its own expression, or push items through a `queue.Queue` so the hand-off is explicit.
- Does a `for` loop over an exhausted generator raise anything?No. The generator raises `StopIteration`, the `for` statement catches that as the normal end-of-sequence signal, and the body simply never executes. That silence is what makes the bug expensive: you get a plausible zero rather than a traceback pointing at the line.
saying these in an interview costs you the question
- Believes a generator object restarts when a new for loop begins
- Expects an error on the second pass instead of silence
- Says list() on an exhausted generator returns the original values
- Thinks assigning it to a second name gives a fresh generator
- Uses one generator object for two aggregates and trusts both