skip to content

Why can a list be looped over repeatedly while a generator object is exhausted after one pass?

level: middleimportance: should knowfreq 70%

answer

  1. Ask what happens on the second pass
  2. Fresh cursor, or the same cursor?
  3. __iter__ returns self versus a new object
  4. Exhausted means empty, never an error
  5. Materialize once, or call the function again

basics

~20 s

Each loop over a list calls iter() and gets a brand-new cursor starting at the first element. A generator object is already the cursor: its iter returns self, so a second loop resumes at the end and finishes without running the body.

solid answer

~50 s

`for x in xs` calls `iter(xs)` first. For a list that returns a fresh iterator with its own position every time, so each loop starts at the beginning and two nested loops never interfere. A generator object *is* an iterator — its `__iter__` returns `self` — so the second loop receives the same, already-drained cursor and completes without producing anything. Crucially nothing raises: an exhausted cursor reports end of iteration, which is indistinguishable from an empty source, so you get a silently wrong result instead of a traceback. The fixes are to materialize once with `list()` or `tuple()` when the data is bounded, to call the generator function again for a fresh generator object, or to make the source an object whose `__iter__` builds a new generator per call. `dict` views, `range` and `str` are re-iterable; generator objects, file objects and `enumerate()` results are not.

code

python · 10 lines
python
def timestamps():
    yield 100
    yield 90


container = [100, 90]
print(sum(container), sum(container))

stream = timestamps()
print(sum(stream), sum(stream))

go deeper

for a junior

Recall the observable behaviour: looping a list twice gives the items twice, looping a generator object twice gives them once and then nothing. Knowing that list() captures the items for reuse is enough at this level.

for a middle

Explain the mechanism, not just the symptom: iter() on a container mints a new cursor, iter() on a generator object returns self, so the position survives across loops. Then name the fixes and what each one costs.

for a senior

Emphasise the silence. No exception is raised, so an accidental double read ships as an empty batch and a wrong result. Say how you would surface it — materialize at the boundary, document laziness, or assert the consumption pattern in a test.

for a principal

Frame it as an interface contract across teams: which layers are allowed to return lazy values, where a stream gets materialized, and what memory budget that boundary implies. An implicit one-shot convention is a recurring defect source rather than a style choice.

### The rule in one sentence `for x in xs` always calls `iter(xs)` first. Whether the loop can be repeated depends entirely on what that call returns: a **fresh cursor** (re-iterable) or **the same cursor you already drained** (one-shot). ### Containers hand out new cursors A list stores its elements. Its `__iter__` builds a new iterator object each time it is called, and each of those iterators carries its own position, starting at index 0. That is why two nested loops over the same list do not interfere, why `sum(xs)` followed by `for x in xs` both see the full contents, and why you can pass the same list to five different functions without anything being consumed. The elements are the list's state; the position is the cursor's state, and cursors are disposable. The same is true of a tuple, a `str`, a `set`, a `range` object and a `dict` view. A `dict` view is a particularly clean demonstration: the view is re-iterable and stays a live window on the dictionary, while an iterator taken *from* the view is single-use. ### A generator object is already the cursor Calling a generator function does not run its body; it returns a **generator object**, which is an iterator. Its `__iter__` returns `self`, because there is nothing else it could return — the "position" is a suspended stack frame, and there is only one of them. Consuming it advances that frame to the end. A second `for` loop over the same name calls `iter()`, gets that same finished object back, immediately sees end-of-iteration, and runs the body zero times. Nothing raises. That is the important half of the answer: an exhausted iterator does not complain, it simply reports that it is done, which is exactly what an *empty* source reports. The protocol has no way to express "already consumed", so the failure surfaces as a silently wrong result rather than a traceback. ```python def timestamps(): yield 100 yield 90 stream = timestamps() print(sum(stream), sum(stream)) # 190 0 ``` ### What it looks like in production Consider a search-index rebuilder that reads documents from a lazy source. The first pass computes the newest timestamp in the batch so the job can detect a clock-skew artefact — a machine whose clock ran ahead, stamping documents into the future and poisoning recency ranking. The second pass then filters the documents against that threshold and writes them into the index. If the source is a generator object bound to one name, the second pass reads an exhausted cursor: zero documents are filtered, zero are written, and the rebuild "succeeds" against an empty batch. Worse, the symptom is masked downstream — the query tier is still serving an 83% cache-hit rate, so most user traffic never touches the newly-empty index, and the incident surfaces days later as slowly degrading results rather than as an alert. The two properties that make this bug expensive are exactly the two the protocol guarantees: laziness and silence. ### The fixes, and when each is right - **Materialize once**: `rows = list(source)`, then loop over `rows` as often as you like. Simplest and fastest for repeated full passes; the cost is holding the whole batch in memory, so it needs a stream you know is bounded. - **Call the generator function again**: `for r in documents(): ...` twice creates two fresh generator objects. Bounded memory, but the body re-runs — re-reading a file, re-issuing a query, or re-consuming an upstream iterator — so it costs the work twice and can observe different data on the second run. - **Make the source re-iterable**: an object whose `__iter__` builds and returns a *new* generator each time it is called behaves like a container to every caller, without materializing anything. - **Restart the underlying resource**: a file object is an iterator, but the file it wraps is seekable, so calling its `seek` method with 0 rewinds it for another pass. ### How to spot it in review Look for a name bound to a lazy object that is *read more than once*: a length or a `sum()` before the real loop, a debug log line that walks the stream, a validation pass, a conditional that peeks. Any of those turns the real pass into an empty one. When you have the object in hand, `iter(x) is x` answers the question directly — `True` means you are holding a one-shot cursor, not a container. And when you are the one designing the function, be explicit about which you return: a documented lazy contract is fine, an accidental one is a defect factory.

  • How would you catch this in code review?
    Look for a name bound to a lazy object that is read more than once: a `sum()` or a count before the real loop, a debug log line that walks the stream, a validation pass, a peek in a conditional. Any of those turns the real pass into an empty one. With the object in hand, `iter(x) is x` settles it — `True` means a one-shot cursor rather than a container.
  • Does calling the generator function a second time give you the same items?
    It gives you a fresh generator object that re-runs the function body from the top, so you get identical items only if the underlying source is unchanged and deterministic. If the body reads a file, queries a service or consumes another iterator, the second run repeats that work and may observe different data. That is precisely the trade against materializing the first pass with `list()`.
  • Why does the second loop not raise an error?
    An exhausted iterator keeps reporting end of iteration, and a `for` loop treats that as a normal finish, so the body simply never runs. The protocol has no way to express 'already consumed', which makes an exhausted stream indistinguishable from an empty one. The failure therefore appears as a wrong result — an empty write, a zero count — rather than as a traceback.

A list is a printed book you can reread; a generator object is a live radio broadcast — once the words have gone past, tuning in again just gives you silence unless you recorded them.

saying these in an interview costs you the question

  • Says a generator object can be reset or rewound
  • Expects the second loop to raise an error
  • Thinks list(gen) makes the generator itself replayable
  • Believes generators cache their output automatically
  • Assumes anything you can loop over is re-iterable

context