When should a Python class's `__iter__` return `self` instead of a fresh iterator?
answer
- The question is really what the object is
- One shot versus many passes
- Ask whether the source can be replayed
- Nested loops share a single cursor
- Iterator ABC requires __iter__ returning self
basics
~20 sReturn self only when the object is itself a cursor over a source that cannot be replayed, such as a stream reader; it is then one-shot. A container that must survive repeated and nested loops builds a new iterator per call.
solid answer
~50 s`__iter__` returning `self` declares *I am an iterator*: the object carries the cursor and is consumed exactly once. That is right for a one-pass source -- a socket or file reader, a database cursor, a paginated feed -- and it is what `collections.abc.Iterator` requires, since the only reason an iterator defines `__iter__` at all is so it can be passed where an iterable is expected. It is wrong for a container. If a class holding data returns `self`, the second `for` over it yields nothing, a nested loop exhausts the shared cursor in the inner pass, and an innocent `in` test or `sum()` silently eats the data before the real consumer runs. A container should return a new generator or iterator object per call, so `list(obj)` twice gives the same answer twice. Decide by asking whether the source is replayable, then state the answer in the type itself.
code
python · 17 linesclass Countdown:
def __init__(self, n):
self.n = n
def __iter__(self):
return self
def __next__(self):
if self.n <= 0:
raise StopIteration
self.n -= 1
return self.n + 1
c = Countdown(3)
print(list(c))
print(list(c))go deeper
Remember the rule of thumb: return self only if the class also defines __next__ and is meant to be walked once. Otherwise hand back a new generator so the object can be looped again.
Explain the mechanics precisely -- one shared cursor versus a new one per call -- and name the four consequences: an empty second loop, broken nesting, consuming in/any() probes, and helpers that iterate twice.
Show you have debugged the silent version of this: a report that came out empty because something upstream had already consumed the object. Talk about declaring the contract with collections.abc.Iterator versus Iterable, and about naming the type so the constraint is obvious.
Own it as an API decision. Whether a shared type is single-pass or re-iterable propagates to every consumer, so decide where materialisation lives -- in the producer, in a caching layer, or at the call site -- and how that choice is enforced in review and typing.
There are exactly two sensible bodies for `__iter__`, and picking between them is picking what your object **is**. ### Returning `self`: the object is an iterator An iterator is a cursor. It owns a position, `__next__` advances it, and when the source runs out it raises `StopIteration` -- permanently. `collections.abc.Iterator` codifies this: you supply `__next__`, and the ABC supplies an `__iter__` that returns `self`. The reason iterators define `__iter__` at all is uniformity -- `for` calls `iter()` on whatever you give it, so an iterator must answer that call with itself or it could never be used directly in a loop. This is why `for line in a file object:` works and why you can pass a generator straight into `sorted()`. Return `self` when the underlying source genuinely cannot be replayed: a socket, a pipe, a cursor over a result set, an API feed you page through, a `stdin` reader. Replaying would mean re-issuing the work, and pretending otherwise would be a lie about cost. ### Returning a fresh iterator: the object is a container A container holds or can regenerate its data, so every `__iter__()` call should hand back a new, independent cursor -- normally a new generator object, or `iter(self._items)`. This is what makes the object re-iterable, and re-iterability is not a nicety. Enormous amounts of Python quietly assume it. ### The four failure modes of returning `self` from a container **1. The silent second loop.** The first `for` consumes everything; the second sees an already-exhausted cursor and yields nothing at all. No exception is raised -- the loop body simply never runs, which is far harder to spot than a crash. **2. Nested loops collapse.** `for a in obj: for b in obj:` shares one cursor. The inner loop drains it on the first outer iteration, and the outer loop then ends too. The pairwise algorithm you thought you wrote produced one partial row. **3. Consuming probes.** `if not any(obj)`, `x in obj`, `len(list(obj))`, a log line printing `list(obj)[:3]` -- each of these calls `iter()` and advances the shared cursor. Adding a debug print changes the program's output. Membership testing is the sharpest case, because `in` falls back to iteration when `__contains__` is missing. **4. Functions that iterate twice.** Plenty of helpers take an iterable and walk it more than once -- compute a total, then normalise by it. Handed a self-returning object they see the full data on the first pass and nothing on the second, producing a wrong answer rather than an error. ### The tempting middle path, and why it is worse A common patch is to reset the cursor inside `__iter__` and still `return self`. Now repeated sequential loops work, so it passes a casual test -- but nested loops are still broken (the inner pass resets the outer one, and the loop may never terminate), and the object still claims to be an iterator while behaving like a container. If it is resettable, it is a container: build a new cursor instead. ### Making the choice visible Say it in the type, not just the docstring. A single-pass type should subclass `collections.abc.Iterator` (define `__next__`, inherit `__iter__`), and its name should read like a cursor -- `Reader`, `Stream`, `Cursor`, `Feed`. A re-iterable type should subclass `collections.abc.Iterable` or `Sequence`, and its name should read like a container. Callers can then check with `isinstance(obj, collections.abc.Iterator)`, which is the standard way library code asks "can I safely walk this twice?" ### What to do when a caller needs two passes over a one-shot object There are only three honest options and each costs something. Materialise it -- `rows = list(source)` -- which is correct and obvious, and buys speed with memory. Use `itertools.tee`, which gives you two independent iterators but internally buffers every element in the gap between the slower and faster consumer, so it saves nothing if one runs to the end first. Or re-create the source and pass over it again, paying the retrieval cost twice. Choosing among them belongs at the call site, where the size of the data is known -- not inside `__iter__`, where the class would be silently buffering on everyone's behalf. The interview-ready summary: `return self` means one-shot cursor, a fresh iterator means re-iterable container, and the bugs from getting it wrong are silent rather than loud.
- Why does `collections.abc.Iterator` define `__iter__` at all, if the object already has `__next__`?So an iterator can be used anywhere an iterable is expected. `for`, comprehensions, `sorted()`, `sum()` and unpacking all call `iter()` on their argument first; without an `__iter__` returning `self`, an iterator would fail that call and could never be looped directly. It is the adapter that keeps one protocol instead of two.
- A colleague resets the cursor at the top of `__iter__` and still returns `self`. What breaks?Nested iteration. Two simultaneous passes share one cursor, and the inner loop's reset rewinds the outer one -- so the pairwise loop either produces wrong results or never terminates. Concurrent consumers in threads have the same problem. If the object can rewind it is a container, so return a new iterator per call instead.
- A caller needs two passes over a single-pass object. What are the options and their costs?Materialise with `list()` -- correct, obvious, and pays memory proportional to the data. Use `itertools.tee` -- two independent iterators, but it buffers every element in the gap between consumers, so it helps only when they advance together. Or rebuild the source and iterate again, paying the retrieval cost twice. Decide at the call site, where the data size is known, rather than buffering inside the class.
Returning self is handing out your own bookmark: whoever reads next moves it for everyone. Returning a fresh iterator is handing each reader their own copy of the book.
saying these in an interview costs you the question
- Says __iter__ should always return self
- Thinks returning self restarts the loop each time
- Cannot explain why nested loops break on a shared cursor
- Resets state in __iter__ and still returns self
- Expects an exception rather than a silently empty second loop
- Believes itertools.tee replays without buffering