Why do len() and indexing fail on a Python generator object?
answer
- Which protocol does a generator implement
- len() looks for one dunder method
- Iterator, not a sized sequence
- Counting it would have to run it
basics
~20 sA generator object implements only iter and next, with no len and no getitem. It also cannot know how many items remain without running to the end, so the size is unknowable, not merely unimplemented.
solid answer
~40 s`len(obj)` calls `type(obj).__len__` and subscription calls `type(obj).__getitem__`; a generator object defines neither, so both raise `TypeError`. That is not an oversight. A generator produces values by executing code, and the only way to learn its length is to run it to exhaustion - which would defeat the laziness and consume the very items you wanted to count. It is an `Iterator`, not a `Sized` `Sequence`. In practice: if you truly need a count or random access, call `list()` on it and accept the memory; if you need one position, `itertools.islice` skips forward without materializing; if you need positions while streaming, use `enumerate`; if you need a count and nothing else, `sum(1 for _ in gen)` works, but it consumes the generator.
code
python · 9 linesimport itertools
gen = (n * n for n in range(10))
try:
len(gen)
except TypeError as exc:
print("no len:", exc)
print(next(itertools.islice(gen, 3, 4))) # the item at offset 3, without a listgo deeper
Remember the practical rule: a generator object supports iteration and nothing else, so len() and square-bracket access both raise TypeError. Convert with list() when you genuinely need a size or random access.
Explain the mechanics: len() dispatches to len and subscription to getitem, and a generator defines neither. Then explain why - the count exists only after running the generator, which would consume it.
Show the alternatives and their costs: enumerate for positions while streaming, itertools.islice to reach an offset, itertools.chain to push a peeked item back, and a deliberate list() when the data is bounded and needed twice.
Frame it as an interface decision: functions should declare whether they need an Iterable or a Sequence, and a TypeError at the boundary is better than a hidden materialization deep inside a pipeline that was designed to stream.
`len()` and `gen[0]` both fail on a generator object, and the reason goes deeper than a missing convenience method. ## What the built-ins actually call `len(obj)` is not a survey of the object - it is a call to `type(obj).__len__()`. Subscription `obj[i]` is a call to `type(obj).__getitem__(i)`. A generator object, the thing an `async`-free `def` containing `yield` returns when called, and the thing `(x for x in src)` evaluates to, defines exactly two protocol methods: `__iter__` (returning itself) and `__next__` (resuming the frame for one more value). No `__len__`, no `__getitem__`. So: - `len(gen)` raises `TypeError: object of type 'generator' has no len()` - `gen[0]` raises `TypeError: 'generator' object is not subscriptable` In the vocabulary of `collections.abc`, a generator object is an `Iterator`. It is not `Sized` and it is not a `Sequence`. Lists, tuples and strings are sequences: they own their elements, so length and position are cheap, constant-time facts about stored data. ## Why the size is unknowable, not just unimplemented Even if the language wanted to give generators a length, it could not compute one. A generator produces values by *running code*, and that code is arbitrary: ```python def lines_until_sentinel(source): for line in source: if line == "STOP": return yield line ``` How many items will that yield? It depends on data the interpreter has not read yet. The generator could be infinite (`itertools.count()`), could depend on a network read, could depend on a random draw. The only way to find out is to run it to the end - and running it consumes it, since a generator makes exactly one pass. A `len()` that destroyed its argument would be a far worse API than no `len()` at all. The same argument kills indexing. `gen[500]` could only be served by pulling and discarding 500 items, which is a linear, destructive operation wearing constant-time clothing. Python refuses to hide that cost behind sequence syntax. ## What to do instead **Need a count, do not need the items.** `sum(1 for _ in gen)` counts without holding anything, but leaves the generator exhausted - if you still need the data afterwards, this is the wrong tool. **Need a count and the items.** Materialize once: `items = list(gen)`, then `len(items)` and indexing both work. You have paid the memory, deliberately and visibly. **Need one position, not the whole thing.** `itertools.islice(gen, 500, 501)` advances past the first 500 items and yields the 501st, without building a list. It is still O(n) work - islice makes the cost honest rather than free. **Need positions while streaming.** `enumerate(gen)` yields `(index, item)` pairs and keeps the stream lazy, which covers most real uses of "I wanted an index". **Need to look ahead without losing the item.** Pull with `next(gen)`, inspect it, and use `itertools.chain([item], gen)` to put it back on the front of the stream. **Need a length in an API you control.** Do not accept a generator object there. Accept a `Sequence`, or take the iterator plus an explicit count, or convert at the boundary. Type annotations help: `Iterable[str]` promises only iteration, `Sequence[str]` promises length and indexing. ## The design lesson The absence of `__len__` is the type system telling you the truth about the object. A function that takes an iterator and calls `len()` on it is a function that has silently assumed materialized data. The failure is loud at the boundary - a `TypeError` on the first call - rather than quiet, which is what you want. This also explains why converting "just to check the size" is such a common performance bug: a pipeline written to stream a huge input in flat memory can be ruined by one `len(list(rows))` added during debugging and never removed. ## Version note None of this has changed across recent releases; a generator object's protocol surface is the same on 3.14 as it was a decade ago.
- How would you count a generator object's items and still use them afterwards?You cannot have both from one pass. Either materialize - `items = list(gen)`, then `len(items)` - and accept the memory, or restructure so the count falls out of the pass you were already making, for example incrementing a counter inside the consuming loop. Counting first and iterating second means two passes, and a generator object only has one.
- Why does len() work on range(1_000_000) but not on a generator over the same values?`range` is a lazy *sequence*, not an iterator: it stores start, stop and step, implements `__len__` and `__getitem__`, and computes any element by arithmetic in constant time. It is reusable and randomly accessible. A generator object stores a suspended frame and can only move forward, so neither length nor position can be answered without running it.
A list is a shelf of books you can count and reach into; a generator object is a conveyor belt - you can take the next item, but asking how many are still coming means running the belt to the end.
saying these in an interview costs you the question
- Says len() fails only because the generator was not primed
- Thinks len() would work after the generator is exhausted
- Believes indexing a generator would be a constant-time operation
- Calls list() on a huge stream just to get a count
- Confuses a generator object with a lazy sequence such as range
- Claims sum(1 for _ in gen) leaves the generator usable