skip to content

enumerate, zip and reversed

The builtins that reshape a loop in place: enumerate hands you an index, zip walks in lockstep and truncates unless strict=True, and reversed needs a real sequence or a __reversed__ method.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

What does enumerate(items, start=1) give you that range(len(items)) does not?

level: juniorimportance: must knowfreq 75%

answer

  1. Pairs, not bare positions
  2. No len() required from the input
  3. Second argument offsets the counter
  4. start=1 numbers, it does not skip
  5. Lazy iterator of (count, value) tuples

basics

~20 s

enumerate yields (index, item) pairs straight from any iterable, so the loop body never indexes back into the sequence and the object needs no len(). The start argument only shifts the counter; it never skips items.

solid answer

~50 s

`enumerate(iterable, start=0)` returns a lazy iterator of `(count, value)` tuples, so `for i, item in enumerate(items):` binds both the position and the element in one step. `range(len(items))` gives you only numbers, forcing an `items[i]` lookup in the body and requiring an object that supports both `len()` and integer indexing — which rules out generators, open file objects and other one-pass iterators that `enumerate` handles fine. The `start` keyword sets the first counter value for human-facing numbering (`start=1` for line numbers or report rows) without arithmetic like `i + 1` scattered through the body. The classic misreading is that `start=1` begins at the second element: it does not, every item is still yielded, only the counter is offset. Reach for `range(len(...))` only when you genuinely need indices alone, for example to assign back into a list in place.

code

python · 7 lines
python
lines = ["open", "reconcile", "close"]
for number, step in enumerate(lines, start=1):
    print(f"{number}: {step}")

rows = [("sku-1", 4), ("sku-2", 9)]
for position, (sku, count) in enumerate(rows, start=1):
    print(position, sku, count)

go deeper

for a junior

Be ready to write the loop from memory and say what each part of the pair is. Know that start shifts the counter only, and that enumerate reads lazily rather than building a list.

for a middle

Explain why the idiom is more general than range(len(...)): no length and no subscripting are required, so generators and file objects work. Name the in-place-mutation case where the index is still needed.

for a senior

Show the off-by-one trap of using a start=1 counter as an index, and be able to spot it in review. Reach for enumerate when numbering a stream you must not buffer.

for a principal

Frame it as a readability and interface rule: functions should accept iterables rather than indexable sequences, which keeps callers free to pass streams. Positional counters that leak into data are a design smell worth naming in a review standard.

## What enumerate actually is `enumerate` is a built-in type, not a function that builds a list. Calling `enumerate(iterable, start=0)` constructs an `enumerate` object: a lazy iterator that holds a reference to an iterator over your input plus one integer counter. Each time it is advanced it pulls the next value from the underlying iterator and hands back a two-tuple `(counter, value)`, then bumps the counter. Nothing is precomputed and nothing is copied, so the memory cost is constant no matter how long the input is. That design is why the idiom is preferred over `range(len(items))`. The `range(len(...))` form asks two things of the object that `enumerate` does not: that it report a length, and that it support integer subscripting. A list, tuple or string satisfies both. A generator object, an open file, a `zip` object, a network cursor or anything else that can only be walked forwards satisfies neither. Writing `for i in range(len(f))` over a file is simply a `TypeError`; `for lineno, line in enumerate(f, start=1)` is the canonical way to number lines while streaming. ## The two arguments The first argument is any iterable — anything `iter()` accepts. The second, `start`, is the initial value of the counter and defaults to `0`. It can be passed positionally or by keyword, and `start=1` is by far the most common non-default: report rows, line numbers, ranked results and anything a human will read are conventionally one-based. The single most important thing to be able to state in an interview is what `start` does *not* do. It does not skip elements, it does not step the counter by more than one, and it does not change which values are produced. `list(enumerate(['a', 'b'], start=1))` is `[(1, 'a'), (2, 'b')]` — both elements, offset numbering. If you need to skip the header of a feed you slice or advance the iterator explicitly; if you need to skip *and* number the rest, you combine the two, and then the counter no longer lines up with positions in the original sequence. That mismatch is worth saying out loud: once `start` is anything but `0`, the counter is a label, not a valid index. Using it as an index is a silent off-by-one. ## Unpacking in the loop header `for i, item in enumerate(items):` works because each yielded value is a tuple and the `for` statement performs ordinary tuple unpacking. That composes: if the elements are themselves pairs, nested unpacking reads cleanly. ```python rows = [("sku-1", 4), ("sku-2", 9)] for position, (sku, count) in enumerate(rows, start=1): print(position, sku, count) ``` If you forget to unpack, the loop variable is the whole tuple, and the usual symptom is a confusing `TypeError` when you try to use it as a string or a number. ## When range(len(...)) is still right The idiom is not banned. Three cases keep it honest. First, when you want positions but not values at all — for instance building a list of indices that satisfy some property computed elsewhere. Second, when you must assign back into the sequence during the walk: `items[i] = transform(items[i])` needs the index, and `enumerate` gives you a *copy of the reference* to the element, not a writable slot, so rebinding the loop variable changes nothing in the list. Third, when you are stepping over a sequence with a stride or pairing an index against a second structure by position, where explicit arithmetic is clearer than a counter. Outside those, `enumerate` is the more general tool: it works on iterables with no length, it removes the repeated subscript from the body, and it keeps the counter and the value visibly in step. ## Cost and behaviour notes The object is an iterator, so it produces each pair on demand; you can wrap it around an unbounded source and stop early with `break`. It composes with anything that takes an iterable — `dict(enumerate(names))` builds a position-to-name mapping, and `sorted`, `min` and `max` accept the pairs directly, comparing counter first and value second, which is occasionally exactly what you want and occasionally a surprise.

  • Can you use enumerate on an open file object, and why does range(len(f)) fail there?
    Yes. A file object is an iterator over lines, and `enumerate(f, start=1)` numbers them while streaming, holding one line at a time. `len(f)` raises `TypeError` because the file has no length and `f[0]` is not supported: `range(len(...))` needs both a length and integer subscripting, which only real sequences provide.
  • If start=1 is passed, is the counter still safe to use as an index into the original list?
    No. The counter is offset by one, so `items[i]` reads the following element and the last iteration walks off the end or silently returns the wrong item. Once `start` is non-zero, treat the number as a human-facing label only; if you need both, keep the default and add one at the point of display.
  • When would you still reach for range(len(items)) instead?
    When the index itself is the payload — assigning back into the list in place (`items[i] = ...`), stepping with a stride, or pairing two sequences by position where explicit arithmetic reads better. Rebinding `enumerate`'s loop variable never writes back into the source, so in-place mutation genuinely needs the index.

saying these in an interview costs you the question

  • Claims start=1 makes the loop skip the first element
  • Says enumerate returns a list of tuples
  • Thinks enumerate requires the input to support len()
  • Uses the start=1 counter as an index into the list
  • Believes rebinding the loop variable edits the source list

context

open as a page

Why can zip() silently drop data, and how does zip(strict=True) prevent it?

level: middleimportance: must knowfreq 60%

basics

~20 s

zip stops the moment its shortest input is exhausted, so trailing items of longer inputs are dropped with no warning. Since Python 3.10, passing strict=True makes zip raise ValueError when the inputs turn out to be unequal in length.

open as a page

Why does reversed() raise TypeError on a generator object but work on a list?

level: middleimportance: should knowfreq 42%

basics

~20 s

reversed() must walk an object backwards, so it needs either a reversed method or the sequence protocol: len plus integer getitem. A generator object offers neither and can only move forward, so reversed() raises TypeError.

open as a page

Why is zip(*rows) a risky way to transpose a 2.4 GB inventory feed?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The * unpacking consumes the entire row source into one argument tuple before zip starts, so a lazy 2.4 GB feed is fully materialized. zip then pairs fields by position, and any ragged row silently truncates every column beyond the shortest.

open as a page