skip to content

When Python's zip() stops early over iterators, what happens to items it already pulled?

level: seniorimportance: should knowfreq 28%

answer

  1. It pulls left to right each round
  2. The round that fails is abandoned
  3. Lists hide it, streams do not
  4. One value never reaches a tuple
  5. The same trick powers pairwise chunking

basics

~20 s

They are discarded. zip pulls from its inputs left to right each round, so when a later input is exhausted mid-round, the items already taken from earlier inputs are dropped and cannot be recovered — visible whenever the inputs are iterators rather than re-iterable sequences.

solid answer

~40 s

`zip` builds each output tuple by pulling one item from every input in argument order. If a later input raises `StopIteration` part-way through a round, `zip` stops immediately and the items it already pulled from the earlier inputs in that round are thrown away. With lists this is invisible, because a list can be iterated again from the start. With **iterators** — generators, open file objects, stream readers — the consumption is permanent: the item is gone from the source, so resuming that iterator after the zip skips a record. It is the same reason the `zip(*[iterator] * n)` chunking idiom works and the same reason it silently drops a partial final chunk. `strict=True`, added in 3.10, reports the mismatch but does not give the swallowed item back.

code

pycon · 6 lines
pycon
>>> a = iter(range(4))
>>> b = iter(range(2))
>>> list(zip(a, b))
[(0, 0), (1, 1)]
>>> next(a)
3

go deeper

for a junior

Know the safe rule: over lists you can ignore this entirely, because a list can be walked again. The subtlety only appears once one of the inputs is a generator or an open file.

for a middle

Explain the round-by-round pull in argument order and predict, on paper, which values are consumed but never yielded. Be able to say why zipping one iterator with itself pairs neighbours.

for a senior

Diagnose it in the wild: an off-by-one in a reconciliation pass caused by reusing a reader after a zip, and a reader left open mid-stream. Show the fixes — context managers, zip_longest, or materialising before pairing.

for a principal

Frame it as a stream-contract question: which pipeline stages may consume shared iterators at all, whether the tail is allowed to be dropped, and where a lossy pairing must be replaced by an explicit batching step with defined tail handling.

## The mechanic `zip` produces its output one round at a time. Within a round it calls `next()` on each input in the order the arguments were written. If every call succeeds, it packs the results into a tuple and yields it. If any call raises `StopIteration`, `zip` stops the whole iteration immediately — and the values it already obtained from the inputs to the left in that unfinished round are simply dropped. They were never packed into a tuple, so nothing downstream ever sees them. Over lists this is a non-event. A list is an *iterable*, and zipping it creates a fresh iterator over it; the list itself is untouched and can be walked again from index zero. Over an *iterator* it is a real, observable side effect, because the iterator holds the position and `zip` has advanced it. ```python a = iter(range(4)) b = iter(range(2)) list(zip(a, b)) # [(0, 0), (1, 1)] next(a) # 3 — the value 2 was pulled and discarded ``` The first two rounds succeed. In the third, `zip` pulls `2` from `a`, then finds `b` exhausted, stops, and drops the `2`. Resuming `a` yields `3`. One record has vanished with no error. ## Where it bites in production The classic case is pairing two live streams. Imagine a payroll import where one reader yields employee identifiers and another yields the hours rows, both lazily from open files, and a loop zips them to build pay records. When the hours reader ends first, the identifier reader has already been advanced past the record it was about to pair. If the code then continues to use that identifier reader — for a reconciliation pass, a tail-handling branch, or a second zip — it starts one record late and never notices. The bug reads as an off-by-one in output that has no visible off-by-one in the code. The same mechanic explains a second surprise: partially consumed readers left mid-stream. A `zip` that stops early leaves both readers open and positioned somewhere in the middle. If those readers were not opened inside a `with` block, the loop ends but the resource does not — the file handle is left unclosed until the object is collected. Drive readers with a context manager so the handle closes on the normal exit, on an early `break`, and on the `ValueError` from `strict=True` alike. ## The idiom that depends on it The discard rule is not purely a hazard; one well-known idiom is built on it. Zipping the *same* iterator with itself groups consecutive items: ```python it = iter([1, 2, 3, 4, 5, 6]) list(zip(it, it)) # [(1, 2), (3, 4), (5, 6)] ``` Because both arguments name one shared iterator, the first `next()` in a round takes `1` and the second takes `2`. Generalise it with `zip(*[iter(seq)] * n)` for fixed-size chunks. Its documented weakness is exactly the discard rule: if the input length is not a multiple of `n`, the final partial chunk is pulled and thrown away rather than emitted. That is acceptable for fixed-width framing and quietly wrong for record batching, where you must not lose the tail. ## What strict=True does and does not fix `strict=True`, added in **Python 3.10**, changes the ending condition: on discovering that one input ended while another still had items, `zip` raises `ValueError` instead of stopping quietly. That converts a silent data loss into a loud failure, which is the right default for parallel data that must line up. What it does not do is restore the swallowed item. By the time the mismatch is detected, `next()` has already been called on the earlier inputs in that round, and there is no way to push a value back into a plain iterator. So `strict=True` tells you the inputs were ragged; it does not let you resume either input at the record boundary you wanted. If you need to keep the tail, `itertools.zip_longest` is the tool — it never abandons a round, padding the exhausted inputs with `fillvalue` instead. ## How to reason about it in review Three questions settle almost every case. Are the inputs iterators or re-iterable sequences? If sequences, the discard is unobservable and you can ignore it. Is any input used *again* after the zip? If yes, assume it has been advanced one item further than the last tuple you saw. Is losing the tail acceptable? If not, you want `zip_longest`, or you want to materialise the inputs before pairing them. And in every one of those cases, if the input is a reader over a real resource, it belongs in a `with` block so that stopping early still closes it.

  • Why does zip(it, it) with one shared iterator pair consecutive items?
    Both arguments name the same iterator, so within a round the first `next()` takes one item and the second takes the following one. Generalised as `zip(*[iter(seq)] * n)` it yields fixed-size chunks. Its caveat is the discard rule: when the length is not a multiple of n, the final partial chunk is consumed and never emitted.
  • Does strict=True let you recover the item that zip already consumed?
    No. It only changes the ending condition, raising `ValueError` when one input ends while another has items left. The `next()` calls for that round have already happened and a plain iterator has no push-back. Use `itertools.zip_longest` if the tail must survive, or materialise the inputs before pairing them.
  • How does this interact with an input opened as a file or other resource?
    Stopping early leaves the reader open and positioned mid-stream; without a `with` block the handle stays unclosed until collection. Wrap the reader in a context manager so it closes on the normal end, on an early break, and on a strict mismatch — and never reuse a reader after a zip without accounting for the extra advance.

saying these in an interview costs you the question

  • Says zip never advances an input past the last tuple
  • Assumes an iterator can be re-zipped from the start
  • Thinks strict=True gives the consumed item back
  • Cannot explain why zip(it, it) pairs neighbours
  • Expects the partial final chunk to be emitted
  • Leaves a stream reader unclosed after an early stop

context