skip to content

Why can zip() silently drop data, and how does zip(strict=True) prevent it?

level: middleimportance: must knowfreq 60%

answer

  1. Stops with the shortest input
  2. No error is raised by default
  3. A keyword makes the mismatch loud
  4. Padding lives in itertools
  5. ValueError, raised lazily at the end

basics

~20 s

zip stops the moment its shortest input is exhausted, so trailing items of longer inputs are dropped with no warning. Since Python 3.10, passing strict=True makes zip raise ValueError when the inputs turn out to be unequal in length.

solid answer

~50 s

`zip` is lazy: each step pulls one item from every input and yields a tuple, and it stops as soon as any input raises `StopIteration`. Anything left in the longer inputs is discarded silently, which is exactly how two parallel lists that drifted out of sync produce a short, plausible-looking result instead of an error. Python 3.10 added the `strict` keyword (PEP 618): `zip(a, b, strict=True)` raises `ValueError` when one input runs out while another still has items, and the message says which argument was shorter. Because `zip` is lazy, that check fires *during* iteration, at the end — the tuples already yielded were already handed to the caller. When ragged input is legitimate rather than a bug, `itertools.zip_longest(a, b, fillvalue=...)` pads to the longest input instead. Comparing `len()` up front is not a substitute: many iterables have no length, and measuring one by consuming it destroys it.

code

python · 10 lines
python
from itertools import zip_longest

print(list(zip([1, 2, 3], "ab")))          # the 3 is dropped silently

try:
    list(zip([1, 2, 3], "ab", strict=True))
except ValueError as exc:
    print("mismatch:", exc)

print(list(zip_longest([1, 2, 3], "ab", fillvalue="-")))

go deeper

for a junior

Recall that zip walks inputs in lockstep and ends with the shortest one, producing tuples lazily. Know that the extra items simply vanish rather than causing an error.

for a middle

Explain the mechanics: one next() per input per step, stop on the first StopIteration, and strict=True since 3.10 raising ValueError. Contrast it with itertools.zip_longest and say why an up-front length check is not general.

for a senior

Show that the strict check fires mid-iteration, so side effects already committed are not undone, and argue for strict=True as a default habit in dict(zip(...)) and other parallel-collection code.

for a principal

Own the standard: decide where silent truncation is acceptable, make the strict form the reviewed default for cross-system data, and consider validating shape at the ingest boundary rather than at each use site.

## The default: truncate, quietly `zip(*iterables)` returns a lazy iterator. On each advance it calls `next()` on every stored iterator in left-to-right order and packs the results into a tuple. The instant one of those calls raises `StopIteration`, `zip` stops and raises `StopIteration` itself. Items already pulled from the earlier inputs on that final round are thrown away, and every item still sitting in the longer inputs is never read at all. No warning is emitted, and that is the whole hazard. `zip` is most often used on two collections the code *believes* are parallel — identifiers and their values, headers and a row, keys read from one system and quantities read from another. When one of them is short by an element, the result is not an exception but a shorter list of pairs that looks entirely reasonable. Downstream code then processes a subset and reports success. The bug surfaces days later as missing records rather than as a stack trace at the point of the mistake. A second, subtler consequence of the same mechanism: if the inputs are themselves one-pass iterators, the items consumed on the aborted final round are gone. Reading the remainder of a partially-zipped iterator afterwards can start one element later than you expect. ## strict=True Python 3.10 added `strict` (PEP 618), defaulting to `False`. With `strict=True`, when one input is exhausted `zip` checks whether the others are exhausted too; if any still yields a value, it raises `ValueError` naming which argument was shorter or longer than the first. The behaviour is symmetric: a longer *first* argument is reported just as clearly as a longer later one. The crucial property to say out loud is *when* it raises. `zip` cannot know the lengths in advance — its inputs may be generators — so the check happens lazily, at the moment one side runs dry. All earlier tuples were already produced and consumed. If the loop body writes rows to a database, sends messages, or mutates state, that work has already happened when the `ValueError` arrives. `strict=True` is therefore a loud failure signal, not a transactional guard: if the operation must be all-or-nothing, materialize and validate first, or make the loop body idempotent. ## zip_longest, and why len() is not the answer When inputs are genuinely allowed to differ, `itertools.zip_longest(*iterables, fillvalue=None)` continues until the *longest* is exhausted, substituting `fillvalue` for the inputs that ended early. It is the right tool for padding a short row, and the wrong tool for validating a contract, because it turns a mismatch into data rather than an error. The reflex to "just check `len(a) == len(b)` first" fails for the cases that matter most. Generators, open files, `map` objects and database cursors have no `len()`, and the only way to measure them is to consume them — which leaves nothing to zip. `strict=True` costs one extra `next()` call at the end of a single pass and works on anything iterable. ## Edge shapes worth knowing `zip()` with no arguments yields nothing. `zip(xs)` yields one-tuples, which surprises people who expect the original values back. Because `zip` is an iterator rather than a sequence, `len()` on the result is a `TypeError` and you cannot index it; wrap it in `list()` or `tuple()` when you need either — and remember that materializing it costs memory proportional to the shortest input. The symmetric operation, `dict(zip(keys, values))`, is the idiom that most often hides a truncation bug: a missing value simply produces a dictionary with fewer keys, and a later lookup fails far from the cause. That call is a strong candidate for `strict=True` as a matter of habit. ## Choosing The decision rule is short. If unequal lengths mean the data is wrong, pass `strict=True`. If unequal lengths are expected and padding is meaningful, use `zip_longest` with an explicit `fillvalue`. Only leave the bare default when truncation is the deliberate intent — pairing a stream against a finite set of slots, for example — and say so in a comment, because a reader cannot otherwise tell a deliberate truncation from an overlooked one.

  • Does zip(strict=True) raise before yielding anything, or partway through?
    Partway through. `zip` never learns the lengths in advance because its inputs may be generators, so it raises `ValueError` only when one input is exhausted while another still yields. Every tuple before that point has already been produced and consumed, so side effects in the loop body have already happened. Treat it as a loud alarm, not a transaction boundary.
  • Why not simply compare len() of both inputs before zipping?
    Because many iterables have no length — generators, open files, `map` objects, cursors — and measuring them means consuming them, leaving nothing to iterate. `strict=True` performs the check inside the single pass you were already making, works on any iterable, and costs one extra `next()` call at the end.
  • When is itertools.zip_longest the right choice instead of strict=True?
    When unequal lengths are legitimate and a filler value is meaningful — padding a short row to a fixed column count, or aligning an optional trailing field. `zip_longest` runs to the longest input and substitutes `fillvalue`. It is wrong for contract validation, because it converts a data error into plausible-looking data.

A clothing zip closes only as far as the shorter side reaches; the extra length on the other side just hangs there, unremarked, unless someone insists on checking that both sides ended together.

saying these in an interview costs you the question

  • Says zip raises an error on unequal lengths by default
  • Thinks strict=True pads the shorter input
  • Claims zip returns a list of tuples
  • Believes zip validates lengths before yielding anything
  • Says itertools.zip_longest raises on mismatched lengths
  • Assumes strict=True has existed in every Python 3 release

context