Why does Python's zip() stop at the shortest input, and how do you catch ragged data?
answer
- The shortest input decides the end
- Nothing is raised by default
- One keyword flag changes the ending
- It arrived in a 3.x release, not 2.x
- A different tool pads instead of stopping
basics
~20 szip() ends as soon as any input is exhausted, so unequal inputs are silently truncated. Pass strict=True, added in Python 3.10, to raise ValueError on a length mismatch, or use itertools.zip_longest with fillvalue to pad instead.
solid answer
~40 s`zip()` walks its inputs in lockstep and stops the moment any one of them raises `StopIteration`, so a shorter input silently truncates the result — no error, just missing rows. That default exists because `zip` must work with infinite and one-pass iterables, where comparing lengths up front is impossible. Since **Python 3.10** (PEP 618) the keyword-only `strict=True` turns a length mismatch into a `ValueError`, raised lazily during iteration at the point the mismatch is discovered, not at call time. When ragged input is legitimate rather than a bug, `itertools.zip_longest(*iterables, fillvalue=...)` runs to the longest input and pads the missing slots with `fillvalue`, which defaults to `None`. Choose deliberately: `strict=True` for data that must line up, `zip_longest` for data that need not.
code
python · 12 linesfrom itertools import zip_longest
employee_ids = ["e-101", "e-102", "e-103"]
hours = [38.5, 41.0]
print(list(zip(employee_ids, hours)))
print(list(zip_longest(employee_ids, hours, fillvalue=None)))
try:
list(zip(employee_ids, hours, strict=True))
except ValueError as exc:
print("mismatch:", exc)go deeper
Remember the one-line rule: the result ends when the first input ends, and nothing complains. Being able to predict that zipping a three-item list with a two-item list gives two tuples is enough here.
Explain why truncation is the default — inputs may be infinite or unsized — and name both remedies with their versions: strict=True from Python 3.10, and the padding function from itertools with its fillvalue.
Show that the strict check fires lazily at the end, so side effects performed per tuple have already happened. Talk about ordering the work, materialising pairs first, or wrapping the write so a late mismatch does not leave half-applied output.
Own it as a data-contract decision: which pipelines must assert equal length, which legitimately pad, and what the fill value asserts about the domain. Decide whether strict=True becomes a lint rule rather than a per-call habit.
## The stopping rule `zip(*iterables)` returns an iterator. Each step pulls one item from every input, left to right, and yields them as a tuple. The moment any input signals exhaustion, `zip` stops. It does not raise, does not pad, does not warn — the output simply ends. Zip over a three-element list and a two-element list produces two tuples. That is a deliberate design choice, not an oversight. `zip` accepts arbitrary iterables, including generators, open files and infinite sequences. It cannot ask them how long they are, because most iterables have no length and asking would mean consuming them. Stopping at the shortest is also what makes the pairing of a finite sequence with an unbounded counter useful at all. ## Why silent truncation is a real defect class The cost of that rule is that a data bug becomes invisible. Consider importing a payroll CSV: one pass builds a list of 83 employee identifiers, another builds the list of hours worked, and a `zip` marries them for the pay run. If a filter drops one hours row — a blank line, a header counted twice, a row rejected by validation — the identifiers list is 83 long and the hours list is 82. `zip` yields 82 pairs and the last employee vanishes from the run. Nothing raises. The output is well-formed, plausible, and wrong. Worse, if the two lists have drifted *out of alignment* rather than merely differing in length, every pair after the drift point is mismatched: the right shape, entirely wrong data. Truncation at least loses the tail; misalignment corrupts everything downstream of it. ## strict=True — Python 3.10 PEP 618 added the keyword-only `strict` parameter in **Python 3.10**: `zip(a, b, strict=True)`. It changes only the ending condition. While every input still has items, behaviour is identical. When one input is exhausted, `zip` checks whether the others are exhausted too; if any is not, it raises `ValueError` with a message naming which argument was longer or shorter. It is keyword-only, so it can never be mistaken for another iterable, and it defaults to `False` so existing code is unaffected. Two properties are worth naming in an interview. First, the check is **lazy**: nothing is validated at call time, because building the iterator consumes nothing. The `ValueError` surfaces only when the consumer reaches the end of the shortest input — so a `zip(..., strict=True)` that is never iterated, or is abandoned early by a `break`, never raises at all. Second, because the error arrives at the end, any work you performed per-tuple has already happened. If the loop body was writing pay rows, those rows are written by the time the mismatch is detected. Where partial output is unacceptable, either materialise the pairs first (`rows = list(zip(a, b, strict=True))`) and only then act on them, or wrap the whole write in a transaction you can roll back. That is exactly the judgement an interviewer is listening for. ## zip_longest — when ragged is legitimate Sometimes the inputs genuinely differ in length and dropping the tail is wrong. `itertools.zip_longest(*iterables, fillvalue=None)` runs until the **longest** input is exhausted, substituting `fillvalue` for every input that has already ended. With `fillvalue=0.0` the missing hours become zero rather than disappearing; with the default `None` you can test for the sentinel and branch. Pick the fill value with care: it lands in your data. A numeric default of `0.0` in a payroll import quietly asserts that the employee worked no hours, which may be a more expensive lie than dropping the row. `None` forces the consumer to make that decision explicitly, which is usually what you want. Note also that `strict` and `zip_longest` are mutually exclusive by nature — one demands equal lengths, the other assumes unequal ones — so the choice is a statement about your data contract. ## What not to do instead The reflex fix is to compare lengths before zipping: `if len(a) != len(b): raise`. That works for lists and fails for everything else, because `len()` does not exist on generators, file objects or database cursors, and materialising a stream just to measure it defeats the reason it was streamed. `strict=True` performs the same check without needing a length, which is precisely why it was added rather than left to callers. ## The decision, stated once Default to `strict=True` whenever the inputs are supposed to line up — parallel columns, keys and values, records and their computed results. Reach for `zip_longest` only when raggedness is part of the contract, and then choose the fill value on purpose. Leave bare `zip` for the case where truncation is the intent, such as pairing a finite sequence against an unbounded source.
- At what moment does zip(a, b, strict=True) raise, and why does that matter?Only during iteration, when the shorter input runs out and another input still has items. Constructing the zip object validates nothing. So a mismatch is never reported if you break out early or never consume the iterator, and if you do consume it, every tuple before the mismatch has already been processed — materialise the pairs first when partial side effects are unacceptable.
- Why not just compare len() on the two inputs before zipping?Because `len()` only exists on sized objects. Generators, open file objects, cursors and other one-pass streams have no length, and consuming one to measure it destroys it. `strict=True` performs the equivalent check as a by-product of the walk it is already doing, which is why the language added it instead of leaving it to callers.
- In a payroll import, the identifier stream and the hours stream come from an open reader. If zip stops early, what else can go wrong?The underlying reader is left mid-stream and, unless it was opened in a `with` block, its file handle is left unclosed — the loop ended but the resource did not. Always drive such readers inside a context manager so the handle closes on the normal exit, the early `break`, and the `ValueError` from strict alike.
- How do you choose the fillvalue for itertools.zip_longest?Treat it as data you are asserting, not as padding. `fillvalue=0.0` in a payroll import claims the employee worked zero hours; the default `None` claims nothing and forces the consumer to branch. Pick the neutral sentinel unless a domain-correct default genuinely exists.
saying these in an interview costs you the question
- Says zip raises an error on unequal lengths
- Thinks zip pads short inputs with None
- Believes strict=True checks lengths at call time
- Claims strict works on Python 3.9
- Compares len() on inputs that are generators
- Uses zip_longest with a numeric fill in financial data