skip to content

How does itertools.zip_longest differ from the built-in zip on unequal-length inputs?

level: juniorimportance: should knowfreq 48%

answer

  1. Shortest wins, silently
  2. Padding versus truncation versus error
  3. The pad value has a default
  4. One shared object fills every gap
  5. Strictness became an option in 3.10

basics

~10 s

The built-in zip stops at the shortest input and silently discards the rest. itertools.zip_longest runs until the longest input is exhausted and substitutes fillvalue, which defaults to None, for every missing element.

solid answer

~50 s

`zip(a, b)` yields tuples until the shortest input runs out; the surplus items in the longer one are dropped with no warning, which is a routine source of quietly truncated data. `itertools.zip_longest(a, b)` instead continues until the longest input is exhausted, substituting `fillvalue` (default `None`) for every missing element, so the number of tuples equals the longest length. Since Python 3.10 the builtin offers a third stance: `zip(a, b, strict=True)` raises `ValueError` when the inputs turn out to be different lengths, which is the right choice when equal lengths are an invariant rather than a hope. Pick deliberately — `strict=True` when a mismatch is a bug, `zip_longest` when a short input is legitimate and padding is meaningful, plain `zip` only when truncation is genuinely intended. Note that a single `fillvalue` object is shared by every padded slot.

code

python · 8 lines
python
import itertools

names = ["ada", "grace", "alan"]
scores = [91, 88]

print(list(zip(names, scores)))
print(list(itertools.zip_longest(names, scores)))
print(list(itertools.zip_longest(names, scores, fillvalue=0)))

go deeper

for a junior

Be ready to state the difference in one sentence: zip stops at the shortest input, zip_longest keeps going and fills the gaps with fillvalue, which is None unless you say otherwise.

for a middle

Explain the three behaviours and their failure modes — silent truncation, padded rows that downstream code must recognise, and the ValueError from strict=True added in 3.10.

for a senior

Show that you treat a length mismatch as a data-quality signal: decide per call site whether a short input is legitimate or a defect, and make padded slots detectable with a sentinel rather than a bare None.

for a principal

Own the convention: which of truncate, pad or raise is the house default for parallel iteration, and how invariants such as equal column lengths get enforced once at the boundary instead of re-checked in every function.

### The default is truncation, and it is silent The built-in `zip` walks its inputs in parallel and stops the moment any one of them is exhausted. `list(zip([1, 2, 3], ['a']))` is `[(1, 'a')]`. No exception, no warning, no log line — two of the three items simply never appear. That behaviour is right when truncation is the point (pairing an unbounded counter with a finite sequence, say), and wrong in the far more common case where the inputs were *supposed* to line up and a mismatch means something upstream is broken. ### zip_longest: pad instead of truncate `itertools.zip_longest(*iterables, fillvalue=None)` runs until the **longest** input is exhausted, substituting `fillvalue` wherever a shorter input has already run out. `list(itertools.zip_longest([1, 2, 3], ['a']))` gives `[(1, 'a'), (2, None), (3, None)]`. The number of tuples is the length of the longest input, and every tuple still has one element per input iterable, so downstream unpacking keeps working. Three details matter in practice. **The default fill is `None`.** That is fine until `None` is also a legitimate value in the data, at which point a padded slot and a real value become indistinguishable. The fix is a private sentinel: `MISSING = object()`, pass `fillvalue=MISSING`, and test with `is MISSING`. Since the sentinel cannot occur in real data, the test is exact. **One fill object is shared by every padded slot.** Passing a mutable object such as `[]` or `{}` as `fillvalue` puts the *same* object into every gap, so mutating one padded entry mutates all of them — the same aliasing trap as a mutable default argument. Fill with `None` or an immutable sentinel and construct a fresh object per row afterwards if you need one. **Padding hides the mismatch rather than reporting it.** zip_longest is the right answer when a short input is legitimate — an optional column, a series that starts later — and the wrong answer when a short input means a data defect, because it hands downstream code well-formed rows full of invented values. ### The third option: strict Python 3.10 (PEP 618) added the keyword-only `strict` parameter to the builtin: `zip(a, b, strict=True)` raises `ValueError` as soon as it discovers that one input is exhausted while another is not, naming which argument was shorter. On 3.14 all three behaviours are therefore available and each has a clear meaning: - `zip(a, b)` — stop at the shortest; truncation is intended. - `zip(a, b, strict=True)` — equal lengths are an invariant; a violation is a bug and should be loud. - `itertools.zip_longest(a, b, fillvalue=...)` — unequal lengths are expected and padding carries meaning. Most parallel iteration over data that *ought* to line up should be strict. It costs nothing, it fails at the point of the mismatch rather than three functions later, and the error message tells you which side was short. Because `strict=True` must confirm exhaustion before raising, the error surfaces when the iterators are consumed, not at call time — `zip` is lazy either way and returns an iterator, so nothing is checked until you actually iterate. ### Reading the output A subtle consequence of padding is that the *number of rows* now reflects the longest input, which is rarely what a length check downstream expects. Code that reports "processed N records" from `len(list(zip_longest(...)))` is counting real rows plus padded ones. Where that matters, count before combining, or filter the padded rows out explicitly with the sentinel test. ### Answering it well State the one-line difference first — shortest-wins versus longest-with-padding — then name `fillvalue` and its `None` default, then show that you know the third option exists and when each is appropriate. Mentioning the shared-fill-object trap and the sentinel technique for distinguishing padding from a genuine `None` turns a screening answer into a convincing one.

  • How do you tell a padded slot from a genuine None in zip_longest output?
    Use your own sentinel: `missing = object()`, then `itertools.zip_longest(a, b, fillvalue=missing)` and test with `is missing`. Because that object cannot appear in the input data, the test is unambiguous, whereas the default `None` collides with any legitimate `None` already present in the inputs.
  • Why is a mutable fillvalue such as an empty list dangerous?
    The same object is placed into every padded slot, so mutating one padded entry is visible through all of them — exactly the aliasing trap of a mutable default argument. Pass `None` or an immutable sentinel, then build a fresh object per row afterwards if the row really needs one.
  • When would you choose zip(..., strict=True) over itertools.zip_longest?
    When equal lengths are an invariant of the data rather than an expectation — parallel columns of one record set, for instance. `strict=True`, added in 3.10, turns a violated invariant into a `ValueError` at the point of the mismatch, while zip_longest would quietly hand downstream code padded rows that look entirely real.

saying these in an interview costs you the question

  • Thinks the built-in zip raises on unequal lengths
  • Believes zip pads short inputs automatically
  • Cannot name None as the default fillvalue
  • Passes a mutable object as fillvalue and shares it
  • Treats a padded None as a real input value
  • Counts padded rows as processed records

context