What does zip(*rows) return in Python, and when does that transpose idiom break?
answer
- Two features composed in four tokens
- Each row becomes its own argument
- Rows in, columns out
- The star cannot unpack an endless source
- A short row costs you a column
basics
~20 sThe star unpacks each row into a separate argument, so zip walks the rows in lockstep and yields one tuple per column — a transpose. It breaks on ragged rows, which truncate silently, and on row sources too large or too lazy to unpack.
solid answer
~40 s`zip(*rows)` uses iterable unpacking: each row becomes its own positional argument, so `zip` receives N iterables and yields one tuple per column position — the transpose of a row-major table. Two things break it. First, unpacking is eager: every row must be materialised as an argument before `zip` is even called, so an unbounded or streaming row source cannot be transposed this way, and a very large table pays full memory for the argument list. Second, ragged rows truncate to the shortest, so a table with one short row silently loses columns; pass `strict=True`, available since **Python 3.10**, to make that a `ValueError`. The result is tuples, not lists, and `zip(*[])` is just `zip()`, an empty iterator.
code
python · 11 linesrows = [[1, 2, 3], [4, 5, 6]]
print(list(zip(*rows)))
print([list(col) for col in zip(*rows)])
ragged = [[1, 2, 3], [4, 5]]
print(list(zip(*ragged)))
try:
list(zip(*ragged, strict=True))
except ValueError as exc:
print("ragged table:", exc)go deeper
Recognise the shape and be able to say the result: rows go in, columns come out as tuples. Reading it correctly is enough here; writing it unprompted is not expected.
Decompose it into its two parts — iterable unpacking in the call, then ordinary lockstep zipping — and name both failure modes: eager unpacking of the row source, and silent truncation on ragged rows.
Show the cost model: the argument count is the row count, so tall tables are expensive and streamed ones cannot be transposed at all. Guard real tables with strict=True and know when the job belongs to an array library instead.
Decide where terseness is worth it. A dense four-token idiom in shared code may deserve a named helper, and a data path that transposes large numeric tables at all is usually a signal that the wrong layer is doing the work.
## What the star does The expression is two features composed. `*rows` is **iterable unpacking in a call**: it consumes `rows` and passes each of its elements as a separate positional argument. If `rows` is `[[1, 2, 3], [4, 5, 6]]`, then `zip(*rows)` is exactly `zip([1, 2, 3], [4, 5, 6])`. Once that substitution is made, the rest is ordinary `zip` behaviour: it pulls one item from each argument per round and yields a tuple. The first round takes the first element of every row — the first *column*. The second round takes the second element of every row. The output is therefore the table transposed, one tuple per column: ```python rows = [[1, 2, 3], [4, 5, 6]] list(zip(*rows)) # [(1, 4), (2, 5), (3, 6)] ``` Apply it twice and you get back the original shape, though as tuples of tuples rather than lists of lists. Convert explicitly if the type matters: `[list(col) for col in zip(*rows)]`. ## Failure one: unpacking is eager The star operator must build the argument list before the call happens. That has three consequences worth stating in an interview. It cannot work on an infinite row source: unpacking would never finish. It cannot work on a lazily streamed one either — a reader that yields rows from a file is fully drained into an argument tuple first, which defeats the streaming and, on a large table, holds every row in memory at once. And the number of arguments is the number of *rows*, so a tall, narrow table is passed as an enormous argument list, whereas a short, wide one is cheap. Transposing is therefore only free-looking; its cost is proportional to the row count, and it is a row-count-shaped cost, not a cell-count-shaped one. If the rows arrive lazily and the table is large, the honest answers are to materialise deliberately and accept the memory, to restructure the producer to emit columns directly, or to hand the job to a dedicated array library whose transpose is a view rather than a copy. Reaching for `zip(*rows)` on a stream is the mistake. ## Failure two: ragged rows truncate Because it is still `zip`, the result ends at the shortest row. A table where one row lost a field produces a transpose with fewer columns than the table has, and no error is raised — the tail columns simply do not exist in the output. Since **Python 3.10** the fix is `zip(*rows, strict=True)`, which raises `ValueError` when the rows are not all the same length. For a transpose that is nearly always the behaviour you want, because a ragged table is not a matrix and transposing it has no defined meaning. If the raggedness is legitimate, `itertools.zip_longest(*rows, fillvalue=None)` pads instead, producing a rectangular result with sentinels in the gaps. ## Edge cases to have ready `zip(*[])` unpacks nothing, so the call degenerates to `zip()` with no arguments — a valid call that yields an empty iterator. An empty table transposes to nothing rather than raising, which can silently swallow an upstream bug if you never check. A list of strings works too, because a string is an iterable of characters: `list(zip(*["abc", "xyz"]))` gives `[('a', 'x'), ('b', 'y'), ('c', 'z')]`. That is occasionally what you want and occasionally a sign you passed the wrong thing. The result is an iterator, so it is consumed once; wrap it in `list()` if you need to walk the columns twice. And the elements are tuples, which are immutable — if downstream code expects to mutate each column in place, convert. ## Reading it, and when to write something else The idiom is dense. `zip(*rows)` is four tokens carrying two separate ideas, and a reader who has not seen it before cannot decode it from the syntax. That is a fair argument for a named helper — `def transpose(rows): return list(zip(*rows))` — in code that is read by people who mostly write something else. Where the operation is genuinely central and the tables are numeric and large, the language-level idiom is the wrong layer anyway, and the work belongs to a library built for arrays. Used knowingly, though, it is exact: pair the columns of a rectangular table, in one expression, with a `strict=True` that turns a malformed table into a loud failure instead of a short one.
- Why can't you transpose a lazily streamed sequence of rows with zip(*rows)?Unpacking is eager: the star must consume the whole row source and build a positional argument list before `zip` is called. An unbounded source never finishes, and a large one is fully materialised in memory, which defeats the streaming. Either materialise deliberately, restructure the producer to emit columns, or use a library designed for arrays.
- What does zip(*rows) do when one row is shorter than the others?It truncates: the transpose ends at the shortest row, so the trailing columns silently disappear with no error. Pass `strict=True`, available since Python 3.10, to raise `ValueError` instead — for a transpose that is almost always right, since a ragged table has no meaningful transpose.
- What type are the columns you get back, and does that matter?Tuples, produced by a one-pass iterator. If downstream code mutates a column in place or iterates the result twice, convert explicitly with `[list(col) for col in zip(*rows)]` or wrap the whole thing in `list()`. Forgetting that the outer result is an iterator is the more common of the two mistakes.
saying these in an interview costs you the question
- Thinks the star transposes rather than unpacks
- Applies the idiom to an unbounded row generator
- Expects lists rather than tuples back
- Assumes ragged rows raise an error by default
- Believes it is memory-free for huge tables
- Expects the result to be re-iterable