skip to content

Why is zip(*rows) a risky way to transpose a 2.4 GB inventory feed?

level: seniorimportance: should knowfreq 30%

answer

  1. Two mechanisms, not one
  2. The star is evaluated before the call
  3. A generator argument buys nothing here
  4. Fields pair by index, not by name
  5. One short row costs every trailing column

basics

~20 s

The * unpacking consumes the entire row source into one argument tuple before zip starts, so a lazy 2.4 GB feed is fully materialized. zip then pairs fields by position, and any ragged row silently truncates every column beyond the shortest.

solid answer

~50 s

`zip(*rows)` looks like a cheap transpose but has two production hazards. First, memory: the `*` unpacking is evaluated at call time, so every row is drained out of the source into a single argument tuple, and `zip` then advances all of them in lockstep — a streaming 2.4 GB feed becomes 2.4 GB resident, and passing a generator does not help. Second, correctness: `zip` pairs by position, not by field name. If one system emits its columns in a different order, the transpose produces well-formed columns holding mismatched values and raises nothing. And because `zip` truncates to its shortest input, a single ragged row silently drops every trailing column. `strict=True` (Python 3.10+) converts the ragged case into a `ValueError`, but nothing detects reordered columns — for that, pair by key with dictionaries and validate the header at ingest.

code

python · 8 lines
python
rows = [(1, "widget", 12), (2, "bolt", 7), (3, "nut")]

print(list(zip(*rows)))            # the quantity column vanishes

try:
    list(zip(*rows, strict=True))
except ValueError as exc:
    print("ragged feed:", exc)

go deeper

for a junior

Recognise zip(*rows) as the transpose idiom and be able to predict its output on a small rectangular list of rows. Know that the result is an iterator of tuples, not a list of lists.

for a middle

Explain the two stacked mechanisms: iterable unpacking evaluated at call time, then zip truncating to the shortest row. Show that a short row costs a whole column and that strict=True turns that into a ValueError.

for a senior

Diagnose the memory blow-up in a streaming pipeline and say why passing a generator does not help, then argue for name-based pairing over positional pairing when data crosses a system boundary.

for a principal

Decide where a pivot belongs at all — in the producing system, in a chunked pipeline, or not at all — and set the ingest contract: schemas validated at the boundary, positional pairing confined to code that owns both sides.

## What the idiom actually does `zip(*rows)` is two mechanisms stacked. The `*` performs iterable unpacking in a call: Python iterates `rows` to exhaustion and builds a tuple of arguments before `zip` is ever invoked. `zip` then treats each row as one input and yields tuples that take one element from every row in turn — the first output tuple is the first field of every row, the second is the second field of every row, and so on. On a rectangular list of rows, that is a transpose. The elegance is real and so is the cost, and reading the two halves separately is what makes both visible. ## Hazard one: nothing about this is lazy The unpacking is eager by definition — a call's arguments must exist before the call. So the row source is fully consumed at the moment of the call, whatever it is: a generator streaming a 2.4 GB inventory feed off disk, a database cursor, a decompressing reader. All of it lands in a tuple of rows that stays alive for as long as the `zip` object does, because `zip` holds an iterator into every single row and needs them all simultaneously to advance in lockstep. That is the counter-intuitive part worth stating clearly: passing a generator to `zip(*...)` buys nothing. `zip` is lazy in its *output* — the columns are produced on demand — but the input was already flattened into memory. A pipeline carefully built to stream row by row collapses to peak-resident at that one line, and it does so silently, showing up as an out-of-memory kill or a machine that starts swapping rather than as an error anyone can attribute to the transpose. The structural reason is unavoidable rather than an implementation wart: the first output tuple needs the first field of the *last* row, so no transpose can emit anything before reading everything. If the data does not fit, the operation must change shape — accumulate per-column into separate sinks as rows stream past, chunk the feed and transpose block by block, or push the pivot into whatever system produced the rows. ## Hazard two: position is not identity `zip` pairs by index. In an inventory sync between two systems, that is exactly the assumption most likely to be violated over time. Both sides emit `(sku, location, quantity)` today; one side adds a column, or reorders its select list, or its export template changes, and the rows still have three fields — so the transpose still succeeds. Column two now contains a mixture of locations and something else, and nothing raises. The failure is not a crash, it is quietly wrong data flowing into whatever consumes the columns, and by the time a reconciliation report disagrees the original rows may be gone. No argument to `zip` detects this, because from `zip`'s point of view nothing is wrong. The defence is structural: carry field names with the data — dictionaries, or a named-record type — so pairing is by key, and validate the header against an expected schema at the ingest boundary before anything positional happens. Positional pairing is safe only where both sides are produced by the same code in the same place. ## Hazard three: ragged rows truncate `zip` stops with its shortest input. In the transposed direction, the "inputs" are the rows, so one short row truncates the number of *columns* produced, dropping the tail of every other row. A feed where one record is missing an optional trailing field loses that field for the entire batch. This one has a direct fix: `zip(*rows, strict=True)`, available since Python 3.10, raises `ValueError` when the rows are not all the same length. It fires during iteration rather than up front, so it is an alarm rather than a guard, but it converts a silent data loss into a stack trace at the transpose — which is what you want in a sync job. Making the strict form the reviewed default for any positional pairing of externally-sourced data is cheap and catches the ragged case permanently. ## What to say in the interview The strong answer separates the two failure modes, because they have different fixes. Memory: caused by argument unpacking, not by `zip`; fixed by not transposing a stream at all. Correctness: caused by positional pairing; fixed by pairing on names and validating the schema, with `strict=True` closing the narrower ragged-row case. Naming `strict=True` alone is a partial answer — it makes the shape error loud but does nothing at all about a column-order change, which is the one that corrupts data instead of stopping the job.

  • How would you transpose a feed that does not fit in memory?
    You cannot do it lazily: the first output column needs a field from the last row. Practical options are to accumulate each column into its own sink as rows stream past, to chunk the feed and transpose block by block with a reduce step, or to push the pivot upstream to whatever produced the rows. Full materialization is a deliberate choice with a stated memory budget, not a default.
  • Does strict=True protect against the two systems emitting columns in a different order?
    No. `strict=True` only checks that the rows have equal length; reordered columns of the same width look perfectly valid to `zip`. Protection has to be structural — carry field names with the records and pair by key, and validate the incoming header against an expected schema at the ingest boundary before any positional operation happens.
  • Is zip(*rows) applied twice guaranteed to return the original rows?
    Only when the rows are rectangular and you accept tuples in place of the original row type. Ragged input loses columns on the first pass, so the round trip returns a rectangular, truncated version of the data. Any per-row type — a list, a custom record — comes back as a plain tuple.

saying these in an interview costs you the question

  • Believes zip(*rows) streams rows lazily
  • Says a generator argument keeps memory constant
  • Assumes ragged rows raise an error
  • Thinks zip pairs fields by column name
  • Claims transposing twice always restores the input
  • Treats strict=True as protection against reordered columns

context