skip to content

Why does [*records, *records] hold each item once when records is a generator object?

level: middleimportance: should knowfreq 30%

answer

  1. The star iterates, it does not copy
  2. Same source text, different operand kinds
  3. A generator object is its own iterator
  4. Exhausted once, empty forever after
  5. Materialise with list, or tee

basics

~20 s

A star unpacking iterates its operand to exhaustion. A generator object is a one-shot iterator, so the first unpacking drains it and the second finds nothing left. Materialise it as a list, which can be iterated again.

solid answer

~40 s

`*` does not copy a container; it obtains an iterator from the operand and consumes it. A list, tuple, set or `range` is re-iterable, so `[*data, *data]` doubles it. A generator object is its own iterator and yields each value once, so the first `*` drains it and the second contributes nothing — `[*gen, *gen]` is just `[*gen]`. The operands are evaluated in the order they appear, so it is genuinely the *first* unpacking that wins; and because unpacking happens at the call site, any side effect of consuming the iterable — reading a file, advancing a cursor — happens before the callee starts. Materialise once with `list(...)` if you need it twice, or use `itertools.tee`, which buffers whatever the slower branch has not read yet.

code

python · 5 lines
python
data = [1, 2, 3]
print([*data, *data])

gen = (n for n in data)
print([*gen, *gen])

go deeper

for a junior

Recall that a star unpacking walks its operand rather than copying it, and that a generator object can only be walked once, so the second unpacking in the same expression finds it empty.

for a middle

Explain the iterable-versus-iterator distinction that causes it, the left-to-right evaluation of the operands, and the two repairs: materialise once with list, or split the source with itertools.tee.

for a senior

Treat consuming an iterator at a call site as an eager side effect with a cost and a failure mode — the data is gone even if the call raises — and weigh a materialised list, a tee buffer and a lazy chain against the size of the stream.

for a principal

Own the interface rule: whether functions in your codebase accept one-shot iterators at all, or require a materialised sequence at public boundaries so that retries, logging and metrics cannot quietly consume the caller's only copy of the data.

## Unpacking is iteration, not copying The `*` in `[*x, *y]` or in `f(*x, *y)` is defined in terms of iteration: CPython obtains an iterator from the operand and drains it, appending each value. That definition is why `*` is so permissive about what it accepts — anything iterable works, including a string, a `range`, a set, a file object, a dict (which yields its keys) and a generator object. It is also why the operand's *iteration state* matters. ## Re-iterable versus one-shot - A list, tuple, set, dict or `range` is an *iterable*: asking it for an iterator gives a fresh one every time, so it can be walked repeatedly. - A generator object is an *iterator*: its `__iter__` returns itself, and once it has run to exhaustion it stays exhausted, yielding nothing forever after. The objects returned by `map`, `filter`, `zip`, `enumerate`, `reversed`, `open()` and the `itertools` functions behave the same way. So `[*data, *data]` on a list gives you the items twice, and the identical expression on a generator object gives them once — the operand looks the same in the source and behaves completely differently. ## The order this happens in Within a display or an argument list, operands are evaluated left to right, so the first `*` really does take everything and the second sees an empty iterator. In a call there is one wrinkle worth knowing: CPython assembles the positional group before the keyword group, so an iterable unpacking is evaluated before a keyword argument written earlier in the source. In `f(x=first(), *second())`, `second()` runs first. It is a small thing, but it is exactly the kind of detail that turns an ordering assumption about side effects into a bug — and it is why the grammar forbids `f(**d, *a)` outright rather than letting you write an order it will not honour. ## Consumption is a side effect, and it happens early Unpacking a generator that reads a network cursor or a file performs that reading at the call site, eagerly, before the function body executes. - If the function would have short-circuited and never touched the argument, unpacking has already paid the cost. - Worse, if the function raises, the iterator has still been consumed, and if it was your only handle on the data it is gone. That is the mechanism behind the classic "my batch is empty the second time round" bug: something computed `len(list(stream))` for a metric, and the real consumer got nothing. ## How to actually get it twice 1. The simplest fix is to materialise once: `data = list(stream)`, then use `data` wherever you need it, paying one traversal and holding the whole thing in memory. 2. When the data is too large for that, `itertools.tee` gives independent iterators over one source, at the cost of buffering whatever the fastest branch has read and the slowest has not; that buffer is unbounded if the branches diverge, so `tee` is not free either. 3. When you only need concatenation and not a second pass, `itertools.chain` streams the pieces lazily and avoids building the combined list at all — `[*a, *b]` is eager and allocates, `itertools.chain(a, b)` is lazy and does not. ## A related trap in the same family - Because `*` iterates, `f(*some_dict)` passes the dict's **keys** as positional arguments — not its values and not its items — while the same operand behind `**` passes the entries as keyword arguments. - And a one-shot iterator that was partially consumed before the unpacking contributes only the remainder: unpacking silently picks up wherever iteration had got to, which is why a debugging `next(it)` left in the code changes the result of the line below it. ## Why the distinction is worth interviewing on The iterable-versus-iterator split governs far more than unpacking — it is: - the same reason a generator cannot be re-used in two `for` statements, - the reason `len()` does not work on one, - and the reason a function that accepts "any iterable" and walks it twice is subtly broken for half its callers. Unpacking just makes the failure vivid, because both traversals sit on one line where they look symmetrical. ## How to answer it State the rule first — unpacking iterates the operand — then classify the operand as re-iterable or one-shot, and only then reach for the fix. Candidates who say "generators are lazy so it works twice" have the laziness right and the exhaustion wrong, and that is precisely the distinction the question exists to draw out.

  • In `f(*a, *b, **d1, **d2)`, in what order are the four operands evaluated?
    Left to right as written, so `a`, `b`, `d1`, `d2`. The one nuance: CPython builds the positional group before the keyword group, so an iterable unpacking is evaluated before a keyword argument that appears earlier in the source — in `f(x=first(), *second())`, `second()` runs first. The grammar removes the worst version of that surprise by rejecting `f(**d, *a)` as a SyntaxError.
  • What does `f(*some_dict)` pass, and how does it differ from `f(**some_dict)`?
    `*` iterates the dict, and iterating a dict yields its keys, so the keys arrive as positional arguments — the values are not passed at all. `**` spreads the entries as keyword arguments instead, and there the keys must be strings or you get `TypeError: keywords must be strings`. To pass pairs positionally you would unpack the dict's items view.
  • When would you prefer itertools.chain over [*a, *b]?
    When you only need to traverse the concatenation once and do not want the combined list in memory. `[*a, *b]` is eager: it allocates a list sized to both operands. `itertools.chain(a, b)` is lazy and yields from each in turn, which matters for large or streaming sources — but the result is a one-shot iterator itself, so it cannot be walked twice either.

A list is a printed page you can reread; a generator object is a ticker tape running past you once. Unpacking it twice is asking someone to read the same tape again after it has already run out.

saying these in an interview costs you the question

  • Thinks a generator object can be iterated twice
  • Says the star copies the iterator rather than draining it
  • Claims unpacking is lazy and happens inside the callee
  • Assumes a star unpacking after a keyword argument runs later
  • Believes *some_dict spreads the dict's values
  • Reaches for itertools.tee without noting its unbounded buffering

context