When is itertools.chain.from_iterable a better flatten than a nested comprehension for a metrics scraper's per-endpoint sample lists?
answer
- One returns a list, one an iterator
- Peak memory, not raw speed
- It takes the container, not unpacked arguments
- Concatenate only; no transform or filter
- Bytes flatten into integers
basics
~20 sWhen you do not want the flat result materialized. itertools.chain.from_iterable consumes the outer iterable one sublist at a time and yields items lazily, so it works over a stream of scrape batches; a nested comprehension always builds the whole list.
solid answer
~40 s`itertools.chain.from_iterable(rows)` takes **one** iterable whose items are iterables and returns an iterator that yields their contents in order. Compared with `[x for row in rows for x in row]`, it (a) never builds the flat result — the caller decides whether to wrap it in `list`, feed it to `sum`, or consume it one item at a time; (b) works when `rows` is itself a generator that pulls batches from somewhere, because it holds only the sublist it is currently draining; and (c) is implemented in C, so per-item overhead is low when no transformation is needed. Use the comprehension the moment you want to transform or filter items, since `chain.from_iterable` only concatenates. Prefer it over `itertools.chain(*rows)`, which unpacks the whole outer iterable into positional arguments first, and remember it flattens exactly one level.
code
python · 11 linesimport itertools
def batches():
yield [1, 2]
yield []
yield [3]
flat = itertools.chain.from_iterable(batches())
assert not isinstance(flat, list)
assert list(flat) == [1, 2, 3]
assert sum(itertools.chain.from_iterable(batches())) == 6go deeper
Know that itertools.chain.from_iterable(rows) is the standard-library way to flatten one level of nesting, and that you must wrap it in list to get a list back rather than an iterator.
Explain the mechanics: it returns an iterator that drains one sublist at a time, it concatenates without transforming, and chain(*rows) is not a substitute because unpacking consumes the outer iterable up front.
Show production judgment: choose it when the outer object is a stream or the flat result feeds straight into a consumer, and be able to say what peak memory each form holds. Name the element-type trap where undecoded bytes flatten into integers.
Own the boundary rule — where a pipeline decides between materialized collections and streamed iterators, and where text is decoded — so flattening never quietly becomes the place the whole dataset is realized or the place an encoding mismatch first shows up.
## What the two forms actually are `[x for row in rows for x in row]` is a nested loop that appends every item to a new list, and it hands back that finished list. `itertools.chain.from_iterable(rows)` is different in kind: it returns an **iterator**. Nothing is flattened until something consumes it, and it holds only the sublist it is currently draining, not the concatenation. That difference is what decides between them, and it is easiest to see on a scraper that polls a set of endpoints and gets one list of samples back per endpoint: ```python import itertools def batches(): for endpoint in endpoints: yield collect(endpoint) # one list of samples per endpoint total = sum(1 for _ in itertools.chain.from_iterable(batches())) ``` Here `batches()` is a generator, so the comprehension form would still work, but it would hold every sample from every endpoint in one list at once. `chain.from_iterable` holds one endpoint's samples at a time. With an 83% cache-hit rate most of those per-endpoint lists are tiny or empty, so the difference is not the average batch — it is the total, which is precisely what the flat list would have had to hold. ## When to reach for each **Reach for `chain.from_iterable` when:** * the outer object is a stream, a file-backed generator, or anything you would rather not materialize; * the result is going straight into a consumer such as `sum`, `max`, `any`, `min` or a `for` loop, so the intermediate list would exist only to be thrown away; * you want concatenation and nothing else, and the C-level loop is worth the saved per-item bytecode. **Reach for the comprehension when:** * you want to transform items — `[f(x) for row in rows for x in row]` says that plainly, while the `chain` version needs a `map` wrapped around it; * you want to filter — a trailing `if` clause is clearer than composing another iterator; * the reader is better served by seeing the loops than by knowing one more `itertools` name; * you genuinely want the list, in which case the two are close enough in speed that clarity should decide. On 3.12 and later, list, set and dict comprehensions are inlined into the enclosing function rather than running in a separate frame, which narrowed the gap further; the remaining edge for `chain.from_iterable` on a plain concatenation is per-item interpreter overhead, not frame setup. ## The unpacking trap `itertools.chain(*rows)` looks equivalent and is not. The `*` unpacks the outer iterable into positional arguments **before** `chain` is called, so the outer iterable is fully consumed up front and every sublist is held simultaneously. On a generator of unknown length that erases the laziness you came for; on an unbounded one it never returns. `from_iterable` exists exactly to take the container instead of its unpacked contents. ## One level only, and the type trap Like the two-clause comprehension, `chain.from_iterable` removes exactly one level of nesting. Given an iterable of iterables of lists, you get lists back. The sharper trap is the element type, and it bites hardest where encodings are mixed. Both flatteners iterate whatever the sublists are. A `str` iterates as one-character strings, so a list of decoded metric names flattens into letters. A `bytes` object iterates as **integers**, so the same code over raw, undecoded lines flattens into a list of byte values. A scraper that decodes some responses and not others will therefore produce a mixture of `str` and `int` from identical-looking code, and the failure shows up far from the flatten — typically at a later join or comparison. Decide at the boundary whether you are handling `str` or `bytes`, and flatten only structures whose element type you have actually established. ## What a senior answer sounds like Do not present this as one being faster. State the categorical difference — a list versus an iterator, and the memory that follows from it — then give the rule: `chain.from_iterable` for pure concatenation, especially over a stream or straight into a consumer; a comprehension when the flatten also transforms or filters. Add the `chain(*rows)` warning, the one-level limit, and the `str` versus `bytes` element-type trap, and the question is fully answered.
- What is the difference between itertools.chain(*rows) and itertools.chain.from_iterable(rows)?The `*` unpacks `rows` into positional arguments before `chain` is even called, so the outer iterable is consumed in full and all sublists are held at once. That defeats laziness on a generator and never terminates on an unbounded one. `from_iterable` accepts the container itself and pulls one sublist at a time, which is why it exists as a separate constructor.
- How would you flatten and transform in one pass without building the intermediate list?Use a generator expression with the same two `for` clauses — swapping the brackets for parentheses gives the lazy form of the identical nesting — or wrap `map` around `itertools.chain.from_iterable`. Both keep the result unmaterialized. The generator expression usually reads better, because the transformation is written where a reader expects it rather than composed from two iterator calls.
- Does itertools.chain.from_iterable flatten more than one level?No. It yields the items of each sublist exactly as they are, so an iterable of iterables of lists flattens to lists. Removing another level needs another application, and unknown or mixed depth needs a recursive generator function with an explicit rule for treating `str` and `bytes` as leaves rather than as iterables of characters or integers.
saying these in an interview costs you the question
- Says chain.from_iterable flattens arbitrarily deep
- Uses chain(*rows) on an unbounded generator
- Thinks chain.from_iterable returns a list
- Claims the comprehension is always faster
- Expects flattening bytes objects to yield one-byte strings
- Picks it for speed while ignoring the memory difference