skip to content

Comprehensions and Generator Expressions

Python's declarative way to build a list, dict or set — or a lazy stream — in one expression. Comprehensions appear in almost every Python interview, usually with a follow-up about memory or scoping.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

16

In the comprehension [x for row in rows for x in row], which for clause is the outer loop?

level: juniorimportance: must knowfreq 70%

answer

  1. Rewrite it as nested statements
  2. Clauses read left to right
  3. Leftmost clause is the outer loop
  4. Output expression runs innermost
  5. One level of flattening only

basics

~20 s

The leftmost one. Clauses read left to right in the same order you would write nested for statements, so for row in rows is the outer loop and for x in row the inner one. The result is one flat list.

solid answer

~50 s

Multiple `for` clauses in a comprehension are a direct transliteration of nested `for` statements written outermost first, so the clauses are read **left to right**: `for row in rows` is the outer loop, `for x in row` is the inner loop, and the output expression `x` runs in the body of the innermost loop. That is why this exact form is the idiomatic one-expression flatten of a list of lists. It flattens **exactly one level** — each additional level of nesting needs another `for` clause, and truly arbitrary depth needs a recursive generator function instead. Writing the clauses in the opposite order fails with a `NameError`, because a clause can only use names bound by clauses to its left. Two clauses is about the readability limit; past that, a plain nested loop or a helper function reads better.

code

pycon · 10 lines
pycon
>>> rows = [[1, 2], [3], [4, 5, 6]]
>>> [x for row in rows for x in row]
[1, 2, 3, 4, 5, 6]
>>> flat = []
>>> for row in rows:
...     for x in row:
...         flat.append(x)
...
>>> flat
[1, 2, 3, 4, 5, 6]

go deeper

for a junior

Be ready to read [x for row in rows for x in row] aloud as two nested loops and say what it returns for a small ragged input. Knowing that the leftmost clause is the outer loop is the whole ask here.

for a middle

Explain the mechanics: each clause becomes a loop nested in the previous one, the output expression runs in the innermost body, and a clause may only use names bound to its left. Show the NameError from reversing the clauses.

for a senior

Demonstrate judgment about when to stop: two clauses is the readability ceiling, unknown depth needs a recursive generator, and strings and bytes flatten into characters and integers respectively. Name the quadratic flattening anti-patterns rather than merely avoiding them.

for a principal

Own the convention: decide where the team draws the line between a flattening expression and a named helper, and how flattening interacts with data whose shape is not guaranteed, so ragged or unexpectedly deep input fails loudly instead of silently producing letters.

## The rule: left to right, outermost first A comprehension with more than one `for` clause is defined as a transliteration of nested statements. Read the clauses left to right and write each one as a loop nested inside the previous one; the output expression, which sits at the far *left* of the comprehension, is what executes in the body of the *innermost* loop. So this comprehension: ```python flat = [x for row in rows for x in row] ``` is exactly this statement form: ```python flat = [] for row in rows: # first clause -> outer loop for x in row: # second clause -> inner loop flat.append(x) # output expression -> innermost body ``` The only thing that moves is the output expression, which is written first but evaluated last. Everything else keeps its order. Once you internalise that single rewrite, every multi-clause comprehension becomes readable mechanically rather than by intuition. ## Why this is the flattening idiom Because the outer loop walks the sublists and the inner loop walks their items, appending each item to one result, the expression turns a list of lists into a single flat list. It is the standard one-expression flatten, and it has some useful properties: * **Ragged input is fine.** The sublists need not be the same length, and an empty sublist simply contributes nothing. * **The sublists need not be lists.** Any iterable works — tuples, sets, ranges, generators — because the inner clause only requires something iterable. * **The outer object need not be a list** either; it just has to be iterable. ## It flattens exactly one level This is the misconception that shows up most often. `[x for row in rows for x in row]` removes **one** level of nesting. Given `[[1, 2], [3]]` you get `[1, 2, 3]`; given `[[[1, 2]], [[3]]]` you get `[[1, 2], [3]]` — still nested. Each extra level costs another `for` clause, and by three clauses the expression is usually harder to read than the loop it replaces. For a structure of unknown or mixed depth, a comprehension is the wrong tool: write a small recursive generator function that yields leaves and recurses into anything iterable, with an explicit stopping rule for strings so they are not exploded into characters. That string case is worth its own warning. A `str` is iterable, and iterating it yields one-character strings, so flattening a list of words gives you a list of letters. A `bytes` object is iterable too, and iterating it yields **integers**, not one-byte objects — so the same code over decoded text and over raw bytes produces two very different results. ## The reversed-order bug A clause may only reference names bound by clauses to its left, because those clauses are the enclosing loops. Writing `[x for x in row for row in rows]` raises `NameError: name 'row' is not defined`: the first clause tries to iterate `row`, which nothing has bound yet. The error is loud and immediate, which is a small mercy — but the fix is not to shuffle the clauses until it runs, it is to write the nested statements out and translate them back. ## Alternatives and what they cost * `list(itertools.chain.from_iterable(rows))` produces the same flat list, consuming the sublists one at a time. * An explicit loop calling `list.extend` per sublist is equally clear and often the most readable option when there is any surrounding logic. * `sum(rows, [])` also flattens, and should be avoided: it builds a new intermediate list on every step, so it is quadratic in the total number of items. `functools.reduce` with list concatenation has the same defect. ## Same syntax, other brackets The identical clause ordering applies to the dict, set and generator forms — `{x for row in rows for x in row}` deduplicates while flattening, and swapping the brackets for parentheses gives the generator form of the same nesting. The nesting rule does not change with the bracket; only what is built from the innermost body does. ## Interview framing Interviewers ask this because it separates people who have memorised the flatten one-liner from people who can derive it. The strong answer is the rewrite: state that the clauses map onto nested `for` statements in written order, produce the four-line loop version, then note the one-level limit and the readability ceiling at two clauses.

  • What happens if you write the two for clauses in the opposite order?
    `[x for x in row for row in rows]` raises `NameError: name 'row' is not defined`. Each clause is an enclosing loop for the clauses to its right, so a clause can only use names bound to its left. Here the first clause tries to iterate `row` before anything has bound it. The failure is immediate rather than a wrong answer, which makes it easy to spot.
  • How would you flatten a structure whose nesting depth is not known in advance?
    Not with a comprehension — each `for` clause removes exactly one fixed level. Write a small recursive generator function that iterates the object, yields items that are not iterable, and recurses into the ones that are, with an explicit guard so a `str` or `bytes` is treated as a leaf rather than exploded into characters or integers.
  • Why is sum(rows, []) a poor way to flatten a list of lists?
    It concatenates left to right, allocating and copying a new list at every step, so the total work is quadratic in the number of items rather than linear. `[x for row in rows for x in row]` or `list(itertools.chain.from_iterable(rows))` each build the result once. On small inputs the difference is invisible, which is exactly why it survives into code that later grows.

The clauses are an indentation ladder written on one line: the leftmost clause is the least-indented loop, and the output expression is the deepest line in the body.

saying these in an interview costs you the question

  • Says the rightmost for clause is the outer loop
  • Claims one comprehension flattens to arbitrary depth
  • Reads multi-clause comprehensions right to left
  • Thinks flattening a list of words yields whole words
  • Recommends sum(rows, []) as a linear flatten
  • Cannot rewrite the comprehension as nested for statements

context

open as a page

Why does a Python list comprehension's loop variable not leak into the enclosing scope?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Since Python 3, a comprehension evaluates its body in its own implicit function scope, so the name bound by its for clause is local to the comprehension and gone afterwards. Python 2 list comprehensions leaked it; generator expressions never did.

open as a page

In Python, how does `(x*x for x in nums)` differ from `[x*x for x in nums]`?

level: juniorimportance: must knowfreq 75%

basics

~10 s

The square-bracket form builds the whole list in memory immediately. The parenthesised form builds a generator object that computes each value only when asked, uses near-constant memory, and can be iterated only once.

open as a page

What are the list, dict and set comprehension forms in Python, and how do they differ?

level: juniorimportance: must knowfreq 85%

basics

~20 s

Square brackets build a list, braces with a colon per item build a dict, and braces without a colon build a set. All three share one shape: an output expression, a for clause, and an optional trailing if filter.

open as a page

In a Python comprehension, when does `if` filter items and when does it choose a value?

level: middleimportance: must knowfreq 62%

basics

~20 s

An if written after the for clause is a filter, so failing items are skipped and the result gets shorter. An if/else written before the for clause is a conditional expression choosing each emitted value, so every item still appears.

open as a page

In the comprehension [y for x in xs for y in f(x)], how many times is f called?

level: middleimportance: should knowfreq 35%

basics

~20 s

Once per item in xs. The second for clause is the inner loop, so its iterable expression is re-evaluated on every pass of the outer loop, which is exactly what lets it depend on the outer name.

open as a page

In [x for row in rows if row for x in row], what does the if clause filter?

level: middleimportance: should knowfreq 45%

basics

~20 s

The rows, not the items. An if clause attaches to the for clause immediately to its left and runs at that loop level, so this one skips empty rows before the inner loop ever starts.

open as a page

Where does a := assignment inside a Python comprehension bind its target name?

level: middleimportance: should knowfreq 40%

basics

~20 s

In the scope containing the comprehension, not the comprehension's own implicit scope. PEP 572 chose that deliberately so a comprehension can export a computed value; afterwards the name holds the value from the last item that reached the assignment.

open as a page

Why does a Python generator expression produce nothing the second time you iterate it?

level: middleimportance: should knowfreq 60%

basics

~20 s

A generator expression is a one-shot iterator. The first full pass drives it to completion, after which it is exhausted and every further loop ends immediately with no items and no error. Rebuild it, or materialize it once with list().

open as a page

When is a list comprehension the wrong choice in Python, and what do you write instead?

level: middleimportance: should knowfreq 48%

basics

~20 s

Use a plain for loop when you want a side effect rather than a container, when you need statements a comprehension cannot hold such as try/except or break, or when the expression has grown past what a reader can take in at once.

open as a page

When is itertools.chain.from_iterable a better flatten than a nested comprehension for a metrics scraper's per-endpoint sample lists?

level: seniorimportance: should knowfreq 40%

basics

~20 s

When you do not want the flat result materialized. itertools.chain.from_iterable consumes the outer iterable one sublist at a time and yields items lazily, so it works over a stream of scrape batches; a nested comprehension always builds the whole list.

open as a page

Why does a generator expression still yield the old list after the source name is rebound in a document-conversion queue?

level: seniorimportance: should knowfreq 32%

basics

~20 s

The outermost iterable is evaluated eagerly, in the enclosing scope, when the generator is created, and the resulting object is handed to the generator. Rebinding the name afterwards moves the name only; the generator still holds the original list.

open as a page

Why does `any(check(r) for r in rows)` stop early while `any([check(r) for r in rows])` does not?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Arguments are evaluated before the call runs. The bracketed form calls check on every row to build the list first, then any scans it. The bare generator expression is consumed only until the first true value, so the remaining calls never happen.

open as a page

Why can a dict comprehension over `zip(ids, scores)` silently lose rows from a 6,800-row batch?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Two silent losses stack: zip stops at the shorter input, dropping the tail, and a repeated key overwrites its earlier entry. Worse, zip pairs purely by position, so if the two lists were built by different passes the mapping is wrong rather than short.

open as a page

Why does a comprehension in a Python class body raise NameError for a class attribute?

level: middleimportance: nice to knowfreq 18%

basics

~20 s

The comprehension body runs in its own implicit function scope, and a class body is not an enclosing scope a nested function can read. Class-level names are therefore invisible inside it, except in the outermost iterable.

open as a page