Why is a list comprehension usually faster than an append loop in CPython?
answer
- What the loop repeats per element
- An attribute lookup plus a call per element
- Hoisting the bound method proves the point
- Still bytecode, not a C loop
- 3.12 inlined it into the enclosing function
basics
~20 sA comprehension appends with a dedicated opcode onto the list it is building, skipping the attribute lookup and method call that out.append(x) pays every iteration. Since 3.12 it is also inlined into the enclosing function.
solid answer
~50 sAn append loop pays, per element, a lookup of the `append` attribute on the list, a call to that bound method, and the discarding of its `None` result. A list comprehension compiles to a dedicated append opcode operating directly on the list on the value stack: no attribute lookup, no call. That typically buys a modest factor for trivial bodies - and hoisting `append = out.append` before the loop recovers much of the gap, which is the proof that the lookup and call were the cost. Since CPython 3.12 (PEP 709) comprehensions are also inlined into the enclosing function rather than creating and calling an implicit function object, removing a frame push per comprehension. The win evaporates when the body calls a Python function, because that call then dominates - and readability, not speed, is the usual reason to choose either form.
code
python · 14 linesimport dis
def build_loop(xs):
out = []
for x in xs:
out.append(x * 2)
return out
def build_comp(xs):
return [x * 2 for x in xs]
dis.dis(build_loop)
print('-' * 40)
dis.dis(build_comp)go deeper
Know the two forms and that the comprehension is the idiomatic way to build a list from a sequence. Be able to write one with a condition and to say that it is usually the faster of the two.
Explain the mechanics: the per-element attribute lookup and method call that the append loop pays, the dedicated append instruction the comprehension uses, and the fact that the body itself is still interpreted bytecode in both.
Show measurement discipline - disassemble or time it rather than asserting - and know the boundaries: a body that calls a Python function erases the difference, map only wins with a C callable, and 3.12 inlining changed both the cost and what appears in a traceback.
Own the readability tradeoff. Argue when a team should standardise on comprehensions for intent, and when an explicit loop is the right call because the body needs statements, error handling or an early exit that a comprehension cannot express clearly.
### The two shapes, instruction by instruction ```python out = [] for x in xs: out.append(x * 2) ``` Every iteration of that loop does: advance the iterator, bind `x`, load the name `out`, look up the attribute `append` on it (which involves the type's lookup machinery and, without specialisation, materialising a bound method object), call it with one argument, and pop the returned `None`. The comprehension ```python [x * 2 for x in xs] ``` compiles instead to a loop whose body ends in a dedicated append instruction that acts on the partially built list already sitting on the interpreter's value stack. There is no name to load, no attribute to resolve, no call to make and no return value to discard. `dis.dis` on the two functions shows the difference directly, and it is the cleanest way to demonstrate the point in an interview. The practical effect for a trivial body is a modest speedup - a factor of well under two on modern CPython, not an order of magnitude. A useful sanity check is the classic hand optimisation: hoist the bound method out of the loop with `append = out.append` and call `append(x * 2)`. That recovers most of the difference, which tells you the cost was the repeated lookup plus the call, not anything mysterious about comprehensions. ### What is still interpreted A comprehension is not a C loop. Its body runs as ordinary bytecode, one instruction at a time, exactly like the explicit loop's body. What has moved into C is only the appending. This is the boundary that separates this idiom from something like `''.join(parts)` or `sum(data)`, where the entire iteration disappears into a single C call. Saying that comprehensions are fast because they are implemented in C is a misconception worth being able to correct. ### The 3.12 change Before CPython 3.12 a list, dict or set comprehension was compiled as a nested function: evaluating the comprehension created a function object, called it with the outer iterable, and pushed a frame. That is what gave comprehensions their own scope, so the loop variable never leaked. It also meant that for short comprehensions the frame push could be a large share of the total cost. PEP 709, shipped in 3.12, inlines those comprehensions into the enclosing function. No function object is created and no frame is pushed; the compiler instead saves and restores any enclosing local that shares a name with the iteration variable, so the isolation semantics are preserved exactly - the loop variable still does not leak. The PEP reports up to roughly a twofold improvement for comprehensions, largest on short ones where the frame overhead dominated. Two visible side effects: a comprehension no longer shows up as a separate frame in a traceback, and it no longer appears as a separate entry in profile output. Generator expressions were deliberately not inlined, since their laziness requires a real frame that outlives the expression. ### map and filter `list(map(f, xs))` moves the iteration itself into C, so it can beat a comprehension - but only when `f` is already a C-level callable, such as a built-in or an unbound method like `str.strip`. Wrap it in a lambda and each element now costs a Python function call plus the map machinery, and the comprehension wins comfortably. The rule that survives contact with real code: use `map` when you are passing a function that already exists and is implemented in C, and a comprehension otherwise. The same reasoning applies to `filter` against a comprehension with an `if` clause. ### When the difference stops mattering If the body calls a Python function, parses a string, or touches anything outside the interpreter, the per-element append overhead is a small fraction of the iteration and the two forms measure the same. That is the common case, and it is why this rewrite should be driven by clarity rather than by speed. A comprehension is the better default because it states intent - build a list from this sequence - in one expression; an explicit loop is better when the body needs several statements, error handling, or early exits. Nesting three comprehensions to avoid a readable loop is a net loss regardless of the microbenchmark. One thing both forms share: they build the whole result list in memory, so peak memory is proportional to the number of elements produced.
- If you hoist append = out.append before the loop, how much of the gap closes?Most of it for a trivial body. That is the diagnostic: the comprehension's advantage is the per-element attribute lookup and bound-method call it never makes, so removing the lookup from the loop recovers nearly all of it. The comprehension still edges ahead because it also avoids the call itself and the discarded return value, but the remaining difference is small.
- When is list(map(f, xs)) actually faster than the equivalent comprehension?When `f` is already a C-level callable - a built-in, or an unbound method such as `str.strip` - because then both the iteration and the per-element call live in C. If you have to wrap the operation in a lambda, each element costs a Python call plus the map machinery and the comprehension wins. So: `map` for a function that already exists in C form, a comprehension for an inline expression.
- Now that comprehensions are inlined in 3.12, does the loop variable leak into the enclosing scope?No. PEP 709 preserved the semantics: the compiler saves any enclosing local that shares the iteration variable's name and restores it afterwards, so the name is unchanged outside the comprehension. What did change is observability - the comprehension no longer appears as its own frame in tracebacks or in profiler output, because there is no longer an implicit function being called.
saying these in an interview costs you the question
- Says comprehensions are faster because they run as C loops
- Claims the speedup holds regardless of what the body does
- Thinks wrapping a lambda in map beats a comprehension
- Believes the comprehension's loop variable leaks after 3.12 inlining
- Nests comprehensions for imagined speed at the cost of readability