skip to content

In a Python comprehension, when does `if` filter items and when does it choose a value?

level: middleimportance: must knowfreq 62%

answer

  1. Two different ifs, two different jobs
  2. Position relative to the for clause
  3. One changes length, one changes values
  4. The filter has no else branch
  5. Ternary before for, filter after the iterable

basics

~20 s

An if written after the for clause is a filter, so failing items are skipped and the result gets shorter. An if/else written before the for clause is a conditional expression choosing each emitted value, so every item still appears.

solid answer

~50 s

Position decides. `[x for x in nums if x > 0]` puts the `if` after the `for` clause: it is a filter, it takes no `else`, and the result is shorter than the input. `[x if x > 0 else 0 for x in nums]` puts the conditional expression in the output slot, before the `for`: it picks a value per item, it *requires* the `else`, and the result is always the same length as the input. That is why `[x for x in nums if x > 0 else 0]` is a `SyntaxError` — the filter position has no else branch. You can combine both, as in `[x if x > 0 else 0 for x in nums if x is not None]`, and you can chain several trailing `if` clauses, which behave as an `and`.

code

python · 9 lines
python
nums = [3, -1, 4, -1, 5]

print([x for x in nums if x > 0])
print([x if x > 0 else 0 for x in nums])

try:
    compile("[x for x in nums if x > 0 else 0]", "<demo>", "eval")
except SyntaxError as exc:
    print("SyntaxError:", exc)

go deeper

for a junior

Memorise the two shapes and read them out loud: filter after the for clause, ternary before it. Know that a trailing if never takes an else.

for a middle

Explain the mechanics — the filter runs before the output expression and can guard it, and the ternary must produce a value so the else is mandatory. Be able to convert one form into the other on request.

for a senior

In review, spot when a filter has silently broken a positional pairing downstream, and insist on the ternary when the caller depends on one output per input. Watch for the loose ternary precedence in non-trivial output expressions.

for a principal

Set the readability rule for the codebase: how many clauses a comprehension may carry before it becomes a named function, and how filtering versus placeholder-mapping is documented so data-shape assumptions do not drift between teams.

Python comprehensions have two syntactically different places an `if` can appear, and they do unrelated jobs. Mixing them up is the single most common comprehension bug at interview and in review. ## The trailing filter ```python [x for x in nums if x > 0] ``` Everything after the iterable is the **filter clause**. It is evaluated once per item, after the loop name is bound and *before* the output expression runs. If it is falsey, the item is skipped entirely and nothing is emitted. The filter takes **no `else`** — there is nowhere for an else branch to go, because the alternative to emitting is emitting nothing. The result may therefore be shorter than the input, and in the extreme it may be empty. Because the filter runs before the output expression, it can guard that expression against bad input: ```python [1 / x for x in nums if x] # never divides by zero ``` You may also chain filters. `[x for x in nums if x > 0 if x % 2]` is legal and behaves as a logical `and`; most people write one `if` with `and` instead, but the chained form is what a multi-clause comprehension expands to. ## The leading conditional expression ```python [x if x > 0 else 0 for x in nums] ``` Here the `if` is not part of the comprehension grammar at all. `A if C else B` is Python's **conditional expression** — the ternary — and it is simply the output expression of the comprehension. Being an expression, it *must* produce a value in every case, so the `else` is mandatory. Every input item therefore yields exactly one output item, and the result is always the same length as the input. That length property is the real reason to choose it. If you are going to line the result back up against its source — pairing it positionally, writing it out row by row, comparing indexes — you need a placeholder for the items that did not qualify, not a hole. A filter would shift everything after the first skipped item. ## Why the wrong order is a SyntaxError ```python [x for x in nums if x > 0 else 0] # SyntaxError ``` The parser reaches the filter position, accepts `x > 0`, and then finds an `else` that the comprehension grammar has no slot for. There is no clever reading of this; it is simply not a form the language has. The mirror-image mistake, `[x if x > 0 for x in nums]`, is also a `SyntaxError`, because a conditional expression without an `else` is not a complete expression. A useful reading rule: **before the `for` is what you emit, after the last iterable is which items survive.** ## Combining the two Both can appear in one comprehension, and then order of evaluation matters: ```python [x if x > 0 else 0 for x in nums if x is not None] ``` The filter runs first and drops the `None` values; the conditional expression then maps each survivor to itself or to zero. Reading it aloud in that order — *for each x in nums, if it is not None, emit x when positive else zero* — is how you check such a line in review. ## Precedence, the quiet trap The conditional expression binds very loosely — more loosely than arithmetic. So `[x + 1 if x else 0 for x in nums]` parses as `(x + 1) if x else 0`, not `x + (1 if x else 0)`. When the output expression is anything more than a bare name, parenthesise the ternary explicitly. This is the kind of line that reads correctly to a human and computes something else. ## In dict and set comprehensions The same two positions exist. A dict comprehension may put a ternary in the key, the value, or both — `{k: (v if v is not None else 0) for k, v in pairs}` — and its trailing filter still goes at the end. A set comprehension behaves like the list form, except that a filter and a ternary can both be masked by deduplication: mapping every failing item to the same placeholder collapses them into a single element. ## What an interviewer is really testing Whether you know that filtering changes the **length** of the result and a ternary changes the **values**. The follow-up is usually a small refactor: *make this keep one output per input*, or *make this drop the bad rows*, and the answer is to move the `if` from one side of the `for` to the other.

  • Which form do you want when the result has to stay aligned with the input row by row?
    The conditional expression. It emits exactly one value per input item, so index `i` of the result still corresponds to index `i` of the source. A trailing filter shifts everything after the first skipped item, which silently breaks any later positional pairing or side-by-side write-out.
  • Can a comprehension carry more than one trailing `if`, and what does that mean?
    Yes. `[x for x in nums if x > 0 if x % 2]` is valid and the clauses combine as a logical `and`, evaluated left to right with short-circuiting. It is equivalent to a single `if` joined with `and`, which is usually the clearer way to write it.
  • Why does `[x + 1 if x else 0 for x in nums]` not add one conditionally?
    The conditional expression has lower precedence than `+`, so it parses as `(x + 1) if x else 0`. The whole `x + 1` is the true branch. To add a conditional amount, parenthesise the ternary: `[x + (1 if x else 0) for x in nums]`.

A filter is a bouncer at the door deciding who gets in; a conditional expression is a wardrobe inside deciding what each guest wears. The bouncer changes the headcount, the wardrobe never does.

saying these in an interview costs you the question

  • Writes a trailing `if` with an `else` attached
  • Thinks the ternary form also filters items out
  • Expects the filtered result to keep the input length
  • Cannot say which side of `for` each if goes on
  • Assumes the output expression binds tighter than the ternary

context