What are the list, dict and set comprehension forms in Python, and how do they differ?
answer
- One syntax, three containers
- Punctuation picks the container
- The colon is what makes it a dict
- Braces without a colon are a set
- Empty braces are a dict, not a set
basics
~20 sSquare brackets build a list, braces with a colon per item build a dict, and braces without a colon build a set. All three share one shape: an output expression, a for clause, and an optional trailing if filter.
solid answer
~40 sPython has three container comprehensions and the punctuation picks the container. `[len(w) for w in words]` builds a list, `{w: len(w) for w in words}` builds a dict because each item is a `key: value` pair, and `{w[0] for w in words}` builds a set. Each one reads the same way: the output expression comes first, then `for name in iterable`, then an optional `if` that filters items out. All three are eager, so the whole container is built before the expression finishes. The dict form drops earlier duplicates because a repeated key overwrites, and the set form deduplicates and gives you no ordering guarantee. `{}` is an empty dict, not an empty set; the empty set is written `set()`.
code
python · 9 lineswords = ["alpha", "beta", "gamma", "beta"]
lengths = [len(w) for w in words]
by_word = {w: len(w) for w in words}
initials = {w[0] for w in words}
print(lengths)
print(by_word)
print(sorted(initials))go deeper
Be able to write all three forms from memory and read them back in words. Know which bracket produces which container, and that set() is the only way to spell an empty set.
Explain the shared anatomy — output expression, for clause, optional trailing if — and the lossy edges: duplicate keys overwrite in a dict comprehension, and a set comprehension deduplicates and drops ordering.
Show judgment about which container the code actually needs: a set when you only test membership, a dict keyed by something provably unique, a list when position matters. Notice silently shrinking results in review.
Own the convention for the codebase: where comprehensions are the house style, where a named helper is required instead, and how you keep a lossy dict build from becoming an unreviewed data-correctness bug.
A comprehension is a single expression that builds an entire container by iterating over something else. Python has three container comprehensions, and the punctuation around the body decides which container you get. ## The three forms ```python words = ["alpha", "beta", "gamma", "beta"] [len(w) for w in words] # [5, 4, 5, 4] -> list {w: len(w) for w in words} # {'alpha': 5, ...} -> dict {w[0] for w in words} # {'a', 'b', 'g'} -> set ``` Every one of them has the same anatomy: 1. an **output expression** — what to emit for each item (for a dict, two expressions separated by a colon: `key: value`); 2. a **`for` clause** that binds a name over an iterable; 3. an optional **trailing `if`** that filters items out before the output expression runs. `[x * x for x in nums if x % 2]` reads as *square each odd number in `nums`*. The literal loop equivalent is `result = []` then `for x in nums:` / `if x % 2:` / `result.append(x * x)`. The comprehension puts the intent — *I am building a list, of squares* — in the first two tokens rather than making the reader reconstruct it from three lines of mutation. ## Braces are shared, and the colon disambiguates Sets and dicts both use braces, so Python decides by looking for a colon in the item. `{w: len(w) for w in words}` is a dict; `{len(w) for w in words}` is a set. This is also why `{}` is an empty **dict**: the literal was a dict long before sets got a literal syntax, and the empty set has to be spelled `set()`. A comprehension never produces an empty-brace ambiguity, because it always has a body to inspect. ## What each container does to your data The **list** form is the faithful one: one output item per surviving input item, in iteration order, duplicates kept. The **dict** form is lossy in a way that surprises people. Keys must be hashable, and a repeated key silently overwrites — `{w: len(w) for w in words}` on the list above yields three entries, not four, because `beta` appears twice. If the key expression is not unique per item, you are writing a last-one-wins reduction, not a mapping. Insertion order is preserved: since Python 3.7 that is a language guarantee, not an implementation detail. The **set** form deduplicates by hash and equality, and the result has no meaningful order — printing one is not a stable sight. Elements must be hashable, so `{[w] for w in words}` raises `TypeError`. A set comprehension is the natural way to answer *which distinct values appear*, e.g. `{w[0] for w in words}` for the distinct initials. ## Iterating a dict inside a comprehension Iterating a dict yields its keys, so `[k for k in d]` is the keys and `{v: k for k, v in d.items()}` is the classic inversion — it needs `dict.items()` to get pairs, and it collapses any duplicate values into one key. ## They are expressions, and they are eager A comprehension can go anywhere an expression can: an argument, a return value, a dict value. It cannot contain statements — no `try`, no `break`, no assignment statement — which is the boundary at which you go back to a `for` loop. All three container comprehensions are **eager**: they run to completion and hand back a fully built object. That is the difference from the parenthesized form, which is a different construct with its own behaviour. ## Speed, and what changed Comprehensions are usually a little faster than the equivalent loop, because the interpreter uses a dedicated append opcode instead of looking up and calling the `list.append` bound method on every iteration — you can see this by disassembling both with `dis.dis`. In **Python 3.12** (PEP 709) list, dict and set comprehensions were further changed to be *inlined* into the enclosing function rather than creating and calling a hidden one-shot function object each time, which roughly halves the per-comprehension overhead. The visible semantics did not change. ## When to reach for which Use a list comprehension when you want a mapped or filtered sequence you will index or iterate more than once. Use a dict comprehension to build a lookup table, invert a mapping, or reshape pairs. Use a set comprehension when you want distinct derived values, or a fast membership set. If what you actually want is a side effect rather than a container, none of the three is the right tool.
- Why does `{}` create an empty dict rather than an empty set, and how do you write an empty set?The brace literal meant a dict long before sets had a literal syntax, so `{}` stayed a dict for compatibility and the empty set has no literal at all — you write `set()`. Non-empty braces are unambiguous because Python looks for a colon: `{1: 2}` is a dict, `{1, 2}` is a set.
- What happens when the key expression in a dict comprehension is not unique across items?Nothing raises. Later items overwrite earlier ones, so the dict ends up shorter than the input and holds the last value seen for each key. If that is not what you intended, compare `len(result)` against the input length, or group with a list value instead of overwriting.
- Does a set comprehension preserve the order of the source iterable?No. A set is unordered and deduplicated, so the result reflects hashing, not iteration order, and you must not rely on how it prints. A dict comprehension does preserve insertion order — that has been a language guarantee since Python 3.7.
One assembly line with three different bins at the end: the same conveyor of items, and the bracket you write decides whether they land in an ordered crate, a labelled filing cabinet, or a bag that quietly throws away repeats.
saying these in an interview costs you the question
- Says `{}` is an empty set
- Thinks braces always mean a dict
- Expects a dict comprehension to keep duplicate keys
- Believes a set comprehension preserves source order
- Claims comprehensions are lazy like the parenthesized form
- Cannot rewrite a simple append loop as a comprehension