What does collections.defaultdict(list) do that a plain dict does not?
answer
- A dict that fills in the blanks
- No KeyError when the key is absent
- One callable, invoked only on a miss
- The subclass overrides __missing__
- default_factory: list, int, set
basics
~10 scollections.defaultdict(list) calls its factory when a looked-up key is missing, stores the fresh empty list under that key and returns it, instead of raising KeyError. Grouping becomes a single append with no membership check.
solid answer
~40 s`collections.defaultdict` is a `dict` subclass carrying one extra attribute, `default_factory`. When `d[key]` misses, `dict.__getitem__` hands off to `__missing__`, which `defaultdict` overrides: it calls the factory with **no arguments**, stores the result under that key and returns it. So `defaultdict(list)` turns grouping into `groups[k].append(v)` and `defaultdict(int)` turns tallying into `counts[k] += 1`, with no `if k not in d` preamble and no `KeyError`. The factory is any zero-argument callable — `list`, `int`, `set`, a class, or a function you write — and it runs only on a miss, never on a hit. Pass nothing and `default_factory` is `None`, at which point a missing key raises `KeyError` exactly like a plain `dict`. Everything else is inherited unchanged: insertion order, the view objects, and equality against a plain `dict`.
code
python · 10 linesfrom collections import defaultdict
rows = [("K", 4.1), ("Na", 139.0), ("K", 3.8)]
by_analyte = defaultdict(list)
for analyte, value in rows:
by_analyte[analyte].append(value) # no membership check needed
print(by_analyte["K"]) # [4.1, 3.8]
print(by_analyte.default_factory) # <class 'list'>go deeper
Be ready to write the grouping loop from memory and to name the three everyday factories: list for grouping, int for tallying, set for distinct values. Knowing that the factory takes no arguments is enough at this level.
Explain the mechanics rather than the recipe: dict.getitem delegates to missing, the factory is called with no arguments, and the result is stored before it is returned. Mention that a None factory restores plain KeyError behaviour.
Show the judgement about where it belongs. An accumulate-then-report mapping is a good fit; a lookup table where a missing key signals a defect is not, because you have traded a loud KeyError for a silent empty value.
Own the convention. Decide whether accumulators hand out defaultdicts or convert to plain dicts at module boundaries, so a mapping that invents entries never escapes into code that assumes reads are pure.
## The problem it removes Building a grouping or a tally on a plain `dict` always starts with the same three lines: look, discover the key is absent, create the container, then finally do the real work. Written out it is `if k not in d: d[k] = []` followed by `d[k].append(v)`. The logic is trivial but it is repeated at every accumulation site, it is easy to get subtly wrong, and it buries the one interesting line under bookkeeping. `collections.defaultdict` moves that bookkeeping into the mapping itself. ## The mechanism, precisely `defaultdict` is a genuine subclass of `dict` — `isinstance(d, dict)` is true, and every `dict` method is inherited. It adds exactly one piece of state, the instance attribute `default_factory`, and one piece of behaviour, an override of `__missing__`. `__missing__` is a hook defined by `dict` itself: `dict.__getitem__` calls `type(d).__missing__(d, key)` when a subscript lookup fails and the subclass defines it. Plain `dict` does not define it, so the lookup raises `KeyError`. `defaultdict.__missing__` does three things in order: 1. If `default_factory` is `None`, raise `KeyError(key)`. 2. Otherwise call `default_factory()` — **with no arguments at all**. 3. Store the returned value under `key`, then return that same object. Step 3 is the whole trick and also the whole trap. The value the caller receives is the value now living in the dict, so mutating it — `.append(v)`, `.add(v)` — mutates the stored value. It is also why a bare lookup grows the mapping, which is the single most-asked follow-up on this type. ## What counts as a factory Anything callable with zero arguments. The three canonical choices map onto the three canonical shapes: - `defaultdict(list)` — group values into per-key lists. - `defaultdict(int)` — tally, because `int()` is `0` and `+= 1` then works on the first sighting. - `defaultdict(set)` — collect distinct values per key. But a class, a `functools.partial`, or a plain function are all equally valid; a function is the right answer whenever the default is more than a bare empty container. What is *not* valid is passing a value rather than a callable: `defaultdict([])` raises `TypeError: first argument must be callable or None`. And even when a callable is accepted, it never learns which key it is producing a value for — the factory signature is fixed at zero parameters. A default that must depend on the key needs a mapping class of your own that overrides `__missing__` itself. ## Construction and conversion The constructor takes the factory first and then anything `dict()` accepts: `defaultdict(list, existing_mapping)` copies items in and attaches the factory. Going the other way, `dict(d)` produces an ordinary `dict` with the same items and no factory — worth doing before you hand the result to code that would be surprised by a mapping that invents entries, or before you serialize it. The factory is not fixed at construction time either. `d.default_factory = None` on a populated `defaultdict` freezes it: from that point a missing key raises `KeyError` again, while everything already accumulated is untouched. That is a cheap way to run an accumulation phase permissively and then make later typos loud. ## What does not change It is still a `dict`, so keys must be hashable, iteration follows insertion order, `keys()`/`values()`/`items()` return the usual views, and `d == {...}` compares only the items — `default_factory` takes no part in equality, so a `defaultdict` and a plain `dict` holding the same pairs compare equal. `repr()` does show the factory, which is why the printed form looks unfamiliar: `defaultdict(<class 'list'>, {'K': [4.1]})`. ## When it earns its place Use it when *every* missing key genuinely deserves the same freshly-built default and the mapping is used that way from creation to disposal. That is the accumulate-then-report shape: read a stream of records, fan them out into per-key buckets, report. Do not reach for it in a lookup table where a missing key means a bug — there, the `KeyError` from a plain `dict` is the feature you are throwing away. ## Why `int` works for counting `counts[k] += 1` is not one operation. Python expands augmented assignment on a subscript into a read, an addition, and a write: fetch `counts[k]`, add `1`, store the result back. The read is a normal subscript, so on a plain `dict` the very first sighting of a key raises `KeyError` and the write never happens. With `defaultdict(int)` the read misses, the factory returns `int()` — which is `0` — that zero is stored, the addition produces `1`, and the write replaces it. The same expansion explains why `defaultdict(list)` is used with `.append` rather than with `+=`: `append` mutates the stored list in place, needing no write-back at all.
- What happens if you build a collections.defaultdict without passing a factory?`default_factory` is `None`, and a missing key raises `KeyError` just as a plain `dict` would — the subclass adds nothing until a factory is attached. You can attach one afterwards by assigning `d.default_factory = list`, and you can remove it again by assigning `None`, which is the usual way to freeze a mapping once its accumulation phase is over.
- Can the default_factory see which key it is being asked to produce a value for?No. `defaultdict` calls it with zero arguments, so the factory has no idea which lookup triggered it. If the default genuinely depends on the key, `defaultdict` is the wrong tool — you need a mapping class of your own that overrides `__missing__` and uses the key it is handed.
- Does a collections.defaultdict compare equal to a plain dict holding the same items?Yes. Equality is inherited from `dict` and compares items only; `default_factory` plays no part, so `defaultdict(list, {"a": 1}) == {"a": 1}` is `True`. Only `repr()` differs, since it shows the factory. `dict(d)` gives you a genuine plain `dict` with the same items and no factory.
Like a filing cabinet that snaps a fresh empty folder into the gap the moment you reach for a label that has none, so you never have to check whether the folder exists before filing.
saying these in an interview costs you the question
- Says a defaultdict can never raise KeyError
- Passes a value, as in defaultdict([]), rather than a callable
- Thinks the factory receives the missing key as an argument
- Believes the factory runs on every lookup, not only on misses
- Claims defaultdict is a separate type unrelated to dict
- Uses defaultdict for a lookup table where a miss means a bug