Why does a collections.defaultdict grow when you only read a missing key?
answer
- Lookup here is not read-only
- Reading can change the length
- The hook stores before it returns
- get() and in never trigger it
- Freeze by clearing default_factory
basics
~20 sSubscripting a defaultdict is not a pure read. A miss triggers missing, which calls default_factory and stores the new value under the key before returning it, so the mapping grows. Use .get() or an in test when a lookup must not insert.
solid answer
~40 s`defaultdict.__missing__` does not merely *produce* a default, it **installs** it: on `d[key]` with `key` absent it calls `default_factory()`, assigns the result to `d[key]`, and returns that stored object — which is exactly why `groups[k].append(v)` works. The consequence is that any code path doing `if d[key]:`, `len(d[key])`, or a bare `d[key]` for inspection silently creates an entry. Only subscripting does this: `key in d`, `d.get(key)`, `d.keys()` and `d.setdefault(key, v)` never call `__missing__`. Symptoms are a `len()` that drifts upward as you inspect, empty lists appearing in serialized output, and `RuntimeError: dictionary changed size during iteration` when the stray lookup happens inside a `for` over the same mapping. Fixes: read with `.get()`, test with `in`, and set `d.default_factory = None` once the accumulation phase is over.
code
python · 9 linesfrom collections import defaultdict
by_analyte = defaultdict(list)
by_analyte["K"].append(4.1)
print(len(by_analyte)) # 1
by_analyte["Cl"] # looks like a read; it inserts
print(len(by_analyte)) # 2
print("Ca" in by_analyte, by_analyte.get("Ca"), len(by_analyte))go deeper
Remember the one-line rule: writing d[key] on a defaultdict can add the key, while d.get(key) and key in d never do. Reach for get when you are only checking what is there.
Explain why the insert is necessary rather than accidental: append works only because the returned object is the stored one. Then name the operations that bypass missing and the two ways to stop the growth.
Diagnose from the symptom. Drifting counts, empty groups in output, or a size-changed-during-iteration RuntimeError in single-threaded code all point at a probing lookup; show that you convert or freeze the mapping at the boundary as a matter of habit.
Frame it as an API-contract question. Decide whether a mapping whose reads mutate may cross a module boundary at all, and make conversion to a plain dict the house rule so callers never need to know which flavour they hold.
## Insertion is the point, not a side effect The reason `groups[k].append(v)` works at all is that the object `__missing__` returns is the object it has just stored. If it handed back a detached empty list, the `append` would mutate a throwaway and the mapping would stay empty forever. So insert-on-miss is load-bearing behaviour, not an oversight — but it means the subscript operator on a `defaultdict` is a *write*, and code review habits built around plain `dict` do not transfer. ## Exactly which operations insert Only `__getitem__` consults `__missing__`. That gives a short, memorable split: - **Inserts:** `d[key]`, and anything built on subscripting — `d[key].append(...)`, `if d[key]:`, `len(d[key])`, `for x in d[key]:`, an f-string interpolating `d[key]`. - **Does not insert:** `key in d`, `d.get(key)`, `d.get(key, fallback)`, `d.keys()` / `d.values()` / `d.items()`, `d.pop(key, fallback)`, and `d.setdefault(key, value)` — `setdefault` inserts the value *you* passed and ignores `default_factory` entirely. That last one surprises people: on a `defaultdict`, `setdefault` behaves exactly as it does on a plain `dict`. ## Where it bites in production Take a clinical-lab result loader ingesting a 6,800-row batch into `results = defaultdict(list)`, keyed by collection date parsed out of each row. The upstream file uses a locale-dependent date format, so a slice of the batch parses into a differently-formatted key string than the rest. Nothing raises. Downstream, a summary routine written as `if results[expected_date]:` probes a handful of dates that were never populated, and each probe *creates* them. Three failures follow, in escalating order of confusion: 1. **Counts drift.** `len(results)` is larger after the report ran than before it. A metric derived from `len(results)` reports more collection dates than the batch contained. 2. **Empty groups leak into output.** `defaultdict` is a `dict` subclass, so serializers accept it happily and write out the manufactured `[]` values as if they were real, empty result sets. A consumer cannot distinguish "no results that day" from "we never looked at that day". 3. **Iteration explodes.** The moment one of those probing lookups happens inside `for date in results:`, the next step of the loop raises `RuntimeError: dictionary changed size during iteration`. This is the good outcome — it is the only one that is loud — and it is routinely misdiagnosed as a threading problem when it is a single-threaded self-inflicted insert. ## How to keep the behaviour and lose the hazard **Read with the non-inserting API.** Anywhere the intent is inspection rather than accumulation, write `d.get(key, ())` or guard with `if key in d`. Reserve the subscript for the accumulation site itself. This is the fix that generalizes; the others are safety nets. **Freeze the factory when accumulation ends.** `results.default_factory = None` after the load loop turns every later miss back into a `KeyError`. Reporting code that probes an unknown date now fails loudly instead of inventing an entry, and the data already accumulated is untouched. **Convert at the boundary.** `dict(results)` returns an ordinary `dict` with the same items and no factory. Handing that out — to a serializer, to another module, to a caller — means no downstream code can trip the behaviour, and no reviewer of that code has to know about it. **Iterate over a snapshot when a lookup is unavoidable.** `for date in list(results):` iterates a copy, so an insert during the loop no longer invalidates the iteration. Prefer removing the insert, but this is the correct patch when the lookup genuinely belongs there. **Validate keys at the edge.** The lab loader's real defect is not `defaultdict`, it is a key format that was never normalized. Normalizing dates at ingest — parsing to a `datetime.date` and keying on that object rather than on whatever string the file carried — removes the mismatch and, with it, the probing that exposed the insert. ## Interview framing The question is asked because it separates people who have *used* `defaultdict` from people who have *debugged* one. The strong answer names the mechanism (`__missing__` stores before returning), draws the insert/no-insert line correctly, and volunteers the freeze-or-convert habit without being prompted. ## Why the design is still right It is tempting to call insert-on-read a wart, but the alternative is worse. A `__missing__` that returned a detached default would make `d[k].append(v)` silently do nothing — a far quieter bug than a mapping that grew. The type is honest about what it is: a mapping whose subscript is a write. What goes wrong in practice is not the semantics but the assumption, carried over from plain `dict`, that reads are free. Treat a `defaultdict` as an accumulator with a narrow lifetime, use the non-inserting API everywhere else, and the behaviour stays where it earns its keep.
- Which mapping operations on a defaultdict never call default_factory?Everything except subscripting. `key in d`, `d.get(key)`, `d.keys()`, `d.values()`, `d.items()` and `d.pop(key, fallback)` all leave the mapping alone. `d.setdefault(key, value)` also ignores `default_factory` completely — it inserts the value you passed, exactly as it would on a plain `dict`. Only `d[key]` reaches `__missing__`.
- How would you make a populated defaultdict raise KeyError again once loading is done?Assign `d.default_factory = None`. `__missing__` raises `KeyError(key)` whenever the factory is `None`, so later misses fail loudly while everything already accumulated stays put. The alternative is to convert with `dict(d)` and hand out the plain mapping, which also stops any downstream code from relying on the behaviour.
- A service logs RuntimeError: dictionary changed size during iteration and there is only one thread. What do you look for?A subscript on the same `defaultdict` inside the loop body. On a `defaultdict` a missing-key lookup inserts, which changes the size mid-iteration and invalidates the iterator. Find the lookup and switch it to `.get()` or an `in` test; if the lookup genuinely belongs there, iterate a snapshot with `for k in list(d):` instead.
saying these in an interview costs you the question
- Says a lookup never mutates a defaultdict
- Uses d[key] to test membership instead of in
- Thinks .get() also triggers default_factory
- Blames the RuntimeError on threading rather than the insert
- Believes setdefault consults default_factory
- Serializes the result without stripping manufactured empty values