skip to content

Why does a collections.defaultdict grow when you only read a missing key?

level: middleimportance: must knowfreq 60%

answer

  1. Lookup here is not read-only
  2. Reading can change the length
  3. The hook stores before it returns
  4. get() and in never trigger it
  5. Freeze by clearing default_factory

basics

~20 s

Subscripting a defaultdict is not a pure read. A miss triggers missing, which calls default_factory and stores the new value under the key before returning it, so the mapping grows. Use .get() or an in test when a lookup must not insert.

solid answer

~40 s

`defaultdict.__missing__` does not merely *produce* a default, it **installs** it: on `d[key]` with `key` absent it calls `default_factory()`, assigns the result to `d[key]`, and returns that stored object — which is exactly why `groups[k].append(v)` works. The consequence is that any code path doing `if d[key]:`, `len(d[key])`, or a bare `d[key]` for inspection silently creates an entry. Only subscripting does this: `key in d`, `d.get(key)`, `d.keys()` and `d.setdefault(key, v)` never call `__missing__`. Symptoms are a `len()` that drifts upward as you inspect, empty lists appearing in serialized output, and `RuntimeError: dictionary changed size during iteration` when the stray lookup happens inside a `for` over the same mapping. Fixes: read with `.get()`, test with `in`, and set `d.default_factory = None` once the accumulation phase is over.

code

python · 9 lines
python
from collections import defaultdict

by_analyte = defaultdict(list)
by_analyte["K"].append(4.1)

print(len(by_analyte))                     # 1
by_analyte["Cl"]                           # looks like a read; it inserts
print(len(by_analyte))                     # 2
print("Ca" in by_analyte, by_analyte.get("Ca"), len(by_analyte))

go deeper

for a junior

Remember the one-line rule: writing d[key] on a defaultdict can add the key, while d.get(key) and key in d never do. Reach for get when you are only checking what is there.

for a middle

Explain why the insert is necessary rather than accidental: append works only because the returned object is the stored one. Then name the operations that bypass missing and the two ways to stop the growth.

for a senior

Diagnose from the symptom. Drifting counts, empty groups in output, or a size-changed-during-iteration RuntimeError in single-threaded code all point at a probing lookup; show that you convert or freeze the mapping at the boundary as a matter of habit.

for a principal

Frame it as an API-contract question. Decide whether a mapping whose reads mutate may cross a module boundary at all, and make conversion to a plain dict the house rule so callers never need to know which flavour they hold.

## Insertion is the point, not a side effect The reason `groups[k].append(v)` works at all is that the object `__missing__` returns is the object it has just stored. If it handed back a detached empty list, the `append` would mutate a throwaway and the mapping would stay empty forever. So insert-on-miss is load-bearing behaviour, not an oversight — but it means the subscript operator on a `defaultdict` is a *write*, and code review habits built around plain `dict` do not transfer. ## Exactly which operations insert Only `__getitem__` consults `__missing__`. That gives a short, memorable split: - **Inserts:** `d[key]`, and anything built on subscripting — `d[key].append(...)`, `if d[key]:`, `len(d[key])`, `for x in d[key]:`, an f-string interpolating `d[key]`. - **Does not insert:** `key in d`, `d.get(key)`, `d.get(key, fallback)`, `d.keys()` / `d.values()` / `d.items()`, `d.pop(key, fallback)`, and `d.setdefault(key, value)` — `setdefault` inserts the value *you* passed and ignores `default_factory` entirely. That last one surprises people: on a `defaultdict`, `setdefault` behaves exactly as it does on a plain `dict`. ## Where it bites in production Take a clinical-lab result loader ingesting a 6,800-row batch into `results = defaultdict(list)`, keyed by collection date parsed out of each row. The upstream file uses a locale-dependent date format, so a slice of the batch parses into a differently-formatted key string than the rest. Nothing raises. Downstream, a summary routine written as `if results[expected_date]:` probes a handful of dates that were never populated, and each probe *creates* them. Three failures follow, in escalating order of confusion: 1. **Counts drift.** `len(results)` is larger after the report ran than before it. A metric derived from `len(results)` reports more collection dates than the batch contained. 2. **Empty groups leak into output.** `defaultdict` is a `dict` subclass, so serializers accept it happily and write out the manufactured `[]` values as if they were real, empty result sets. A consumer cannot distinguish "no results that day" from "we never looked at that day". 3. **Iteration explodes.** The moment one of those probing lookups happens inside `for date in results:`, the next step of the loop raises `RuntimeError: dictionary changed size during iteration`. This is the good outcome — it is the only one that is loud — and it is routinely misdiagnosed as a threading problem when it is a single-threaded self-inflicted insert. ## How to keep the behaviour and lose the hazard **Read with the non-inserting API.** Anywhere the intent is inspection rather than accumulation, write `d.get(key, ())` or guard with `if key in d`. Reserve the subscript for the accumulation site itself. This is the fix that generalizes; the others are safety nets. **Freeze the factory when accumulation ends.** `results.default_factory = None` after the load loop turns every later miss back into a `KeyError`. Reporting code that probes an unknown date now fails loudly instead of inventing an entry, and the data already accumulated is untouched. **Convert at the boundary.** `dict(results)` returns an ordinary `dict` with the same items and no factory. Handing that out — to a serializer, to another module, to a caller — means no downstream code can trip the behaviour, and no reviewer of that code has to know about it. **Iterate over a snapshot when a lookup is unavoidable.** `for date in list(results):` iterates a copy, so an insert during the loop no longer invalidates the iteration. Prefer removing the insert, but this is the correct patch when the lookup genuinely belongs there. **Validate keys at the edge.** The lab loader's real defect is not `defaultdict`, it is a key format that was never normalized. Normalizing dates at ingest — parsing to a `datetime.date` and keying on that object rather than on whatever string the file carried — removes the mismatch and, with it, the probing that exposed the insert. ## Interview framing The question is asked because it separates people who have *used* `defaultdict` from people who have *debugged* one. The strong answer names the mechanism (`__missing__` stores before returning), draws the insert/no-insert line correctly, and volunteers the freeze-or-convert habit without being prompted. ## Why the design is still right It is tempting to call insert-on-read a wart, but the alternative is worse. A `__missing__` that returned a detached default would make `d[k].append(v)` silently do nothing — a far quieter bug than a mapping that grew. The type is honest about what it is: a mapping whose subscript is a write. What goes wrong in practice is not the semantics but the assumption, carried over from plain `dict`, that reads are free. Treat a `defaultdict` as an accumulator with a narrow lifetime, use the non-inserting API everywhere else, and the behaviour stays where it earns its keep.

  • Which mapping operations on a defaultdict never call default_factory?
    Everything except subscripting. `key in d`, `d.get(key)`, `d.keys()`, `d.values()`, `d.items()` and `d.pop(key, fallback)` all leave the mapping alone. `d.setdefault(key, value)` also ignores `default_factory` completely — it inserts the value you passed, exactly as it would on a plain `dict`. Only `d[key]` reaches `__missing__`.
  • How would you make a populated defaultdict raise KeyError again once loading is done?
    Assign `d.default_factory = None`. `__missing__` raises `KeyError(key)` whenever the factory is `None`, so later misses fail loudly while everything already accumulated stays put. The alternative is to convert with `dict(d)` and hand out the plain mapping, which also stops any downstream code from relying on the behaviour.
  • A service logs RuntimeError: dictionary changed size during iteration and there is only one thread. What do you look for?
    A subscript on the same `defaultdict` inside the loop body. On a `defaultdict` a missing-key lookup inserts, which changes the size mid-iteration and invalidates the iterator. Find the lookup and switch it to `.get()` or an `in` test; if the lookup genuinely belongs there, iterate a snapshot with `for k in list(d):` instead.

saying these in an interview costs you the question

  • Says a lookup never mutates a defaultdict
  • Uses d[key] to test membership instead of in
  • Thinks .get() also triggers default_factory
  • Blames the RuntimeError on threading rather than the insert
  • Believes setdefault consults default_factory
  • Serializes the result without stripping manufactured empty values

context