skip to content

How do you nest collections.defaultdict, and what breaks when you do?

level: seniorimportance: nice to knowfreq 25%

answer

  1. The factory returns another mapping
  2. defaultdict(defaultdict) is the near-miss
  3. An anonymous factory has no findable name
  4. Serialization needs a module-level function
  5. One lookup can create a whole branch

basics

~20 s

Nest by making the factory build the inner mapping, usually defaultdict(lambda: defaultdict(list)). The common breakage is that a lambda factory cannot be pickled, so the structure will not cross a process boundary or into a cache; a module-level named function fixes it.

solid answer

~50 s

The factory must be a callable returning the inner mapping, so `defaultdict(list)` cannot be nested directly — you write `defaultdict(lambda: defaultdict(list))`, or for arbitrary depth a named function that returns `defaultdict` of itself. Three things then break. **Pickling**: a `lambda` is not picklable, so `pickle.dumps` fails and anything that pickles under the hood — sending the structure to a worker process, caching it to disk — fails with it; a module-level `def` is picklable and fixes it. **Silent branch creation**: one probing lookup at the outer level now manufactures a whole empty subtree, and because `defaultdict` is a `dict` subclass, serializers write those phantom branches out as real data. **Readability**: nested `repr` output is dense and the structure has no schema, so nothing catches a typo in either level's key. Convert with a recursive `dict()` before the structure leaves the function.

code

python · 16 lines
python
import pickle
from collections import defaultdict

def batch_index():
    return defaultdict(list)

nested = defaultdict(batch_index)
nested["2026-03-11"]["K"].append(4.1)
print(pickle.loads(pickle.dumps(nested))["2026-03-11"]["K"])

anon = defaultdict(lambda: defaultdict(list))
anon["2026-03-11"]["K"].append(4.1)
try:
    pickle.dumps(anon)
except pickle.PicklingError as exc:
    print("cannot pickle:", exc)

go deeper

for a junior

Know that nesting means the factory returns another mapping, and that defaultdict(lambda: defaultdict(list)) is the usual two-level spelling. Recognising the shape when you read it is enough at this level.

for a middle

Explain why defaultdict(defaultdict) fails, and be able to write the recursive factory. Know that the factory is stored on the mapping and therefore travels with it when the structure is copied or serialized.

for a senior

Show that you anticipate the failure far from its cause: a lambda factory breaks pickling, and pickling happens implicitly whenever the structure crosses a process boundary or reaches a disk cache. Convert to plain dicts before the structure leaves your function.

for a principal

Judge whether the nest should exist. Past two levels an unschematized nested mapping is a data model nobody wrote down; decide when a tuple key, a small class, or a typed record is the right shape for data that outlives one function.

## Building the nest `default_factory` is called with no arguments and its return value becomes the stored value, so nesting is just a factory that returns another mapping. `defaultdict(defaultdict)` almost works and is the classic near-miss: `defaultdict()` with no argument produces an inner mapping whose own factory is `None`, so the second level raises `KeyError`. The two shapes that do work: ```python # fixed two levels by_date = defaultdict(lambda: defaultdict(list)) by_date["2026-03-11"]["K"].append(4.1) # arbitrary depth, by recursion def tree(): return defaultdict(tree) ``` The recursive form is often called autovivification, borrowed from another language's mappings: any chain of subscripts you write into existence simply exists. ## Breakage 1: the factory has to be picklable Pickling a `defaultdict` pickles `default_factory` **by reference**, i.e. by module and qualified name. A `lambda` has no findable name, so `pickle.dumps` raises `PicklingError` — `Can't pickle <function <lambda>>`. `list`, `int`, `set` and any module-level `def` are all findable and pickle fine. This matters more than it sounds, because pickling happens implicitly. Sending the structure to a worker process serializes it: on 3.14 the `multiprocessing` default start method is `forkserver` on Unix other than macOS and `spawn` on macOS and Windows, and both of those transfer arguments by pickling them — only the explicitly-requested `fork` method sidesteps it by inheriting memory. Writing the structure to a disk cache serializes it. Passing it across a process pool boundary serializes it. The failure therefore shows up far from the line that chose the `lambda`, and the fix is a one-line rewrite to a module-level factory function. ## Breakage 2: a lookup now creates a whole subtree The insert-on-miss behaviour compounds with depth. On a single-level accumulator, a probing lookup costs you one empty list. On `defaultdict(lambda: defaultdict(list))`, `by_date[some_date]` creates an entire empty inner mapping, and `by_date[some_date][some_analyte]` creates the inner mapping *and* an empty list inside it. A clinical-lab loader ingesting a 6,800-row batch keyed by collection date and analyte will, after one report routine probes a handful of dates that a locale-dependent date format never produced, hold phantom branches indistinguishable in shape from real ones. And nothing downstream objects. `defaultdict` is a `dict` subclass, so serializers accept it and write the phantom branches out as genuine empty result sets. The structure has no schema and no validation layer, so a typo in either key level is invisible until a human reads the output and asks why a date nobody collected on has an entry. ## Breakage 3: it is hard to read and hard to hand over `repr` on a nested `defaultdict` interleaves factory objects with data and becomes unreadable past two levels. More importantly the type communicates nothing about the intended shape — `defaultdict(lambda: defaultdict(list))` names neither level, so a reader has to reconstruct "date to analyte to values" from the accumulation code. ## What to do instead, and when to accept it **Convert at the boundary.** Recursively rebuild as plain `dict`s once accumulation is finished — walk the structure and call `dict()` on every level. What leaves the function then has ordinary lookup semantics, an ordinary `repr`, and no factory to serialize. **Name the factory.** Even for a fixed two levels, a module-level `def` beats a `lambda`: it is picklable, it is greppable, and its name (`analyte_index`, `per_date_buckets`) documents the level it builds. **Stop at two levels.** Beyond that, a flat mapping keyed by a tuple — `results[(date, analyte)]` — is usually easier to reason about, easier to aggregate, and immune to half-created branches. A dedicated small class, or a mapping of dataclass values, beats both when the shape is fixed and known. **Accept the nest** when the work is genuinely a short-lived in-function accumulation that never crosses a process boundary and is converted before it is returned. That is a real and common case, and there the nested form is the clearest thing to write. `copy.deepcopy` handles nested `defaultdict`s correctly, factories included, so in-process copying is not among the hazards — it is specifically serialization and the silent branch creation that bite. ## Other factory shapes worth knowing The factory is only required to be callable with no arguments, which leaves more room than the `lambda` habit suggests. `functools.partial` binds arguments ahead of time and is picklable when the underlying function is, so `partial(dict.fromkeys, ANALYTES)`-style defaults travel fine. A class works directly, since instantiating it takes no arguments when `__init__` has defaults, and it gives the inner level a name and a schema in one move. And `dict` itself is the right factory when the inner mapping should behave normally — the outer level fills in, the inner level raises `KeyError` on a bad key, which is often exactly the boundary you want.

  • Why does defaultdict(defaultdict) not give you a working two-level mapping?
    Because the factory is `defaultdict` itself, called with no arguments. That produces an inner `defaultdict` whose own `default_factory` is `None`, so the first level fills in fine and the second raises `KeyError`. You need a factory that returns an inner mapping *with* a factory attached — a `lambda` returning `defaultdict(list)`, or a named function that does the same.
  • How would you detect phantom branches that a probing lookup created in a nested structure?
    Look for the factory's fingerprint: a branch whose leaves are all empty containers. That is heuristic, though, because a genuinely empty group looks identical. The reliable move is prevention — freeze the outer factory with `default_factory = None` once loading ends, or convert the whole structure to plain `dict`s, so later probing raises `KeyError` instead of creating anything.
  • When would you drop the nesting entirely and key on a tuple instead?
    Once the shape exceeds two levels, or when you need to aggregate across the inner dimension. A flat mapping keyed by `(date, analyte)` cannot hold half-created branches, iterates in one pass, and groups by either component with a single expression. The nested form pays off only when the inner mappings are themselves handed around as units.

saying these in an interview costs you the question

  • Writes defaultdict(defaultdict) and expects two working levels
  • Assumes any factory pickles, lambdas included
  • Forgets that one lookup creates a whole empty branch
  • Serializes the nest without converting to plain dicts
  • Nests four levels deep rather than keying on a tuple

context