skip to content

Subclassing Built-in Types

Subclass dict or list and the C-level methods keep bypassing your overrides, so update() never calls your __setitem__. collections.UserDict exists for exactly that, and interviewers like the trap.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why does d.copy() on a dict subclass return a plain dict while copy.copy(d) does not?

level: middleimportance: must knowfreq 45%

answer

  1. Two copy routes, not one
  2. C constructors build the exact type
  3. Extra attributes live in the instance __dict__
  4. The copy module goes through the reduction protocol
  5. Check type() of the result before trusting it

basics

~20 s

dict.copy(), dict(d) and {**d} are built-in constructors: they allocate an exact dict and copy the items, so your class and its instance attributes are lost. copy.copy(d) goes through the reduction protocol and rebuilds an instance of your subclass.

solid answer

~40 s

There are two copying machineries. `d.copy()`, `dict(d)`, `{**d}` and the `|` merge operator are implemented on the concrete `dict` type in C: they allocate a brand-new exact `dict` and fill it with the pairs they read, never asking what class `d` was. Your subclass is gone, and so is anything you stored on the instance `__dict__`, because a plain `dict` has nowhere to put it. `copy.copy(d)` takes a different route: it looks for a copy hook, then falls back to `d.__reduce_ex__()`, which reports your class, the instance state and an item iterator. `copy` rebuilds via `cls.__new__(cls)` — note `__init__` is skipped — restores the attributes, and re-inserts the items through ordinary subscript assignment. `copy.deepcopy()` and `pickle` preserve the class for the same reason. `fromkeys()` also keeps it, being an inherited classmethod.

code

python · 12 lines
python
class Headers(dict):
    def __init__(self, source, /, **kw):
        self.source = source
        super().__init__({k.lower(): v for k, v in kw.items()})

    def __setitem__(self, key, value):
        super().__setitem__(key.lower(), value)


h = Headers("upstream", Accept="json")
print(type(h.copy()), type(dict(h)), type({**h}))
print(hasattr(h.copy(), "source"))

go deeper

for a junior

Remember the headline: copying a dict subclass with .copy(), dict(d) or {**d} normally hands you a plain dict. Check type() of anything you copy before assuming your class and its extra attributes came along.

for a middle

Be ready to explain both routes: the built-in constructors allocate an exact dict in C and never consult the source's class, while copy.copy() goes through __reduce_ex__, rebuilds with cls.__new__(cls), restores the instance __dict__ and re-inserts the items.

for a senior

Show where the downgrade actually bites in a running system — a normalizer or cache layer that returns dict(arg), an invariant that stops being enforced at the next write, an isinstance check failing far from the copy — and how you defend: composition, collections.UserDict, or an explicit copy() override with a test.

for a principal

Own the standing question: should this codebase carry dict subclasses at all, given that every boundary can silently downgrade them and type checkers will not flag it? Be ready to argue for an explicit mapping type or plain composition as the convention, and to say what you would migrate first.

### Two different copying machineries A `dict` subclass has two entirely separate ways of being copied, and they disagree about what the result *is*. The first is the built-in machinery: the bound method `dict.copy()`, the `dict(...)` constructor and the `{**d}` unpacking form. All three are implemented in C on the concrete `dict` type. They allocate a brand-new *exact* `dict` object and fill it with the key/value pairs they read out of the source mapping. Nothing in that path asks what class the source was. So for ```python class Headers(dict): def __init__(self, source, /, **kw): self.source = source super().__init__({k.lower(): v for k, v in kw.items()}) def __setitem__(self, key, value): super().__setitem__(key.lower(), value) ``` `type(h.copy())`, `type(dict(h))` and `type({**h})` are all plain `dict`. The `source` attribute is gone too, because it never lived in the dict storage at all: it lives in the instance `__dict__` that your subclass gets on top of the dict, and a fresh exact `dict` has no instance `__dict__` to put it in. The `|` merge operator behaves the same way — `h | {"x": 1}` is a plain `dict`. The one built-in that does keep your class is `fromkeys()`, because it is inherited as a classmethod and is therefore called with `cls` bound to your subclass. The second machinery is the `copy` module. `copy.copy(d)` looks for a copy hook on the class, then falls back to the pickle reduction protocol: it calls `d.__reduce_ex__(4)`, which for a `dict` subclass returns a constructor callable plus your class, the instance `__dict__` as state, and an item iterator. `copy` then builds a new object with `cls.__new__(cls)`, restores the state onto the new instance's `__dict__`, and re-inserts the items one at a time. The result is an instance of your subclass, carrying `source`, and — because the re-insertion is ordinary subscript assignment — the class's own `__setitem__` runs for every item. `copy.deepcopy()` follows the same route, and `pickle` round-trips preserve the class for exactly the same reason. ### The consequences worth naming **`__init__` is not called.** Reconstruction goes through `__new__`, not `__init__`. A subclass whose `__init__` opens a connection, registers the instance in a table, or computes a derived field will not do any of that on a copy; only whatever ended up in the instance `__dict__` is carried over. That is usually what you want, and occasionally a nasty surprise. **The invariant is not the type.** In the example above the class exists to keep every key case-folded. `plain = h.copy()` still *contains* folded keys — the folding already happened — but the returned object no longer enforces anything, so `plain["X-Trace"] = "1"` stores an unfolded key and a later `plain["x-trace"]` lookup misses. The bug does not appear at the copy; it appears at the next write, possibly in a different module, which is what makes it expensive to find. **Every boundary is a downgrade point.** Any function that takes your mapping and returns `dict(arg)`, `{**arg}` or `arg.copy()` — normalizers, cache layers, serializers, "defensive copy" helpers, functions building keyword arguments — hands back a plain `dict`. You cannot audit all of them, and a type checker will not flag it, because `dict.copy()` is annotated as returning `dict[K, V]`, which is a perfectly valid supertype. Code downstream that does `result.source` or `isinstance(result, Headers)` then fails at runtime. **Serialization is a partial exception.** `pickle` uses the same reduction protocol as `copy`, so a pickled and unpickled instance is still your subclass, attributes and all. Text formats are not so kind: a JSON encoder walks the mapping and writes the pairs, so `source` and the class simply are not in the output, and what comes back is a plain `dict`. If your subclass carries meaning, that meaning has to be encoded as data somewhere, not left implicit in the type. ### How this shows up in practice The tell is a `TypeError` or `AttributeError` on a line nowhere near the copy — code doing `result.source` on something that is now a bare `dict`, or a lookup that misses because the folding quietly stopped happening. A short REPL check settles it: build an instance, run each of `d.copy()`, `dict(d)`, `{**d}` and `copy.copy(d)`, and print `type(...)` of each. Do that once for any subclass you are considering shipping, before you rely on the type surviving. ### What to do about it If the subclass exists only to hold extra attributes, it should probably not be a `dict` subclass: compose — hold a `dict` in an attribute and expose exactly the operations you need. If it exists to change mapping behaviour, `collections.UserDict` is the honest base: it stores everything in a plain `dict` under `self.data`, all mutations go through Python-level `__setitem__`, and its `copy()` is implemented in Python so it returns your class. Even then `dict(u)` still yields a plain `dict`, because that is the built-in constructor consuming a mapping — `UserDict` fixes the methods your class inherits, not what other people's constructors do with it. If you must keep a `dict` subclass, make the copying explicit: override `copy()` to return `type(self)(...)`, and test it. A one-line test asserting `type(d.copy()) is MyDict` is worth more than the docstring explaining why it might not be.

  • What do the `|` merge operator and `fromkeys()` give back when the operand is a dict subclass?
    `h | {"x": 1}` returns a plain `dict` — the merge operator is implemented on the concrete type and builds an exact `dict`, exactly like `dict(h)`. `MyDict.fromkeys(keys)` is the exception: it is inherited as a classmethod, so it is called with `cls` bound to your subclass and constructs an instance of it. That asymmetry is worth remembering, because it means "built-in methods always downgrade" is not quite the rule.
  • If I switch the class to collections.UserDict, which of those copies keep my class?
    `UserDict.copy()` and `copy.copy()` both return your class — `UserDict` is ordinary Python, keeps its items in `self.data`, and its `copy()` delegates to the copy module. `dict(u)` still gives a plain `dict`, because that is the built-in constructor consuming a mapping, and nothing about your base class changes what someone else's constructor builds. `UserDict` fixes the methods you inherit, not the calls other code makes on you.
  • Why does copy.copy() not call the subclass's __init__, and when does that bite?
    Reconstruction goes through `cls.__new__(cls)`, then the instance `__dict__` is restored from the reduction state and the items are re-inserted. `__init__` never runs. That is fine when all your state is plain attributes, and a real problem when `__init__` does work — opening a resource, registering the instance somewhere, deriving a cached field from arguments that were not stored. Those side effects silently do not happen for a copied object.

Photocopying a labelled folder: the built-in copy gives you the pages in a blank folder, with the label and the filing rules left behind, while the copy module reproduces the folder itself, label and all.

saying these in an interview costs you the question

  • Says dict.copy() returns the subclass because copy is inherited
  • Assumes {**d} or dict(d) preserves the subclass type
  • Expects instance attributes to survive a built-in copy
  • Calls copy.copy a slower alias for the built-in copy method
  • Believes copy.copy calls the subclass __init__ on the new object
  • Claims switching to UserDict makes dict(d) return the subclass

context

open as a page

Where does a collections.UserDict subclass store its items, and why does that matter?

level: juniorimportance: should knowfreq 30%

basics

~10 s

In an ordinary dict held on the instance as its data attribute. Because UserDict is written in Python and funnels every operation through __getitem__/__setitem__, your overrides are honoured even by the constructor and update().

open as a page

Why do dict.get() and the in operator bypass an overridden __getitem__ in a dict subclass?

level: middleimportance: should knowfreq 32%

basics

~20 s

dict.get() and the in operator are C methods that read the instance's hash table directly; only d[key] dispatches to a Python-level __getitem__. Subclassing dict lets you override one entry point, not the whole read path.

open as a page

A TTL cache subclasses dict but a 6-hour nightly invoice-PDF run keeps serving stale values — how do you choose between dict, collections.UserDict and composition?

level: seniorimportance: should knowfreq 35%

basics

~20 s

The inherited dict methods never consult the expiry logic, so any caller using update(), setdefault() or the constructor writes past it. Choose by invariant: no invariant, subclass dict; overrides must always run, use collections.UserDict; a restricted surface, compose.

open as a page