skip to content

Where does a collections.UserDict subclass store its items, and why does that matter?

level: juniorimportance: should knowfreq 30%

answer

  1. It holds a container, it is not one
  2. One public attribute carries the storage
  3. Written in Python, so overrides always run
  4. The attribute is literally named data
  5. isinstance against dict returns False

basics

~10 s

In an ordinary dict held on the instance as its data attribute. Because UserDict is written in Python and funnels every operation through __getitem__/__setitem__, your overrides are honoured even by the constructor and update().

solid answer

~40 s

`collections.UserDict` keeps the real mapping in a plain dict exposed as the instance's `data` attribute, and implements the whole mapping interface in Python on top of `__getitem__`, `__setitem__`, `__delitem__`, `__iter__` and `__len__`. That indirection is the point: override one of those five and every other operation — construction, `update()`, `setdefault()`, `get()`, `pop()` — goes through it, unlike a `dict` subclass whose inherited C methods bypass overrides. `collections.UserList` and `collections.UserString` follow the same design with a `list` and a `str` in `data`. The cost is that the object is *not* a `dict`: `isinstance(u, dict)` is False, so code demanding a real dict (JSON serialisation, `exec()` globals, C-level APIs) rejects it, and each operation costs several times a raw dict write. Note also that writing to `data` directly bypasses your own overrides.

code

python · 11 lines
python
from collections import UserDict

class CaseInsensitiveDict(UserDict):
    def __setitem__(self, key, value):
        super().__setitem__(key.lower(), value)

headers = CaseInsensitiveDict(Alpha=1)   # constructor routes through __setitem__
headers.update({"BETA": 2})              # so does update()
headers["Gamma"] = 3
print(headers.data)                      # {'alpha': 1, 'beta': 2, 'gamma': 3}
print(isinstance(headers, dict), type(headers.data).__name__)

go deeper

for a junior

Remember two things: the storage is an ordinary dict on the instance's data attribute, and collections.UserDict is what you pick when your overrides must actually be used everywhere. Be able to show a two-line subclass.

for a middle

Explain the delegation: UserDict implements five primitives over data and inherits the rest of the mapping interface from an abstract base that is written in terms of them, which is why update() honours your __setitem__.

for a senior

Weigh the costs out loud — not an instance of dict, so serialisation and C-level APIs reject it; several times slower per operation; data is public and can be written past your own invariant — and say where you would put the conversion boundary.

for a principal

Set the house rule for custom mappings: when a team should use UserDict, when to implement collections.abc.MutableMapping over private storage, and when a mapping interface is the wrong abstraction to expose at all.

## What UserDict is `collections.UserDict` is a mapping written in pure Python whose entire storage is an ordinary `dict` kept on the instance under the public attribute name `data`. `collections.UserList` and `collections.UserString` are the same idea for `list` and `str`. They exist for one reason: to give you a container whose behaviour you can genuinely override. The contrast is with subclassing the built-in. When you subclass `dict`, the methods you inherit are C functions that operate on the concrete hash table; an overridden `__setitem__` is never consulted by `update()`, `setdefault()` or the constructor. `UserDict` has no such shortcut available, because it *is not* a dict — it holds one. Its `update()` is inherited from `collections.abc.MutableMapping` and is written as a loop of `self[key] = value`, which means it goes through whatever `__setitem__` your subclass defines. Same for `setdefault()`, `pop()`, `get()`, `clear()` and the rest. ## The five methods that matter `UserDict` derives from `collections.abc.MutableMapping`, so the abstract base supplies the whole mapping surface as mixin methods defined in terms of five primitives: `__getitem__`, `__setitem__`, `__delitem__`, `__iter__` and `__len__`. `UserDict` implements those five against `data` and inherits the rest. Your subclass overrides one or more of the five, and the mixins keep working — that is the delegation contract you are buying. A concrete example: a mapping that lower-cases its keys. Override `__setitem__` to call `super().__setitem__(key.lower(), value)`, and then the constructor (`UserDict(Alpha=1)`), a later `update({"BETA": 2})` and plain assignment all normalise. Written as a `dict` subclass, only the plain assignment would. ## What it costs **It is not a dict.** `isinstance(u, dict)` is False. Anything that checks for a real `dict` rather than for the mapping interface will refuse it. Standard-library examples you can verify in a REPL: JSON serialisation raises `TypeError: Object of type ... is not JSON serializable`, and `exec()` rejects it with `exec() globals must be a dict`. Where a real dict is required, pass `u.data` or `dict(u)` at the boundary. Conversely `isinstance(u, collections.abc.MutableMapping)` is True, so code that checks the ABC — the better habit — accepts it. **It is slower.** Every read and write is Python-level: an attribute lookup for `data` plus a method call, instead of one C-level hash operation. Measured on a modern CPython the gap for simple writes is a small multiple, not a rounding error. That is irrelevant for a config object rebuilt once per request and very relevant for a hot inner loop. **`data` is public and unguarded.** It is deliberately part of the API — you use `self.data` inside your overrides, and callers can reach it for interoperability. The flip side is that `u.data[key] = value` writes past your own normalisation. If the invariant must be absolute, keep the storage private with `collections.abc.MutableMapping` and your own attribute instead, or hide it behind composition. ## UserList and UserString `UserList` stores a `list` in `data` and derives from `collections.abc.MutableSequence`; `UserString` stores a `str` in `data` and behaves as an immutable sequence whose methods return new `UserString` objects rather than `str`. `UserList` has one edge worth knowing: its `__getitem__` rebuilds a slice as `self.__class__(...)`, which raises `TypeError` if your subclass's `__init__` requires extra arguments. That is the mirror image of the built-in problem — where `list` slicing loses your subclass, `UserList` slicing tries to preserve it and can fail. ## When to reach for it Use `UserDict` when you need a small number of behavioural overrides on a mapping, the object stays inside your own code, and performance is not the constraint: case-insensitive headers, key namespacing, a mapping that records every write, a mapping with validation. Reach for `collections.abc.MutableMapping` directly when you want control over the storage (not necessarily a dict at all — a file, a remote store, two dicts in a layered lookup). Reach for composition when the object should not present a mapping interface in the first place. And subclass `dict` itself only when you genuinely need a real dict for a C-level or serialisation boundary and are prepared for inherited methods to bypass you. ## Answering it in an interview Say `data`, say it is a plain dict, and say *why*: pure-Python delegation means overrides are honoured everywhere. Then volunteer the cost — not an instance of `dict`, several times slower — because that pair is what separates recall from understanding.

  • What breaks when you hand a collections.UserDict to code that expects a real dict?
    Anything that type-checks with `isinstance(obj, dict)` or requires the C dict layout rejects it: JSON serialisation raises TypeError, and `exec()` refuses it as a globals mapping. The fix at the boundary is to pass `u.data` or `dict(u)`. Code that checks `collections.abc.MutableMapping` instead accepts it, which is the more robust habit for library authors.
  • When would you use collections.abc.MutableMapping instead of UserDict?
    When you want to own the storage rather than inherit a dict in a public `data` attribute — a mapping backed by a file, a remote store, or two dicts in a layered lookup — or when the invariant must be airtight, since `data` is reachable and writing to it bypasses your overrides. MutableMapping gives the same mixin guarantee but requires you to implement the five primitives yourself.

A dict subclass is a rebuilt engine with factory parts you cannot reach; UserDict is a new dashboard bolted onto a stock engine sitting in plain view — every control you add really is in the loop.

saying these in an interview costs you the question

  • Says UserDict subclasses dict, so isinstance passes
  • Cannot name the data attribute at all
  • Thinks UserDict is faster than a plain dict
  • Believes UserDict stores items in __dict__
  • Claims dict and UserDict are interchangeable everywhere

context