Why does dict.update() skip an overridden __setitem__ in a dict subclass?
answer
- Only one syntax reaches your method
- Inherited C methods take a shortcut
- Bulk insertion paths go behind your back
- update, the constructor, setdefault, |=
- Wrap a dict instead of inheriting one
basics
~20 sBecause dict.update() is C code that writes straight into the dictionary's internal storage instead of calling the Python-level setitem you defined. dict.init, setdefault and the |= operator behave the same way, so an override is honoured only for plain d[key] = value assignment.
solid answer
~40 sA `dict` subclass inherits methods implemented in C, and those methods manipulate the hash table directly rather than dispatching back through the type's `__setitem__` slot. So `d[key] = value` reaches your override, but `dict.update()`, the `dict(...)` constructor, `dict.setdefault()` and `|=` all insert behind its back — and symmetrically `dict.get()`, `in` and the view methods bypass an overridden `__getitem__`. Nothing in the language guarantees that a built-in method will call the subclass's dunder; only the explicit `obj[key]` syntax does. The consequences are silent: half your entries are normalized, validated or logged and half are not. If you need every write to go through your code, override all the entry points explicitly, or wrap a real dictionary — `collections.UserDict` — instead of inheriting from `dict`.
code
python · 10 linesclass Tracked(dict):
def __setitem__(self, key, value):
print("tracked", key)
super().__setitem__(key, value)
d = Tracked()
d["a"] = 1 # prints: tracked a
d.update({"b": 2}) # silent - dict.update bypasses the override
d |= {"c": 3} # silent as well
print(sorted(d)) # ['a', 'b', 'c']go deeper
Recall the concrete fact: d[key] = value calls your __setitem__, but update() and dict(...) do not. Being able to name that one trap, and to show it with three lines at a whiteboard, is enough at this level.
Explain the mechanism, not just the symptom: inherited methods are C implementations that write into the hash table without dispatching through the type's slot. Name several bypassing methods on both the write and read sides.
Show you have debugged the partial-correctness failure this causes in a running system, and argue for a design that closes it — override the full surface, wrap a real dictionary, or expose a narrow domain API with no mapping contract at all.
Own the API-contract argument: subclassing a built-in container publishes a full mapping contract your code only partly implements, and every consumer that reaches past your methods is coupled to that gap. Decide when a container type is the right public surface at all.
## The mechanism When you write `class Tracked(dict)`, you inherit a type whose methods are C functions in CPython. `dict.update()`, `dict.__init__()`, `dict.setdefault()` and `dict.__ior__()` (the `|=` operator) are written against the concrete hash-table struct. They call the internal insertion routine directly. They do **not** perform an abstract `PyObject_SetItem` call, which is what would look up `__setitem__` on the object's type and find your override. By contrast, the *syntax* `d[key] = value` compiles to a `STORE_SUBSCR` instruction, which does go through the type's `__setitem__` slot — and that slot is your Python function. That single asymmetry is the whole phenomenon: ```python class Tracked(dict): def __setitem__(self, key, value): print("tracked", key) super().__setitem__(key, value) d = Tracked() d["a"] = 1 # tracked a d.update({"b": 2}) # silent d |= {"c": 3} # silent d.setdefault("e", 4) # silent Tracked({"f": 5}) # silent - dict.__init__ inserts directly ``` The same holds on the read side. An overridden `__getitem__` is used by `d[key]`, but `dict.get()`, the `in` operator, `dict.pop()`, `dict.items()` and iteration all read the table directly and never call it. ## Why the language works this way Two reasons, and neither is an accident. The first is speed: dictionaries are the substrate of attribute lookup, globals and keyword arguments, so their built-in methods take the shortest possible path and cannot afford a dynamic attribute lookup per insert. The second is that this is simply the ordinary rule for built-in types written in C — CPython's C-level implementations of built-ins are not obliged to re-dispatch through the Python type. The `dict` docs say so plainly: a subclass that wants consistent behaviour must override every method it cares about, because the base implementations are not guaranteed to call each other. The rule is not even uniformly *one* way, which is why memorising a list beats reasoning from a principle. `dict.fromkeys()` on a subclass, for instance, does route inserts through `__setitem__`, because it builds the new object through the abstract item-setting protocol. Deriving the behaviour from intuition will mislead you; testing it, or reading the source, will not. ## What breaks in practice The failure mode is always *partial* correctness, which is the worst kind. A case-normalizing mapping normalizes keys assigned one at a time and stores the raw keys arriving through a bulk `update()`. A validating mapping validates interactive writes and lets a batch load in unvalidated data. An auditing mapping records some writes. Nothing raises; the container is simply inconsistent, and the inconsistency depends on which call site touched it. There are also secondary surprises. `dict.copy()` on a subclass returns a plain `dict`, not your class — so does a `{**d}` expansion — which means state you added on the subclass silently disappears at the first copy. ## The three ways out **Override everything.** Keep inheriting from `dict` — you keep the C-speed reads and the fact that every consumer expecting a real `dict` accepts your object — but explicitly redefine `update`, `__init__`, `setdefault`, `__ior__`, `pop`, `copy` and anything else you rely on, funnelling each through your own `__setitem__`. It works, it is easy to leave a hole in, and each new dict method in a future release is a new hole. **Wrap instead of inherit.** `collections.UserDict` holds a real dictionary in an attribute and implements every mapping operation in Python on top of `__getitem__`/`__setitem__`, so a single override covers the whole surface. You pay in speed and in the fact that the result is not an instance of `dict`. **Do not subclass a container at all.** Often the honest answer: expose a small class with the two or three domain operations you actually need — `record()`, `lookup()` — and keep a private dictionary inside it. You then owe callers no mapping semantics whatsoever, and no one can reach past your API with an inherited method. ## What to say in an interview State the mechanism (built-in C methods bypass the Python-level slot), name at least `update`, the constructor and `setdefault` as concrete bypassers, note that the read side has the same hole, and finish with the design point: inheriting from a built-in container advertises a full contract you are only partly implementing. That last sentence is what separates a candidate who has read about the trap from one who has been bitten by it.
- Does the same hole exist on the read side?Yes. `dict.get()`, the `in` operator, `dict.pop()`, `dict.items()` and plain iteration all read the hash table directly and never call an overridden `__getitem__`; only `d[key]` does. A subclass that decorates reads therefore returns decorated values through subscription and raw values through every other accessor, which is the same partial-correctness failure seen on writes.
- If you must inherit from dict, which methods do you have to override?At minimum `__init__`, `update`, `setdefault` and `__ior__` on the write side, plus `get`, `pop`, `popitem` and `copy` if your logic touches reads or must preserve the subclass type. `dict.copy()` returns a plain `dict` unless you override it. The list is long, it grows with each release, and forgetting one is silent — which is the argument for wrapping rather than inheriting.
- Is there any dict method that does honour an overridden __setitem__?`dict.fromkeys()` does, when called on a subclass: it constructs the instance and then sets items through the abstract item-setting protocol, so your override runs. That inconsistency is the point — the behaviour is per-method implementation detail, not a rule you can derive, so verify rather than assume.
- Why doesn't CPython just make the built-in methods dispatch through __setitem__?Cost. Dictionaries underpin attribute access, module globals and keyword-argument passing, so every insertion path is performance-critical; an attribute lookup and Python-level call per item would be paid by every program. The language's answer is to offer a composition-based mapping in the standard library instead, and to document that built-in methods are not obliged to call one another.
Inheriting from dict is like repainting the front door of a warehouse: deliveries that come through the door get your treatment, but the loading bay at the back — update, the constructor, setdefault — was never told about you.
saying these in an interview costs you the question
- Claims an override is always honoured by inherited methods
- Thinks only performance, not correctness, is at stake
- Believes dict.copy() returns the subclass
- Assumes the read side is safe while writes are not
- Overrides __setitem__ and calls the container done
- Says subclassing dict is always the right way to extend a mapping