skip to content

Why do dict.get() and the in operator bypass an overridden __getitem__ in a dict subclass?

level: middleimportance: should knowfreq 32%

answer

  1. Overriding one method is not overriding the type
  2. Which reads actually run Python code?
  3. C methods read the hash table directly
  4. Only the subscript hits the type slot
  5. Mapping mixin methods are written as self[key]

basics

~20 s

dict.get() and the in operator are C methods that read the instance's hash table directly; only d[key] dispatches to a Python-level __getitem__. Subclassing dict lets you override one entry point, not the whole read path.

solid answer

~50 s

Only subscript syntax dispatches to your override. `d[key]` goes through the `__getitem__` slot on the type, so it runs your code. `dict.get`, `dict.__contains__` (which is what `key in d` uses), the `.keys()`/`.values()`/`.items()` views, `dict.copy`, `dict(d)` and `{**d}` are all C functions on `dict` that read the instance's hash table directly — a subclass instance is still a real `dict` underneath, so they never re-enter Python to consult an override. `json.dumps` is the same story: its C encoder iterates the table, so it serialises the raw stored values. The lesson is that overriding `__getitem__` intercepts one entry point, not the read path. If every mapping read must go through your logic, use `collections.UserDict` — its `Mapping` mixin methods are written as `self[key]` — or drop the mapping interface entirely and expose an explicit method by composition.

code

python · 13 lines
python
import json


class Vault(dict):
    def __getitem__(self, key):
        return super().__getitem__(key).upper()


v = Vault(token="secret")
print(v["token"])          # SECRET
print(v.get("token"))      # secret
print(list(v.values()))    # ['secret']
print(json.dumps(v))       # {"token": "secret"}

go deeper

for a junior

Remember the concrete result: d[key] runs your overridden __getitem__, while d.get(key) and key in d do not. Be able to say that a dict subclass keeps the built-in's own storage and its own fast methods.

for a middle

Explain the mechanism: subscript syntax dispatches through the type's __getitem__ slot, whereas dict.get, dict.__contains__ and the .values()/.items() views are C code reading the instance's hash table. Name collections.UserDict as the variant whose mixin methods go through self[key].

for a senior

Show why this ships as a bug: the hand-tested subscript path looks right while serialisation, copying and membership leak raw values. Be ready to say how you would detect it and when you would take the per-call cost of UserDict versus dropping the mapping interface for an explicit method.

for a principal

Own the design position: inheriting from a built-in container buys C speed and inherits its shortcuts, so it cannot carry semantics reliably. Argue for composition with a named API at the boundary, and for keeping decryption or lazy loading out of a type that other code is entitled to treat as a plain dict.

Two very different operations both get called "a dict lookup", and only one of them goes through your class. **Syntax goes through the type; C methods do not.** When you write `d[key]`, the interpreter resolves the subscript slot on `type(d)` and calls it. If your subclass defines `__getitem__`, that is the function that runs. When you write `d.get(key)`, `key in d`, `d.values()`, or hand `d` to a C consumer such as the JSON encoder, none of those consult `type(d).__getitem__`. They are separate C functions defined on `dict` that walk the hash table stored inside the object and hand back whatever is stored there. Overriding `__getitem__` overrides *one entry point* into that table; it does not replace the table or the other doors into it. **Why CPython works this way.** An instance of a `dict` subclass really is a `dict` at the C layer: same storage, same layout, and C code can check for it cheaply and then read it directly. `dict.get`, `dict.__contains__`, `dict.setdefault`, `dict.pop`, `dict.copy` and the view objects are all implemented over that structure. Making each of them re-enter the interpreter to look for a possibly-overridden `__getitem__` would put a full Python-level call on every read of every dict in the program, for the benefit of the rare subclass that overrides it. So they read the table. This is a CPython implementation trait rather than a language guarantee, but it has held for the entire 3.x line and is unchanged in 3.14. **Concretely, for a subclass whose `__getitem__` transforms values on read:** ```python import json class Vault(dict): def __getitem__(self, key): return super().__getitem__(key).upper() v = Vault(token="secret") v["token"] # 'SECRET' -> override ran v.get("token") # 'secret' -> bypassed "token" in v # True -> dict.__contains__, override irrelevant list(v.values()) # ['secret'] -> the view reads the table json.dumps(v) # '{"token": "secret"}' dict(v), {**v}, v.copy() # all raw, and copy() returns a plain dict repr(v) # raw as well ``` `json.dumps` is the one that bites in production. With default arguments it uses the C-accelerated encoder, and for anything that passes an `isinstance(o, dict)` check it iterates the underlying table directly — it never calls `.items()`, `.get()` or `__getitem__`. So the object you serialise into a log line or an HTTP body carries the raw stored values, while the hand-written code path you tested in the REPL with `v["token"]` looked perfect. **The diagnostic rule.** Subscript syntax (`d[key]`, `d[key] = x`, `del d[key]`) reaches `__getitem__` / `__setitem__` / `__delitem__`. Anything spelled as a method on the built-in type, and any consumer that treats the object as a concrete dict, does not. Note that `key in d` *looks* like syntax but resolves to `__contains__`, which `dict` already implements in C — so overriding `__getitem__` has no effect on it. If you want certainty for a given call, instrument the override with a counter or a print and exercise the real code path; guessing which stdlib function takes the fast path is how this bug ships. **The two honest fixes.** 1. **`collections.UserDict`.** It is not a `dict` at the C layer — it is a Python `MutableMapping` that keeps the real mapping in a plain attribute and implements the mapping interface in Python. The mixin methods it inherits from `collections.abc.Mapping` and `collections.abc.MutableMapping` are written in terms of `self[key]`, so `get()`, `pop()`, `setdefault()`, `values()` and `items()` all route through your override. The costs are real: every operation is a Python-level call, and because the object is not a `dict`, C consumers that demand a concrete dict reject it — `json.dumps` raises `TypeError` unless you serialise the inner mapping or pass a `default` hook. One wrinkle even here: `UserDict` defines `__contains__` as a direct check against its inner mapping, so membership still skips your override. 2. **Composition with an explicit API.** Stop advertising a mapping. Hold the storage privately and expose a named method — a `read(key)` that decrypts, a `fetch(key)` that lazily loads. Nothing in the runtime has a fast path for a method you invented, so nothing can bypass it, and callers can no longer accidentally reach the raw values by writing `.get()`. **The point an interviewer is testing.** Subclassing a built-in container inherits the C type's speed and its shortcuts together. You get to intercept only the operations routed through type slots, which is a small and non-obvious subset. When the override carries semantics — decryption, lazy loading, auditing, redaction — partial interception is more dangerous than none, because the class behaves correctly in the paths you exercised by hand and silently leaks the raw values everywhere else.

  • Why does collections.UserDict behave differently here?
    `UserDict` is not a `dict` at the C layer — it is a Python `MutableMapping` holding the real mapping in a plain attribute. The methods it inherits from `collections.abc.Mapping` and `MutableMapping` are written in Python as `self[key]`, so `get()`, `pop()`, `setdefault()`, `values()` and `items()` all dispatch to your `__getitem__`. The price is a Python-level call per operation. One exception: `UserDict` defines `__contains__` as a direct check against its inner mapping, so membership still skips the override.
  • Why does json.dumps() emit the raw stored values for such a subclass?
    With default arguments `json.dumps` uses the C-accelerated encoder. For an object that passes an `isinstance` check against `dict` it iterates the instance's own hash table rather than calling `.items()` or `__getitem__`, so it serialises exactly what is stored. Hand it a `collections.UserDict` instead and it raises `TypeError`, because that object is not a `dict` — you then serialise the inner mapping or supply a `default` hook on `json.JSONEncoder`.
  • How would you prove which reads actually reach your override?
    Instrument it: have `__getitem__` bump a counter or print, then run the real code path — the caller, the serialiser, the template render — rather than a REPL subscript. Anything that increments the counter dispatches; anything that returns a value without incrementing it took the C fast path. That measurement is cheap and beats guessing which stdlib function happens to treat your object as a concrete dict.

You put a receptionist on the front door of the building, but the loading bay, the fire exit and the service lift all open onto the same warehouse floor. Anyone who knows those doors walks straight past your receptionist.

saying these in an interview costs you the question

  • Assumes overriding __getitem__ covers every read of the mapping
  • Thinks dict.get() is implemented as a call to __getitem__
  • Believes `key in d` calls __getitem__ and catches KeyError
  • Expects .values() and .items() to show transformed values
  • Says json.dumps() reads a dict through its public methods
  • Claims a dict subclass copy() returns the subclass

context