Why are collections.abc.MutableMapping's inherited methods slow in a hot ETL export loop, and which should you override?
answer
- Completeness is not the same as speed
- The inherited code only knows your five methods
- Misses pay for an exception
- Emptying it repeatedly rescans from the front
- Override the hot methods to delegate
basics
~20 sEvery inherited method is generic Python code routed through your five abstract methods: membership and get catch a KeyError from getitem, update assigns one key at a time, and clear pops repeatedly, which is quadratic. Override the hot ones to delegate to the backing store.
solid answer
~40 sThe mixins buy completeness, not speed. `__contains__`, `get`, `pop` and `setdefault` are `try: self[key] except KeyError`, so a miss pays for raising and catching an exception; `update` performs one `__setitem__` per pair with no bulk path; `__eq__` materialises a `dict` from both sides; and `clear` calls `popitem` in a loop, each call doing `next(iter(self))` over a store that is being emptied — measurably quadratic. In a microbenchmark on 3.14 a dict-backed subclass runs roughly 3x slower than the raw `dict` on hits and 6x on misses. Profile first, then override only the hot methods to delegate straight to the backing store — they are ordinary methods, overriding them is expected. If nearly all of them need overriding, the base class is buying you nothing and you should wrap a `dict` directly.
code
python · 29 linesfrom collections.abc import MutableMapping
class RowBuffer(MutableMapping):
def __init__(self):
self._store = {}
# the five abstract methods
def __getitem__(self, key): return self._store[key]
def __setitem__(self, key, value): self._store[key] = value
def __delitem__(self, key): del self._store[key]
def __iter__(self): return iter(self._store)
def __len__(self): return len(self._store)
# overrides that skip the generic mixins on the hot path
def __contains__(self, key): return key in self._store
def get(self, key, default=None): return self._store.get(key, default)
def keys(self): return self._store.keys()
def items(self): return self._store.items()
def values(self): return self._store.values()
def clear(self): self._store.clear() # inherited clear() is quadratic
def update(self, other=(), /, **kwds): self._store.update(other, **kwds)
buf = RowBuffer()
buf.update({"order_id": 41, "region": "eu"})
print("order_id" in buf, buf.get("missing"), len(buf))
buf.clear()
print(len(buf))go deeper
Know that the methods you did not write still run real Python code on every call, so a class that inherits a full mapping interface is not automatically as fast as the dictionary underneath it.
Be able to open the inherited implementations in your head: membership catching KeyError, update assigning key by key, clear looping over popitem. Explaining those three is what makes the slowness predictable rather than mysterious.
Show the diagnosis, not just the fact: profile pointing into the standard library, a ratio against a raw dict with the same contents, and a short list of overrides ordered by measured cost — with the semantics of each override argued, especially where setitem normalises keys.
Own the call on whether the interface promise is worth the per-call tax at your volumes, and say when the honest answer is to drop the abstract base class and expose a narrower type instead of overriding almost every method it gave you.
## The scenario A nightly ETL export to a warehouse buffers each row's columns in a class that inherits `collections.abc.MutableMapping` over a plain `dict`, flushes at a batch boundary and calls `clear()` before the next batch. It was fine at ten thousand rows. At a few million it dominates the job, and a profile blames three names: `__contains__`, `update` and `clear`. Nothing in the class looks slow — which is exactly the point, because the slow code is the code nobody wrote. ## Why the inherited methods cost what they cost Every method the base class hands you is Python-level source in the standard library, expressed only through your five abstract methods. It cannot do better: it has no idea what your storage is. Concretely: * `__contains__(key)` is `try: self[key]` / `except KeyError: return False`. A hit costs an extra Python frame; a **miss** costs raising and catching an exception, which is far more expensive than a failed hash lookup. * `get`, `pop` and `setdefault` have the same shape and the same miss penalty. * `update(other)` iterates the other mapping and performs one `self[key] = value` per pair. There is no bulk path, so a thousand-column update is a thousand Python-level calls into `__setitem__`. * `popitem()` does `next(iter(self))`, reads that key, deletes it and returns the pair. * `clear()` calls `popitem()` in a loop until `KeyError`. Each call builds a fresh iterator over a store that is being emptied from the front, so the scan to reach the first surviving key grows as the loop proceeds. Timed on 3.14 over a dict-backed subclass, clearing 10k / 20k / 40k keys took roughly 0.017s / 0.056s / 0.207s — four times the work per doubling. **That is quadratic**, and it is the one that turns a batch loop into an outage. * `__eq__` builds a plain `dict` from each side's items before comparing, so every comparison allocates two dictionaries. * `keys()`, `items()` and `values()` return the generic view objects, whose membership tests route back into your Python-level methods rather than into the C hash table. In a microbenchmark on CPython 3.14, a dict-backed subclass against the raw `dict` measured about 3x slower on membership hits, about 6x on misses, and about 5x on `update`. The multipliers are modest per call and lethal per million. ## What to override The mixins are ordinary methods, not sealed machinery — overriding them is the intended escape hatch, and the base class still enforces the five abstracts. Order the work by profile, and typically: 1. `__contains__`, `get` — delegate to the store; this removes the exception path entirely. 2. `update` — delegate to the store's own bulk update, but only if you have no per-key normalisation in `__setitem__` that must still run. If you do, that is the whole reason you built the class, and you keep the loop. 3. `clear` — delegate; this converts quadratic to linear and is usually the single biggest win. 4. `keys`, `items`, `values` — return the store's own views when your keys are stored verbatim. 5. `__eq__` — only if you actually compare mappings in a loop. The trap in every one of those is semantics. If `__setitem__` normalises keys, a `update` that bypasses it silently stores un-normalised keys, and an off-by-one at the batch boundary that flushes a partial buffer twice then writes rows under two spellings of the same column. Any override must preserve exactly what the mixin promised, including which exception it raises; the test for it is the mapping surface, not the fast path. ## When the base class is the wrong tool If you end up overriding almost everything, the base class is contributing an extra frame and no behaviour. At that point wrap a `dict` in a plain class and expose only the operations your callers really use, or store the `dict` directly and put the custom behaviour where it is actually needed. Keep the base class when the interface promise matters — when arbitrary code will be handed your object and must be able to treat it as a mapping — and when the hot path is short enough to override deliberately. ## Diagnosing it A deterministic profiler points straight at the standard library's mapping module, which is the tell: your class is spending its time in code you inherited. Confirm by timing the same operation against a raw `dict` with the same contents; the ratio, not the absolute number, tells you whether the interface layer or the storage is the problem.
- Why is the inherited clear() quadratic on a dict-backed mapping?`clear` calls `popitem` until it raises `KeyError`, and `popitem` does `next(iter(self))` — a brand-new iterator each time. Deleting from the front of the backing store leaves emptied slots the next iterator must scan past before it reaches a live key, so the scan grows with each removal. Timed across 10k, 20k and 40k keys the cost quadruples per doubling; overriding `clear` to delegate makes it linear.
- Which inherited methods are safe to leave alone?The ones off the hot path. `popitem`, `setdefault` and `pop` are fine when they run once per request rather than once per row, and `__eq__` is fine if you never compare mappings in a loop. The rule is a profile, not a reflex: each override is code you now have to keep semantically identical to the one you replaced.
- Why can overriding update() be more dangerous than overriding get()?Because `update` is one of the mixins that routes through `__setitem__`, and `__setitem__` is usually where the class earns its keep — key normalisation, validation, write-through. Delegating `update` straight to the backing store bypasses all of that, so keys arrive un-normalised and reads that go through `__getitem__` stop finding them. Override it only when `__setitem__` is a plain assignment.
saying these in an interview costs you the question
- Claims a MutableMapping subclass performs the same as a dict
- Thinks the inherited mixins are implemented in C
- Blames the abstract base class registration checks for the per-call cost
- Says clear() must be linear because each deletion is O(1)
- Rewrites the five abstract methods instead of overriding the hot mixins
- Overrides update() while __setitem__ still normalises keys