skip to content

Dunder Protocols

Protocols are just special methods the interpreter calls on your behalf: printing, equality and hashing, indexing and iteration, arithmetic, calling. Implementing the right ones makes a class a first-class citizen.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

22

Which dunder methods make a custom class work with len(), obj[i] and `in`?

level: juniorimportance: must knowfreq 60%

answer

  1. Built-ins ask the type, not a base class
  2. Three separate hooks: size, subscript, membership
  3. len() has a dedicated dunder, not an attribute
  4. One method serves index and slice
  5. Membership falls back to iteration

basics

~20 s

len answers len(obj), getitem answers obj[i] and slicing, and contains answers x in obj. Defining them on any class makes the built-ins work, because Python looks for the methods, not for a particular base class.

solid answer

~40 s

Container behaviour in Python is protocol-based: the built-ins call special methods on the object's type, so any class that defines them takes part. `len(obj)` calls `__len__`, which must return a non-negative `int`. `obj[key]` calls `__getitem__(key)`, and that same method receives a `slice` object for `obj[a:b]`; `obj[key] = value` calls `__setitem__` and `del obj[key]` calls `__delitem__`. `x in obj` calls `__contains__` when it exists; otherwise Python iterates the object via `__iter__`, and failing that falls back to the legacy protocol of calling `__getitem__` with 0, 1, 2 and so on until `IndexError`. That last fallback is why a class defining only `__getitem__` is already iterable. None of this requires inheriting from `list` or `dict`, which leaves you free to store the data however you like.

code

python · 16 lines
python
class Postings:
    def __init__(self, doc_ids):
        self._ids = list(doc_ids)

    def __len__(self):
        return len(self._ids)

    def __getitem__(self, index):
        return self._ids[index]

    def __contains__(self, doc_id):
        return doc_id in self._ids


p = Postings([4, 9, 15])
print(len(p), p[-1], 9 in p, list(p))

go deeper

for a junior

Recall the mapping: len() to len, square brackets to getitem, in to contains. Be ready to write a five-line class that wraps a list and supports all three without inheriting from list.

for a middle

Explain the mechanics: the lookup happens on the type, the same getitem serves both an index and a slice, and membership falls back to iter and then to the legacy index protocol before it gives up with TypeError.

for a senior

Show judgement about which hooks to implement. Defining contains turns a linear scan into whatever your storage can do cheaply, and the error type you raise, IndexError versus KeyError, decides whether the rest of the language behaves correctly around your object.

for a principal

Own the interface question: a protocol-implementing wrapper keeps storage private and the surface deliberate, whereas subclassing a built-in inherits behaviour that teams then spend years overriding. Decide which container operations a shared type should promise at all.

## Protocols, not base classes Python's built-in operations are defined in terms of *protocols*: named special methods ("dunders") that the interpreter looks up on an object's **type** and calls on your behalf. `len()`, subscription with square brackets, and the `in` operator are three separate protocols, and a class opts into each one independently by defining the matching method. There is no `Container` base class you are obliged to inherit from; the built-ins never ask "is this a list?", they ask "does this type define the method I need?". That is what people mean when they call Python duck-typed, and it is the reason a class holding its data in a dict, a file, or a remote service can still look and behave like a sequence. ## The size hook: `__len__` `len(obj)` calls `type(obj).__len__(obj)`. The return value must be a non-negative integer: returning a negative number raises `ValueError`, and returning something that is not an integer raises `TypeError`. If the type has no `__len__`, `len()` raises `TypeError: object of type 'X' has no len()`. Two things commonly surprise newcomers. First, `len()` never reads an attribute such as `.length` or `.size` — only the dunder counts. Second, `__len__` does double duty: it is also the fallback used to decide truthiness, so an object with `__len__` returning 0 is falsy in an `if` statement. ## The subscription hooks: `__getitem__`, `__setitem__`, `__delitem__` `obj[key]` calls `__getitem__(key)`, `obj[key] = value` calls `__setitem__(key, value)`, and `del obj[key]` calls `__delitem__(key)`. The *key* is whatever was written between the brackets: an integer for `obj[3]`, a string for `obj['name']`, and a single `slice` object for `obj[1:5]`. Your implementation decides what keys are legal and what to raise when one is not. The convention matters, because other machinery depends on it: sequence-style classes raise `IndexError` for an out-of-range integer, mapping-style classes raise `KeyError` for a missing key. Getting that wrong quietly breaks the legacy iteration fallback described below, and breaks `dict.get`-style patterns built on `try/except KeyError`. ## The membership hook and its two fallbacks `x in obj` is the most layered of the three. Python first looks for `__contains__` and calls it, coercing the result to a bool. If there is no `__contains__`, it iterates the object using `__iter__` and compares each element, using identity first and equality second — the check is effectively `x is element or x == element`. If there is no `__iter__` either, Python drops to the old sequence protocol and calls `__getitem__` with 0, 1, 2 and upward until `IndexError` ends the scan. Only when none of those exist does `x in obj` raise `TypeError`. That chain has two practical consequences. Defining `__contains__` is an *optimisation and a semantics* hook: a set-backed class can answer membership in constant time instead of scanning, and a range-like class can answer arithmetically. And a class that defines only `__getitem__` already supports both `for` loops and `in` without you writing another line — convenient, but worth replacing with a real `__iter__` in production code, because the fallback is slower and only works for integer-keyed objects. ```python class Postings: def __init__(self, doc_ids): self._ids = list(doc_ids) def __len__(self): return len(self._ids) def __getitem__(self, index): return self._ids[index] ``` With just those three methods, `len(p)`, `p[0]`, `p[-1]`, `p[1:3]`, `for d in p`, `9 in p`, `list(p)` and tuple unpacking all work, because each of them is built on the protocols above. ## What it looks like when a hook is missing Each protocol fails on its own terms, and recognising the messages is half of debugging them. Calling `len()` on a type without `__len__` gives `TypeError: object of type 'X' has no len()`. Subscripting without `__getitem__` gives `TypeError: 'X' object is not subscriptable`, while assigning without `__setitem__` gives `TypeError: 'X' object does not support item assignment` — two distinct methods, two distinct errors. And `x in obj` on a type with none of the three membership routes gives `TypeError: argument of type 'X' is not a container or iterable`, whose wording names the fallback chain directly. Notice that these are all `TypeError`: the operation is not merely failing, the type does not implement the protocol at all. ## Why this is the idiomatic route The alternative — subclassing `list` or `dict` — inherits an implementation you may not want and drags in behaviour you then have to override. Implementing the protocol keeps your storage private and your surface deliberate: you expose exactly the operations that make sense for the object. It also composes with static typing and with the abstract base classes in `collections.abc`, which are written in terms of the very same dunders. The interview point is simply this: Python's containers are an interface described by special methods, and any class can implement that interface.

  • Does defining `__getitem__` alone really make instances iterable, and should you rely on it?
    Yes. With no `__iter__`, Python falls back to the legacy sequence protocol and calls `__getitem__` with 0, 1, 2 and upward until `IndexError` stops it. It only works for integer keys, and it silently loops forever if you raise `KeyError` instead of `IndexError` or never raise at all. Treat it as a compatibility path: write `__iter__` explicitly in real code so iteration is independent of indexing.
  • What happens if `__len__` returns a negative number or a non-integer?
    `len()` validates the result. A negative value raises `ValueError` with a message saying `__len__()` should return a value greater than or equal to zero, and a non-integer such as a string raises `TypeError`. The contract is a non-negative `int`, so a computed length must be clamped, not merely estimated.
  • When `in` falls back to iteration, how does it compare elements?
    Identity first, then equality: each element is accepted if `x is element or x == element`. That shortcut is why a NaN value can be found in a list that contains that very object, even though NaN never compares equal to itself. It also means `in` calls your `__eq__`, so an expensive or side-effecting equality method makes membership expensive.

It is like a power socket rather than a family tree: an appliance works because its plug has the right shape, not because it was made by the same manufacturer.

saying these in an interview costs you the question

  • Claiming you must subclass list or dict to be a container
  • Thinking len() reads a .length or .size attribute
  • Believing the in operator only works on built-in types
  • Saying __getitem__ gets separate start and stop arguments for a slice
  • Forgetting __len__ must return a non-negative int
  • Raising KeyError for an out-of-range sequence index

context

open as a page

What is the difference between `__repr__` and `__str__` on a Python class?

level: juniorimportance: must knowfreq 78%

basics

~10 s

__repr__ returns the unambiguous developer-facing view, used by repr(), the REPL and debuggers. __str__ returns the readable user-facing view used by print() and str(). A missing __str__ falls back to __repr__, never the reverse.

open as a page

What contract must `__hash__` satisfy relative to `__eq__` in Python?

level: middleimportance: must knowfreq 64%

basics

~10 s

Objects that compare equal must return equal hash values, and an object's hash must not change while it lives. The converse is not required: equal hashes are a collision, resolved by ==.

open as a page

Why does defining `__eq__` make a class's instances unusable as dict keys?

level: middleimportance: must knowfreq 72%

basics

~10 s

When a class body defines eq without hash, Python sets hash to None in that class, so hash() raises TypeError: unhashable type. Define hash over the same fields to restore it.

open as a page

Why does assigning `obj.__len__ = lambda: 5` leave `len(obj)` raising TypeError?

level: middleimportance: must knowfreq 48%

basics

~20 s

Because len() resolves len on type(obj), not on the object. The lambda lands in the instance dictionary, which implicit special method lookup never reads, so the class still has no len and len() raises TypeError.

open as a page

What does `__iadd__` control when Python evaluates `x += y`?

level: middleimportance: must knowfreq 60%

basics

~10 s

x += y calls type(x).__iadd__(x, y) and then rebinds the name x to whatever that returns. In-place methods normally mutate the object and return self; returning None silently rebinds x to None.

open as a page

When a class's __add__ returns NotImplemented, what does Python do next?

level: middleimportance: must knowfreq 60%

basics

~10 s

Python reads the return as a decline and calls the right operand's reflected method, radd, with the operands swapped. If that also returns NotImplemented, the interpreter raises TypeError: unsupported operand type(s) for +.

open as a page

With no `__eq__` defined, how does `==` behave on class instances?

level: juniorimportance: should knowfreq 58%

basics

~10 s

A class with no eq inherits object.eq, which compares identity: x == y is True only when both names refer to the very same object. Two instances built from identical data are still unequal.

open as a page

When `str(obj)` runs, does Python look up `__str__` on the instance or on its type?

level: juniorimportance: should knowfreq 38%

basics

~10 s

Python finds special methods on the object's type, never on the object itself. str(obj) calls type(obj).str(obj), so a str sitting in the instance dictionary is ignored. Special methods belong in the class body.

open as a page

What does defining `__call__` on a Python class let you do with its instances?

level: juniorimportance: should knowfreq 45%

basics

~10 s

It makes the instances callable: writing obj(args) runs type(obj).__call__(obj, args). The object behaves like a function while keeping its own state in ordinary attributes, and callable(obj) reports True.

open as a page

What does Python's NotImplemented mean, and how is it different from NotImplementedError?

level: juniorimportance: should knowfreq 40%

basics

~20 s

NotImplemented is a singleton value that a binary operator dunder returns to decline the other operand, letting Python try the reflected method. NotImplementedError is an exception raised from an abstract or unfinished method. Return one; raise the other.

open as a page

Why must a custom sequence's `__getitem__` handle negative indexes itself?

level: middleimportance: should knowfreq 38%

basics

~20 s

Python passes the subscript straight through, so a custom class receives -1 unchanged. Only built-in types like list and str wrap negative values around internally. Your method must add the length itself, or the lookup fails.

open as a page

What does `__getitem__` receive when a custom object is sliced as obj[2:9:2]?

level: middleimportance: should knowfreq 42%

basics

~20 s

A single slice object, slice(2, 9, 2), with start, stop and step attributes that may each be None. getitem always takes exactly one key argument, so a slice arrives as one object rather than as separate bounds.

open as a page

What does the `functools.total_ordering` class decorator generate?

level: middleimportance: should knowfreq 38%

basics

~10 s

It fills in the rich-comparison methods a class is missing. Define __eq__ plus any one of __lt__, __le__, __gt__ or __ge__, and the decorator derives the other three ordering methods from that pair.

open as a page

Why does printing a list of your objects show `__repr__` output instead of `__str__`?

level: middleimportance: should knowfreq 58%

basics

~20 s

print(items) calls str() on the list, and a list builds its own text by calling repr() on every element. Containers never call __str__ on their items, so an element's __str__ is skipped one level down.

open as a page

Which method does an f-string call for `{obj}`, `{obj!r}` and `{obj:>10}`?

level: middleimportance: should knowfreq 44%

basics

~10 s

{obj} and {obj:>10} both call type(obj).__format__(obj, spec), with an empty spec in the first case. {obj!r} applies repr() first and formats that plain string, so a custom __format__ is bypassed.

open as a page

A lazy result wrapper defines `__len__`; why does `if wrapper:` force a full count, and how do you avoid it?

level: seniorimportance: should knowfreq 36%

basics

~10 s

Truth testing calls bool first, and with none defined it falls back to len, comparing the count against zero. On a lazy wrapper every if statement therefore materialises everything. Define bool, or drop len.

open as a page

Why does mutating an object already used as a `dict` key make later lookups miss it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The container placed the key using the hash it had at insertion. Mutating a hashed field changes the hash, so lookups search the wrong place: the key tests as absent while iteration still yields it.

open as a page

Why does a proxy that forwards attributes via `__getattr__` still raise TypeError on `with proxy:`?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Implicit invocations skip getattribute and getattr completely: the with statement resolves enter and exit on type(proxy). The forwarding hook is never consulted, so a proxy must declare on its own class every special method it wants to support.

open as a page

When does overloading a Python operator hurt more than a named method?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Overload only when the operator has one obvious meaning for the type, such as arithmetic on a value object. If a reader must look up what + does, or the operator hides mutation of shared state, a named method is clearer.

open as a page

How do you write a `__repr__` that stays useful in production logs and debuggers?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Show the class name and the few fields that identify the instance, keep it cheap, side-effect free and unable to raise, redact secrets, and bound large fields while printing the real size next to any truncated sample.

open as a page

When does Python try the right operand's __radd__ before the left operand's __add__?

level: seniorimportance: nice to knowfreq 18%

basics

~20 s

When the right operand's type is a proper subclass of the left operand's type and actually overrides the reflected method. Python gives the more specific type first refusal, because the base class was written before the subclass existed.

open as a page