skip to content

Which dunder methods make a custom class work with len(), obj[i] and `in`?

level: juniorimportance: must knowfreq 60%

answer

  1. Built-ins ask the type, not a base class
  2. Three separate hooks: size, subscript, membership
  3. len() has a dedicated dunder, not an attribute
  4. One method serves index and slice
  5. Membership falls back to iteration

basics

~20 s

len answers len(obj), getitem answers obj[i] and slicing, and contains answers x in obj. Defining them on any class makes the built-ins work, because Python looks for the methods, not for a particular base class.

solid answer

~40 s

Container behaviour in Python is protocol-based: the built-ins call special methods on the object's type, so any class that defines them takes part. `len(obj)` calls `__len__`, which must return a non-negative `int`. `obj[key]` calls `__getitem__(key)`, and that same method receives a `slice` object for `obj[a:b]`; `obj[key] = value` calls `__setitem__` and `del obj[key]` calls `__delitem__`. `x in obj` calls `__contains__` when it exists; otherwise Python iterates the object via `__iter__`, and failing that falls back to the legacy protocol of calling `__getitem__` with 0, 1, 2 and so on until `IndexError`. That last fallback is why a class defining only `__getitem__` is already iterable. None of this requires inheriting from `list` or `dict`, which leaves you free to store the data however you like.

code

python · 16 lines
python
class Postings:
    def __init__(self, doc_ids):
        self._ids = list(doc_ids)

    def __len__(self):
        return len(self._ids)

    def __getitem__(self, index):
        return self._ids[index]

    def __contains__(self, doc_id):
        return doc_id in self._ids


p = Postings([4, 9, 15])
print(len(p), p[-1], 9 in p, list(p))

go deeper

for a junior

Recall the mapping: len() to len, square brackets to getitem, in to contains. Be ready to write a five-line class that wraps a list and supports all three without inheriting from list.

for a middle

Explain the mechanics: the lookup happens on the type, the same getitem serves both an index and a slice, and membership falls back to iter and then to the legacy index protocol before it gives up with TypeError.

for a senior

Show judgement about which hooks to implement. Defining contains turns a linear scan into whatever your storage can do cheaply, and the error type you raise, IndexError versus KeyError, decides whether the rest of the language behaves correctly around your object.

for a principal

Own the interface question: a protocol-implementing wrapper keeps storage private and the surface deliberate, whereas subclassing a built-in inherits behaviour that teams then spend years overriding. Decide which container operations a shared type should promise at all.

## Protocols, not base classes Python's built-in operations are defined in terms of *protocols*: named special methods ("dunders") that the interpreter looks up on an object's **type** and calls on your behalf. `len()`, subscription with square brackets, and the `in` operator are three separate protocols, and a class opts into each one independently by defining the matching method. There is no `Container` base class you are obliged to inherit from; the built-ins never ask "is this a list?", they ask "does this type define the method I need?". That is what people mean when they call Python duck-typed, and it is the reason a class holding its data in a dict, a file, or a remote service can still look and behave like a sequence. ## The size hook: `__len__` `len(obj)` calls `type(obj).__len__(obj)`. The return value must be a non-negative integer: returning a negative number raises `ValueError`, and returning something that is not an integer raises `TypeError`. If the type has no `__len__`, `len()` raises `TypeError: object of type 'X' has no len()`. Two things commonly surprise newcomers. First, `len()` never reads an attribute such as `.length` or `.size` — only the dunder counts. Second, `__len__` does double duty: it is also the fallback used to decide truthiness, so an object with `__len__` returning 0 is falsy in an `if` statement. ## The subscription hooks: `__getitem__`, `__setitem__`, `__delitem__` `obj[key]` calls `__getitem__(key)`, `obj[key] = value` calls `__setitem__(key, value)`, and `del obj[key]` calls `__delitem__(key)`. The *key* is whatever was written between the brackets: an integer for `obj[3]`, a string for `obj['name']`, and a single `slice` object for `obj[1:5]`. Your implementation decides what keys are legal and what to raise when one is not. The convention matters, because other machinery depends on it: sequence-style classes raise `IndexError` for an out-of-range integer, mapping-style classes raise `KeyError` for a missing key. Getting that wrong quietly breaks the legacy iteration fallback described below, and breaks `dict.get`-style patterns built on `try/except KeyError`. ## The membership hook and its two fallbacks `x in obj` is the most layered of the three. Python first looks for `__contains__` and calls it, coercing the result to a bool. If there is no `__contains__`, it iterates the object using `__iter__` and compares each element, using identity first and equality second — the check is effectively `x is element or x == element`. If there is no `__iter__` either, Python drops to the old sequence protocol and calls `__getitem__` with 0, 1, 2 and upward until `IndexError` ends the scan. Only when none of those exist does `x in obj` raise `TypeError`. That chain has two practical consequences. Defining `__contains__` is an *optimisation and a semantics* hook: a set-backed class can answer membership in constant time instead of scanning, and a range-like class can answer arithmetically. And a class that defines only `__getitem__` already supports both `for` loops and `in` without you writing another line — convenient, but worth replacing with a real `__iter__` in production code, because the fallback is slower and only works for integer-keyed objects. ```python class Postings: def __init__(self, doc_ids): self._ids = list(doc_ids) def __len__(self): return len(self._ids) def __getitem__(self, index): return self._ids[index] ``` With just those three methods, `len(p)`, `p[0]`, `p[-1]`, `p[1:3]`, `for d in p`, `9 in p`, `list(p)` and tuple unpacking all work, because each of them is built on the protocols above. ## What it looks like when a hook is missing Each protocol fails on its own terms, and recognising the messages is half of debugging them. Calling `len()` on a type without `__len__` gives `TypeError: object of type 'X' has no len()`. Subscripting without `__getitem__` gives `TypeError: 'X' object is not subscriptable`, while assigning without `__setitem__` gives `TypeError: 'X' object does not support item assignment` — two distinct methods, two distinct errors. And `x in obj` on a type with none of the three membership routes gives `TypeError: argument of type 'X' is not a container or iterable`, whose wording names the fallback chain directly. Notice that these are all `TypeError`: the operation is not merely failing, the type does not implement the protocol at all. ## Why this is the idiomatic route The alternative — subclassing `list` or `dict` — inherits an implementation you may not want and drags in behaviour you then have to override. Implementing the protocol keeps your storage private and your surface deliberate: you expose exactly the operations that make sense for the object. It also composes with static typing and with the abstract base classes in `collections.abc`, which are written in terms of the very same dunders. The interview point is simply this: Python's containers are an interface described by special methods, and any class can implement that interface.

  • Does defining `__getitem__` alone really make instances iterable, and should you rely on it?
    Yes. With no `__iter__`, Python falls back to the legacy sequence protocol and calls `__getitem__` with 0, 1, 2 and upward until `IndexError` stops it. It only works for integer keys, and it silently loops forever if you raise `KeyError` instead of `IndexError` or never raise at all. Treat it as a compatibility path: write `__iter__` explicitly in real code so iteration is independent of indexing.
  • What happens if `__len__` returns a negative number or a non-integer?
    `len()` validates the result. A negative value raises `ValueError` with a message saying `__len__()` should return a value greater than or equal to zero, and a non-integer such as a string raises `TypeError`. The contract is a non-negative `int`, so a computed length must be clamped, not merely estimated.
  • When `in` falls back to iteration, how does it compare elements?
    Identity first, then equality: each element is accepted if `x is element or x == element`. That shortcut is why a NaN value can be found in a list that contains that very object, even though NaN never compares equal to itself. It also means `in` calls your `__eq__`, so an expensive or side-effecting equality method makes membership expensive.

It is like a power socket rather than a family tree: an appliance works because its plug has the right shape, not because it was made by the same manufacturer.

saying these in an interview costs you the question

  • Claiming you must subclass list or dict to be a container
  • Thinking len() reads a .length or .size attribute
  • Believing the in operator only works on built-in types
  • Saying __getitem__ gets separate start and stop arguments for a slice
  • Forgetting __len__ must return a non-negative int
  • Raising KeyError for an out-of-range sequence index

context