skip to content

A lazy result wrapper defines `__len__`; why does `if wrapper:` force a full count, and how do you avoid it?

level: seniorimportance: should knowfreq 36%

answer

  1. Truth testing has an ordered fallback chain
  2. One method stands in for another
  3. Objects defining neither are always true
  4. Emptiness is cheaper than counting
  5. One item answers the real question

basics

~10 s

Truth testing calls bool first, and with none defined it falls back to len, comparing the count against zero. On a lazy wrapper every if statement therefore materialises everything. Define bool, or drop len.

solid answer

~50 s

Python decides truthiness in three steps: it calls `__bool__` if the type defines one, otherwise it calls `__len__` and treats a non-zero length as true, and if neither exists the object is always true. So adding `__len__` to a lazily-evaluated wrapper silently makes `if wrapper:` a materialising operation — on a search-index rebuild streaming a 2.4 GB working set, one innocuous truth test pulls the entire batch into memory. There are two clean fixes. Define `__bool__` explicitly so it answers cheaply, typically by pulling a single element and buffering it for later iteration. Or refuse to define `__len__` when the size is unknown or expensive, and make callers say what they mean with `if wrapper is not None:` or `next(iter(wrapper), None)`. `__len__` is a contract for a cheap, exact, non-negative count — if you cannot honour that, do not offer it.

code

python · 20 lines
python
class Sized:
    def __len__(self):
        return 0


class Plain:
    pass


class Odd:
    def __bool__(self):
        return 1


print(bool(Sized()), bool(Plain()))

try:
    bool(Odd())
except TypeError as exc:
    print(type(exc).__name__, exc, sep=": ")

go deeper

for a junior

Recall the order: bool first, then len, and otherwise true. Know that an empty list is falsy only because its length is zero, and that a plain object with no such methods is always truthy.

for a middle

Explain that len silently doubles as the truthiness implementation, and that bool must return a real bool or raise TypeError. Be able to say what bool() returns for a class defining neither method.

for a senior

Show the production judgement: truthiness is an implicit control-flow path, so anything reachable from it must be cheap and exact. Diagnose the materialising if-guard on a lazy wrapper and offer both fixes, a peeking bool or omitting len entirely.

for a principal

Own the interface promise. Exposing len commits a shared type to a cheap exact count forever, and reviewers should treat if x: versus if x is not None: as a semantic decision about how an empty result must behave across every consumer.

## The truth-value protocol Every object in Python has a truth value, and `if obj:`, `while obj:`, `not obj`, `and`, `or` and `bool(obj)` all resolve it the same way, in a fixed order: 1. If `type(obj)` defines `__bool__`, call it. The result must be an actual `bool`; returning `1` or a truthy string raises `TypeError: __bool__ should return bool`. 2. Otherwise, if the type defines `__len__`, call it. Zero means false, any non-zero length means true. 3. Otherwise the object is true. This is why a plain `object()` and an instance of a class with no protocol methods are always truthy. That ordering explains the familiar built-in behaviour: `[]`, `''`, `{}` and `set()` are falsy purely because their length is zero, while `None` and `False` are falsy through their own types. ## Why `__len__` is the trap on a lazy object Step 2 is where the design bug lives. `__len__` looks like a harmless convenience — callers get `len(wrapper)` for free — but it silently doubles as the truthiness implementation. On an eager, in-memory container that costs nothing. On a wrapper that produces items lazily from a stream, the only way to answer "how many" is to consume everything. Consider a search-index rebuilder that streams postings out of a source and processes them in batches, with a 2.4 GB working set when a batch is fully materialised. Someone writes the perfectly ordinary guard: ```python if batch: index.write(batch) ``` There is no `__bool__`, so Python calls `__len__`, which drains the source into a list to count it. Memory that the streaming design existed to avoid is now resident, and if the wrapper's iteration is one-shot, `index.write(batch)` may then see nothing at all. The guard that was supposed to be free is the most expensive line in the loop. A related failure appears when someone tries to make `__len__` cheap by *estimating* it — for example scaling a total by a sampled ratio, `int(total * ratio)`. Floating-point rounding drift means the estimate can land on zero for a stream that genuinely has items, and the object becomes falsy while holding data. The batch is skipped silently, no exception is raised, and the missing documents surface much later as a gap in the index. `__len__` is not a hint; it is a promise of an exact non-negative integer, and an estimate must never be published through it. ## Fix one: implement `__bool__` cheaply Emptiness is a *much* weaker question than size, and it is usually answerable by looking at one element. Pull a single item, keep it in a small buffer, and chain the buffer back in front of the source when iteration starts: ```python import itertools class LazyPostings: def __init__(self, source): self._source = iter(source) self._head = [] def __bool__(self): if not self._head: self._head = list(itertools.islice(self._source, 1)) return bool(self._head) def __iter__(self): return itertools.chain(self._head, self._source) ``` Now `if wrapper:` costs one item, `__bool__` is idempotent, and nothing is lost from the stream. Note that `__bool__` short-circuits step 2 entirely, so a class may legitimately have both a cheap `__bool__` and an expensive `__len__` and still be safe in an `if`. ## Fix two: do not offer `__len__` at all If the length is unknown, unbounded or expensive, the honest interface omits it. `len(wrapper)` then raises `TypeError` — loudly, at the call site, where someone can decide what they actually meant. Callers testing for presence use `next(iter(wrapper), None) is not None`; callers testing that a value was returned at all use `if wrapper is not None:`. That last distinction is worth stating in a review: `if wrapper:` and `if wrapper is not None:` differ exactly on the empty case, and a great many bugs are a truthiness test written where an identity test was intended. ## The contracts to keep straight `__bool__` must return a real `bool` and should be cheap, since it is invoked implicitly by control flow. `__len__` must return an exact, non-negative `int`; a negative value raises `ValueError` and a non-integer raises `TypeError`. And an object that defines neither is unconditionally true — so a wrapper class that people expect to be falsy when empty will not be, unless one of the two methods says so. The senior judgement being tested is recognising that truthiness is an implicit, hot, control-flow path, and that whatever you attach to it must be cheap and exact.

  • What is the truth value of an instance whose class defines neither __bool__ nor __len__?
    True, always. Python only overrides the default when one of the two methods says so, which is why a plain `object()` is truthy and why a custom collection that forgets both is never falsy when empty. If emptiness should be meaningful for your type, you must implement one of them deliberately.
  • A wrapper computes __len__ as int(total * ratio) so it can stay cheap. What breaks?
    Floating-point rounding can push the estimate to zero for a stream that actually contains items, which makes the object falsy and silently skips it in any `if` guard. `__len__` promises an exact non-negative count, not an approximation; an estimate belongs in a separately named method that no protocol invokes implicitly.
  • When should code prefer `if wrapper is not None:` over `if wrapper:`?
    Whenever the question is "did I get an object?" rather than "does it contain anything?". The two differ precisely on the empty case, and a truthiness test written where an identity test was meant will discard a legitimately empty result. The identity form also avoids invoking `__bool__` or `__len__`, so it never triggers work.

Asking whether a queue is empty should mean glancing at the front of the line, not counting every person in it before you decide to open a second till.

saying these in an interview costs you the question

  • Thinking every object is truthy unless it is None or False
  • Assuming __bool__ is required for truth testing to work
  • Defining __len__ on a stream of unknown or expensive length
  • Returning 1 rather than True from __bool__
  • Writing if wrapper: when the real question is is not None
  • Publishing an estimated count through __len__

context