A telemetry collector's `__len__` scans a 2.4 GB buffer; why does `if collector:` trigger that scan?
answer
- The check is a conversion, not a peek
- A missing hook redirects the question
- Rewriting the call site changes nothing
- Answer a coarser question more cheaply
- An implicit call must not fail
basics
~10 sBecause the class defines no __bool__, truth testing falls back to __len__, so the emptiness check pays for an exact count it then compares with zero. Add a cheap __bool__ that answers emptiness directly.
solid answer
~40 s`if collector:` is `bool(collector)`: with no `__bool__` on the type, Python falls back to `__len__` and the full 2.4 GB scan runs just to learn whether the count is zero. Rewriting it as `if len(collector) > 0:` calls the same method, so the fix belongs on the class: define `__bool__` to answer the coarser question cheaply — `any(self._chunks)`, a counter maintained on append, or an in-memory head pointer — and leave `__len__` exact for callers that genuinely need the count. Every `if`, `while`, `not` and `and`/`or` in the codebase then takes the cheap path with no call sites edited. Keep implicitly invoked hooks unable to fail, too: an exception out of `__len__` propagates from the `if` line, and a surrounding broad `except` turns it into a silent "looks empty" that drops readings.
code
python · 15 linesclass TelemetryBuffer:
def __init__(self, chunks):
self._chunks = chunks
def __len__(self):
print("full scan")
return sum(len(chunk) for chunk in self._chunks)
def __bool__(self):
return any(self._chunks) # stops at the first non-empty chunk
buf = TelemetryBuffer([[], [1, 2, 3]])
print(bool(buf)) # True, no scan
print(len(buf)) # prints "full scan", then 3go deeper
Take away one rule: truth-testing an object can run your own code. If a class defines __len__, if obj: calls it, so an emptiness check is only as cheap as that method.
Explain the fallback order and why the fix goes on the class rather than the call site, since if len(x) > 0: calls the same method. Know that defining __bool__ short-circuits __len__ completely.
Demonstrate the production judgement: keep implicitly invoked hooks cheap and unable to raise, keep __len__ exact, and show how you would find the problem in a profile where the hot method has no visible caller.
Own the API contract: which implicit protocols a type should opt into at all. Weigh the ergonomics of if obj: against invisible cost and invisible failure, and set the house rule for when an explicit predicate beats a dunder.
### Why `if collector:` runs the scan at all Truth-testing an object is not free and is not special-cased. `if collector:` calls `bool(collector)`, which looks for `__bool__` on the type; the collector does not define one, so the interpreter falls back to `__len__` and treats the result as false only when it is `0`. Your `__len__` is the code that walks the 2.4 GB working set, so the emptiness check pays for an exact count it then throws away after comparing it with zero. A statement that reads like a null check is in fact a full traversal. The most common "fix" makes nothing better: `if len(collector) > 0:` calls the identical method. Nor does hoisting it out of the loop help if the loop is where emptiness changes. The cost is structural — the only cheap emptiness test is one the class provides. ### The fix: give the type a cheap `__bool__` `__bool__` short-circuits `__len__` entirely, so define it to answer the question actually being asked — *is there anything at all?* — with the cheapest evidence available: ```python class TelemetryBuffer: def __len__(self): # exact, expensive, rarely called return sum(len(chunk) for chunk in self._chunks) def __bool__(self): # cheap: stops at the first non-empty chunk return any(self._chunks) ``` `__len__` stays exact — `len()` is a promise and callers, `list()` and slicing logic depend on it — while every `if`, `while`, `not` and `and`/`or` in the codebase now takes the cheap path automatically, with no call sites edited. Other cheap sources of the same answer: a counter maintained on append and pop, a "first chunk" pointer, a `head is not None` check on a linked structure, or an EOF flag already tracked by the reader. If even that is not cheap, the honest move is to stop implying it is: drop `__len__` from the public shape, expose an explicit `has_readings()`, and let the object be plainly truthy so no implicit conversion can ambush you. When you want a rough size for pre-sizing rather than a truth value, `__length_hint__` — read via `operator.length_hint()` — is the hook that is *allowed* to be approximate. ### The second failure: the truth test can raise An expensive `__len__` is usually also a *fallible* one — it touches memory-mapped files, a spooled temp file, or a device handle. Whatever it raises propagates straight out of the `if` line, which is where the second half of this bug lives. Code shaped like ```python try: if collector: flush(collector) except Exception: pass # swallows the failure and the branch with it ``` turns a raised `__len__` into "the collector looked empty". Nothing is flushed, nothing is logged, and the readings are dropped silently — the classic swallowed exception, made worse because the implicit call is invisible in the source. Two disciplines contain it: never let a truth-conversion hook do work that can fail (a cheap `__bool__` reading an in-memory flag cannot raise), and never wrap a conditional in a bare `except Exception: pass`. If the emptiness check genuinely can fail, make the failure explicit with a method call that the reader can see, so the `try` block wraps something that looks like it might raise. ### Diagnosing it in a running service The symptom is a profile where a trivial-looking line dominates. Deterministic profiling (`cProfile`, or `sys.setprofile`-based tooling) attributes the time to `__len__` with a caller that contains no visible `len` call — that mismatch between the profile and the source is the tell. A sampling profiler shows the same stack. A cheap local check: add a `print` or a counter to `__len__` and run the code path once; if the count is far higher than the number of `len()` calls you can find by grepping, implicit truth tests are the difference. Because the scan touches the whole working set, the secondary damage is memory-behaviour rather than CPU alone — page faults and cache eviction across gigabytes, which is why the wall-clock cost is worse than the operation count suggests. ### What senior judgement sounds like here Say the mechanism first (`bool()` → `__bool__` → `__len__` → default true), then the fix at the type rather than at the call sites, then the invariant you are protecting: `__len__` stays exact, and `__bool__` may be an approximation only in the sense that it answers a coarser question. Add the operational point — a hook invoked implicitly must be cheap and must not raise, because callers cannot see it being called — and you have covered what the question is really testing.
- Does rewriting the check as `if len(collector) > 0:` avoid the cost?No — it calls exactly the same `__len__`, just explicitly. It is arguably better only in that the cost is now visible in the source. The cost is structural: an exact count is what the method computes, so the only way to make emptiness cheap is to give the type a `__bool__` that answers emptiness without counting.
- Why not just make `__len__` return a cheap approximate count instead?Because `len()` is a promise of exactness that callers, `list()` construction and slicing logic rely on, and an approximation there produces wrong results far from the change. If you want a cheap, explicitly approximate size for pre-sizing, that is what `__length_hint__` is for — read through `operator.length_hint()` — and it is allowed to be wrong.
- How would you spot this in a profile of a running service?The tell is a mismatch between the profile and the source: time attributed to `__len__` with callers that contain no visible `len` call. A sampling profiler shows the same stack. A quick local confirmation is a counter or a print inside `__len__` — if invocations far outnumber the `len()` calls you can grep for, implicit truth tests are the difference.
- What if the emptiness check genuinely cannot be made cheap or infallible?Then stop implying it is. Do not implement `__len__` or `__bool__` at all; expose an explicit method such as `has_readings()` and let instances be plainly truthy. An explicit call is visible at the call site, can be wrapped in a targeted `try`, and cannot be triggered by an innocuous-looking `if`, `not` or `and`.
Asking a warehouse "is there anything in there?" and getting back a full inventory count is technically an answer, but you paid for a stocktake to learn whether the lights should be on.
saying these in an interview costs you the question
- Claims `if obj:` is a cheap identity or null check
- Suggests `if len(obj) > 0:` as the performance fix
- Makes `__len__` approximate to speed up the check
- Thinks only `if` converts, not `not` or `and`
- Wraps the truth test in a bare `except Exception: pass`
- Adds caching without saying when it is invalidated