Why does `sorted()` raise TypeError on a 6,800-row batch whose duration field is sometimes None?
answer
- Python 2 was permissive here
- Python 3 refuses to guess
- Only ordering raises, not equality
- The comparison has no defined meaning
- Give every row one comparable type
basics
~20 sPython 3 refuses to order unrelated types: comparing None with an int raises TypeError, because neither defines an ordering against the other. Equality never raises. Fix it with a key function mapping every row onto one comparable type.
solid answer
~50 sSorting compares elements pairwise, so a single `None` duration among the numeric ones is enough to attempt `None < 12`, and Python 3 raises `TypeError: '<' not supported between instances of 'NoneType' and 'int'`. Python 2 gave any two objects a consistent-but-arbitrary ordering; **Python 3.0 removed it**, and on 3.14 ordering unrelated types is an error. Note the asymmetry: `==` and `!=` never raise, falling back to identity, so `None == 12` is simply `False` - which is exactly why bad data survives every equality-based validation and only detonates at the sort. The failure is data-dependent, so a video-metadata extractor can pass every test batch and fail on the first production batch containing an unparsed duration. Fix it in the key function, not with a try/except: `key=lambda r: (r['duration'] is None, r['duration'] or 0)` gives a total order and parks the missing values at one end deterministically.
code
python · 11 linesrows = [{"title": "intro", "duration": 12}, {"title": "outro", "duration": None}]
try:
rows.sort(key=lambda r: r["duration"])
except TypeError as exc:
print("sort failed:", exc)
print(None == 12)
rows.sort(key=lambda r: (r["duration"] is None, r["duration"] or 0))
print([r["title"] for r in rows])go deeper
Recall that Python 3 raises TypeError when you order values of unrelated types, such as None against a number, and that a key= function is how you tell a sort what to compare. Reading the exception message, which names both types, is half the fix.
Explain the mechanism: the ordering methods decline, Python tries the reflected operation, and with nowhere left to go it raises - while equality falls back to identity and stays silent. Write the tuple key that gives a total order over rows with missing values.
Show production judgment: the failure is data-dependent, so it clears every fixture and lands on the first real batch. Diagnose at the ingest boundary rather than the sort, and reject try/except-and-retry, which converts an informative crash into unbounded memory growth.
Own the boundary policy: where untyped external data is normalised, whether a missing value is rejected or given a documented sentinel, and how ordering keys are reviewed so nobody 'fixes' a crash with a silently wrong lexicographic sort.
## What Python 3 actually promises about ordering Ordering comparisons are dispatched to the operands' types. When the left operand's type cannot order itself against the right operand's type, it signals that by returning `NotImplemented`; Python then tries the reflected operation on the right operand, and when that also declines, the interpreter raises `TypeError` with a message naming both types: ``` TypeError: '<' not supported between instances of 'NoneType' and 'int' ``` Python 2 behaved differently: it fell back to an arbitrary but *consistent* ordering, typically by type name, so a mixed list sorted without complaint into a meaningless order. **Python 3.0 deleted that fallback deliberately**, on the grounds that a silently meaningless sort is worse than a loud failure. On 3.14 this is unchanged, and it is the behaviour to assume in any interview answer that does not name a version. ## The asymmetry that hides the bad data Only *ordering* raises. Equality is defined for every pair of objects, falling back to identity when the types decline to compare, so `None == 12` evaluates to `False` and `None != 12` to `True`. Nothing raises. That asymmetry is why a `None` slips through an entire ingest pipeline - equality checks, membership tests, deduplication, `if value == expected` guards all tolerate it - and the first operation that *orders* the data is where it surfaces. In a metadata extractor the sort, the `min`/`max` over durations, a `heapq` push, or a `bisect` insertion are all candidates; whichever runs first takes the blame for a defect introduced much earlier. The failure is also data-dependent rather than deterministic across environments. A 6,800-row batch raises only if it actually contains a `None`; the fixture batches used in development may never have had one. Two rows are enough - any list of length two or more forces at least one comparison of every element - so there is no safe "small batch" threshold, only a lucky one. ## Diagnosing it The exception message names both offending types, which is most of the diagnosis. From there: * Confirm which field is heterogeneous by inspecting the distinct types present, for example with a `collections.Counter` over `type(row[field]).__name__`. * Check the ingest boundary, not the sort. Ask where the untyped value entered - an absent key defaulting to `None`, an empty string from a parser, a JSON `null`, a failed regex returning no match. * Reproduce with the smallest possible list. `sorted([12, None])` is the whole bug. Beware of "fixing" it by wrapping the sort in `try/except TypeError` and diverting failed batches to a retry list. The data does not change on retry, so nothing ever drains and the retry list grows without bound until the process is killed - swapping a loud, informative crash for a slow memory leak with no diagnostic in it. ## Making the ordering total The durable fix is to give every element one comparable type. In rough order of preference: 1. **Normalise at the boundary.** If a duration is genuinely unknown, decide at ingest what that means - reject the row, or store a documented sentinel of the right type - so nothing downstream has to know. 2. **Use a key function that produces a total order.** `key=lambda r: (r["duration"] is None, r["duration"] or 0)` sorts on a tuple whose first element is a bool: all present values first, missing ones last, and within each group the numbers order normally. Tuple comparison is lexicographic and short-circuits, so the second element is only compared between rows with the same missing-ness. 3. **Partition explicitly.** Filter the missing rows out, sort the rest, and concatenate. It is more code, but the placement of missing values is impossible to misread. Two tempting fixes are worse than they look. `key=str` never raises but orders lexicographically - `"100"` sorts before `"99"` - which is a *silent* wrong answer, the very thing Python 3 stopped doing for you. And a sentinel like `float("inf")` works for numbers but quietly claims that an unknown duration is the longest one. ## The same error wearing other clothes The pattern generalises well beyond `None`: numeric strings that were never parsed being ordered against numbers; a naive and an aware `datetime.datetime` compared, which raises for ordering while equality returns `False`; complex numbers, which support equality but no ordering at all. In every case the question to ask is the same - what single type is this sequence supposed to order on, and where did something else get in?
- Why does `None == 12` return False instead of raising, when `None < 12` raises?Equality is defined for every pair of objects: when the types decline to compare, Python falls back to identity, so unrelated objects are simply unequal. Ordering has no such fallback in Python 3 - if neither type knows how to order against the other, there is no meaningful answer to invent, so it raises. That asymmetry lets bad data survive equality-based validation.
- A teammate wraps the sort in try/except TypeError and pushes failed batches onto a retry list. What goes wrong?The data is unchanged on retry, so those batches fail forever. The retry list only grows, giving unbounded memory growth and eventually an out-of-memory kill - with the original, perfectly informative `TypeError` swallowed so nothing points at the real defect. Trading a loud crash for a silent leak is strictly worse; fix the ordering or reject the row at ingest.
- Why is `key=str` a poor fix for a numeric field with missing values?It never raises, which is exactly the problem: it silently orders lexicographically, so `'100'` sorts before `'99'` and `'None'` lands among the letters. You get a plausible-looking, wrong order with no error to investigate - reintroducing precisely the silently meaningless sort that Python 3 removed.
- Which other operations surface this same TypeError besides sorting?Anything that orders: `min`, `max`, `heapq` pushes, `bisect` insertions, a chained comparison such as `lo < value < hi`, and a plain `<` in a filter. They all raise on the same unordered pair. Equality-based operations - membership, deduplication via a set, `==` guards - accept the mixed data quietly, which is why the ordering operation is usually blamed for a defect introduced upstream.
Sorting mixed types is like ranking runners against swimmers: Python 3 refuses to invent a common yardstick, where Python 2 would happily hand you a meaningless league table.
saying these in an interview costs you the question
- Says Python sorts None as smaller than any number
- Claims `None == 12` also raises TypeError
- Wraps the sort in try/except instead of fixing the data
- Thinks the sort is unstable rather than type-incompatible
- Believes a type annotation prevents this at runtime
- Suggests `key=str` without noticing the lexicographic order