skip to content

Why can an `is` comparison of parsed codes pass every test yet silently drop rows in production?

level: seniorimportance: should knowfreq 40%

answer

  1. Fixtures and real data are not the same shape
  2. Literals in tests get shared objects
  3. Parsed values are freshly allocated
  4. The compiler already warned about the comparison
  5. A broad handler turned a bug into missing rows

basics

~20 s

Test fixtures use small literals that CPython shares — cached integers up to 256 and interned identifier-like strings — so identity holds. Real data arrives as run-time strings and larger integers, which are distinct objects, so every identity check quietly returns False.

solid answer

~50 s

The comparison never tested what the author thought. Fixtures written as literals in the test module get shared objects: integers in the cached -5..256 window, identifier-like string constants from the intern table, and equal constants deduplicated inside one compiled code object. Values parsed from files at run time get fresh objects, so `is` returns False on data that is equal. If the mismatch branch is wrapped in a broad `except` that skips the row, a nightly load finishes green with a shrinking row count and no stack trace — the worst failure shape, because it looks like upstream data loss. The tell is in the source: CPython emits a `SyntaxWarning` at compile time for `is` against a literal, naming the literal type and suggesting `==`. Fix by comparing values, turning that warning into an error in CI, and making the swallowing handler log and re-raise.

code

python · 6 lines
python
line = "HB,402,normal"
_, code_text, status = line.split(",")
code = int(code_text)

print(code == 402, code is int("402"))
print(status == "normal", status is "normal".upper().lower())

go deeper

for a junior

Take away the concrete lesson: comparing values with is can be True for small literals and False for equal data read at run time, so use == for values and keep is for None, True, False and sentinels.

for a middle

Explain all three reasons the tests passed — the cached small integers, interned identifier-like constants, and constants stored once per compiled code object — and name the compile-time warning that flags the literal form of the mistake.

for a senior

Show the diagnosis: reproduce on real input rather than fixtures, promote the warning to an error, log both operands with their identities at the reject branch, and treat the over-broad handler as the defect that let a one-line bug run for six hours.

for a principal

Own the systemic fix: warnings as build errors, a lint rule against identity comparisons on non-sentinel operands, a reject-rate threshold that fails a batch job, and a test policy that no comparison is validated only against in-module literals.

### The failure, in order A loader for clinical-lab results runs nightly for about six hours. Its row filter compares a parsed status code with `is` — either against a literal, as in `if code is 0`, or against a value it looked up. The test suite builds fixtures inline as literals and passes. In production the same comparison is False for values that are equal, the row takes the reject branch, a broad `except Exception: continue` around the row loop swallows the resulting error, and the job exits zero having written fewer rows than it read. Nobody sees a traceback. The symptom surfaces days later as missing results, and the first suspicion falls on the upstream feed. ### Why the tests passed Three independent CPython mechanisms conspire to make identity look reliable in a test module: 1. **The small-integer cache.** Integers from -5 through 256 inclusive are pre-allocated and shared, so a fixture status of 0 or 200 is the same object everywhere it is produced. A real code of 402 is not. 2. **String interning.** Identifier-like string constants — ASCII letters, digits and underscores — are interned by the compiler, so a fixture value of `"normal"` is one canonical object. A value read from a file, decoded, split or stripped at run time is a fresh object even when the text matches. 3. **Per-code-object constant storage.** Within one compiled code object the compiler stores each distinct constant once, so two occurrences of the same literal in one test function are the identical object regardless of interning. Every one of those is an unspecified CPython implementation detail. Together they mean an identity check on literals is close to a tautology, and an identity check on parsed data is close to always False. ### The tell the interpreter already gave you Writing `is` or `is not` against a numeric, string or bytes literal produces a `SyntaxWarning` at compile time that names the literal's type and suggests `==`. Warnings are printed once per location by default and are easy to lose in a job's output, which is why they should be promoted: run the job with `PYTHONWARNINGS=error::SyntaxWarning`, or the equivalent `-W` command-line option, so the module fails to import rather than mis-comparing for six hours. Note the limit — the warning fires only for the literal form. `code is expected` where both sides are names compiles silently, so a static check that flags identity comparisons on non-sentinel operands, or a review rule, is still needed. ### The swallowed exception is the larger defect The identity bug is a one-line fix; the reason it survived to production is the handler. A bare or over-broad `except` around a per-row loop converts every defect — a comparison bug, a schema change, a decoding failure — into a silently smaller output. Two things make that safe: catch a specific exception type rather than everything, and make any skip observable — count it, log the row key at warning level, and fail the run when the skip rate crosses a threshold. A six-hour job that quietly discards rows is worse than one that crashes at minute three, because the crash is discovered immediately. ### Diagnosing it after the fact Given a shrinking row count and no traceback, the sequence is: reproduce with real input rather than fixtures, since the whole bug is that fixtures behave differently; re-run with `SyntaxWarning` promoted to an error; instrument the reject branch to log the compared values and their `id()`s, which shows two equal values with different identities immediately; and diff the counters between read, accepted and rejected to prove where rows are lost. Comparing `id()` values is a legitimate diagnostic here even though identity is never legitimate program logic. ### The durable rules Compare values with `==` and reserve `is` for `None`, `True`, `False` and unique sentinel objects created for that purpose. Never let test fixtures be the only data shape a comparison sees; parse a real sample in at least one test. Treat compile-time warnings as build failures in CI. And keep exception handlers narrow enough that a logic bug announces itself on the first row rather than after the run.

  • The suspect comparison is `code is expected`, with names on both sides. Does the compile-time warning help?
    No. The `SyntaxWarning` fires only when one operand is a literal, so a name-to-name identity comparison compiles silently. That case needs a static check that flags `is` on operands that are not `None`, `True`, `False` or a known sentinel, plus the habit of reading identity comparisons as suspicious in review.
  • How would you prove the mismatch at run time rather than by reading the code?
    Log both operands with their `id()` at the reject branch. Two equal values with different `id()`s is conclusive: `==` would have accepted the row and `is` did not. Doing it on real input matters, because fixtures made of literals produce shared objects and reproduce nothing.
  • What change would have made this a five-minute incident instead of a multi-day one?
    Making the skip observable: catch a specific exception, count rejects, log the row key, and fail the run when the reject rate exceeds a small threshold. The identity bug would then have surfaced on the first batch with an explicit reason rather than as a quietly short output days later.
  • Is interning ever a legitimate part of the fix here?
    Only as a performance measure, never as correctness. Interning the repeated status values with `sys.intern` at parse time saves memory across a large load and speeds dictionary lookups, but the comparison must still be `==`. Making identity hold on purpose so `is` works would rebuild the same trap on a different foundation.

saying these in an interview costs you the question

  • Blames the upstream feed rather than the comparison
  • Says `is` works for strings because Python interns them
  • Suggests calling `sys.intern` so `is` becomes reliable
  • Keeps the broad handler and only fixes the comparison
  • Claims tests were adequate because they exercised the branch
  • Assumes an exit code of zero means the run was correct

context