A nightly ETL export runs under python -O in production but not in CI; what breaks?
answer
- The tested artefact is not the shipped one
- A flag can arrive from an image, not the code
- Deleted statements take their side effects along
- Guards that only ever ran in CI
- Align the level; keep assert expressions pure
basics
~20 sProduction executes different bytecode from the one tested: every assert and if debug block is gone, so invariant checks never fire and any side effect inside an assert never happens. Bugs surface only in production, unreproducibly.
solid answer
~50 sThe two environments are not running the same program. At level 1 the compiler emits nothing for `assert` statements or `if __debug__:` blocks, so an invariant that fails loudly in CI passes silently in production — and if the assert expression *did* something, such as claiming a shared cursor or advancing a manifest, that work simply does not happen. That is how an export that concurrent workers race on shared state can corrupt a warehouse load only in production: the guard that serialised or validated it existed only in the tested build. The fix has three parts. Make the levels identical everywhere — the cheapest move is dropping `-O`, because it buys no speed. Promote anything that must hold in production out of `assert` into an ordinary check that raises. And log `sys.flags.optimize` at startup so the level of a running process is observable rather than inferred from a deploy manifest.
code
python · 11 linesclaimed = []
def claim(partition):
claimed.append(partition)
return True
def export(partition):
assert claim(partition), "partition already claimed"
return f"exported {partition}"
print(export("day=2026-09-05"), claimed)go deeper
Recall that -O deletes assert statements at compile time, so any check written as an assertion is simply absent in a production process started with that flag.
Explain the concrete mechanism and its trap: the statement is not evaluated at all, so a call inside the assert expression never runs and the two environments take different code paths.
Show the diagnosis and the fix: log sys.flags.optimize at startup, reproduce with the same flag, audit assertions with side effects, and align CI with production or drop the flag entirely.
Treat the optimization level as a platform property alongside the interpreter version — set once for the whole service graph, verified in a smoke check, and never varied per service without a stated reason.
This is a configuration-divergence bug, and it belongs to a family: production runs an artefact that no test ever executed. What makes the optimization level a particularly nasty member of that family is that the difference is invisible in the source, invisible in the process listing if the level came from `PYTHONOPTIMIZE`, and produces no error of its own. ## What actually differs At level 1 the compiler emits no bytecode for `assert` statements and drops `if __debug__:` blocks as dead code. Three consequences follow. **Invariants stop being enforced.** The check that made CI fail fast on a malformed batch is absent in production, so a bad batch flows onward and fails much later, somewhere with less context — a constraint violation in the warehouse, or worse, silently wrong data. **Side effects inside assert expressions disappear.** This is the one that produces genuinely different control flow. `assert claim_partition(worker_id)` acquires the claim under test and does nothing in production. Two export workers that were serialised in CI now overlap, and the race on shared state that the assertion was quietly preventing becomes real. The bug is not "an assertion was skipped"; it is "a function was never called". **Expensive debug work vanishes.** That part is intended, and it is the only legitimate reason the flag exists. ## Why it is hard to catch The level can be set far from the code. A base image, a Kubernetes manifest or a job template can export `PYTHONOPTIMIZE`, and nothing in the repository records it. In a fan-out pipeline — say a 17-service dependency graph where the export job pulls a shared transform library that eleven other services also import — a single service whose image sets the variable makes that shared library behave differently in exactly one place. Reviewers reading the library see assertions; one deployment does not run them. The reproduction step fails too. An engineer runs the failing job locally, without the flag, and it behaves correctly, which sends the investigation toward data, timing or the warehouse rather than toward the interpreter's startup configuration. ## How to diagnose it Start by making the level observable. Log `sys.flags.optimize` and `sys.version` once at startup for every service; a one-line addition turns the whole class of question into a query. If you suspect it retroactively, reproduce the job with the same flag rather than the same code — the difference in behaviour under `-O` versus level 0 is the confirmation. Then audit the assertions. The dangerous ones are those whose expression calls something. A short script that walks the abstract syntax tree with the `ast` module and reports every `assert` whose test contains a call is enough to find them across a repository, and a third-party linter will flag several of the same shapes. Fix each one by lifting the call onto its own statement and asserting on the result — or, more often, by deciding it was never an assertion in the first place. ## How to prevent it **One level everywhere.** Whatever level production uses, CI must use it too, including for the integration jobs. If they cannot be aligned, do not use the flag. **Prefer level 0.** `-O` provides no speed; the only real payoff on the ladder is the memory that `-OO` reclaims by dropping docstrings. If nobody can name the benefit being bought, the flag is pure risk. **Reserve `assert` for internal invariants.** An assertion says "this is impossible if the program is correct", and its cost of failure is a crash during development. Anything a correct program must still check at runtime — a partition claim, an argument that came from outside, a precondition that protects downstream data — needs an ordinary `if` and a raised exception. **Keep assert expressions pure.** Read state, compare it, produce a bool. If deleting the line would change what the program does, it is not an assertion. **Make it a platform rule, not a per-service one.** The level is a property of the runtime image and belongs in the same tier as the interpreter version: set once, verified in a smoke check, and identical across the graph. ## What the interviewer is listening for The weak answer is "assertions are disabled in production". The strong answer adds that the statement is *removed at compile time*, that removal takes its side effects with it, that the level can arrive from the environment without touching the code, and that the remedy is alignment plus moving real checks out of `assert` — not adding a try/except around an `AssertionError` that will never be raised.
- If assertions can be compiled away, is assert ever the right tool?Yes, for internal invariants: statements that are true whenever the program is correct, checked so that a broken assumption fails loudly during development. Their contract is that deleting them changes nothing but the speed of failure. The moment a check protects against untrusted input, guards downstream data, or must run in production regardless of flags, it is not an assertion and needs a real raised exception.
- How would you audit a large repository for assertions that carry side effects?Parse each file with the `ast` module and report every `Assert` node whose test contains a call, an assignment expression, or an attribute mutation — a few dozen lines, and it covers code no test executes. A third-party linter catches several of the same shapes as a rule. Then triage: lift the call onto its own line and assert the result, or convert the whole check into a raised exception.
- Would running CI under -O have caught this on its own?It would have caught the divergence, but not necessarily the bug — with the assertions gone, CI simply stops checking those invariants too, so a data defect can pass quietly in both places. The right combination is CI running at the production level for the integration suite, plus checks that genuinely matter written as ordinary raised exceptions so they survive every level.
It is like shipping a bridge whose load sensors were only wired up in the test rig: every measurement you trust was taken on a different structure from the one carrying traffic.
saying these in an interview costs you the question
- Says assertions are only skipped, not removed at compile time
- Wraps the call in try/except AssertionError to fix it
- Adds -O to production expecting a speed improvement
- Leaves side effects inside assert expressions
- Assumes the flag must be in the repo, not in the image
- Relies on assert to enforce a production invariant