Why does a nightly ETL export print a DeprecationWarning only on its first batch?
answer
- One log line is not one call
- The action decides how often, not just whether
- Something remembers what it already showed
- Keyed by message, category and line
- Rerun with always, or record and count
basics
~20 sThe default warning action shows a message once per location — the text, category and reported line — and records that in a per-module registry, so later identical warnings from the same line are skipped entirely. Switch the filter to always to see every occurrence.
solid answer
~40 sNothing about the export changed between batches; the warning was suppressed. Under the `default` action, `warnings` records each warning it shows in a registry attached to the module it was reported from, keyed by the message text, the category and the line number, and it checks that registry **before** consulting the filters. Batch two hits the same key and returns silently. That is why one log line is not evidence of one call site — and, with a library warning at `stacklevel=1`, every call site in the job shares a single key. To measure the real blast radius, re-run the job with `-W always::DeprecationWarning` or wrap it in `warnings.catch_warnings(record=True)` with `warnings.simplefilter('always')` and count the recorded objects. Mutating the filter list invalidates those registries, which is exactly how `catch_warnings` gives each block a clean slate.
code
python · 9 linesimport warnings
def export_batch():
warnings.warn('row order is not guaranteed', DeprecationWarning)
warnings.simplefilter('default')
for _ in range(3):
export_batch()
print('done')go deeper
Remember that a warning can be printed once and then suppressed, so seeing it a single time does not mean it happened a single time. Re-running with warnings turned up is how you check.
Explain the actions — always, default, module, once, ignore, error — and that three of them record what they already showed in a registry keyed by message, category and reported line.
Show the measurement: re-run with -W always or record with catch_warnings, group by filename and line to find real call sites, and explain why escalating to errors belongs in CI rather than in the nightly job.
Own the reporting policy behind the numbers — which environments run with warnings opened up, where occurrences are aggregated, and how a migration is judged done when the counts are filtered artefacts.
### The symptom A nightly export builds batches of rows and pushes them to a warehouse. On the first batch the log carries one `DeprecationWarning` about a helper whose replacement no longer preserves row order; for the remaining several hundred batches there is nothing. The natural reading — "one call site, one occurrence, low risk" — is wrong, and the ordering assumption baked into the downstream tables is riding on that misreading. ### The once-per-location registry After the filter chain selects an action, three of the actions record what they showed so it is not shown again: - `default` — once per unique (message text, category, module, line number). - `module` — once per (message, category, module), regardless of line. - `once` — once per (message, category) for the whole process. The bookkeeping lives in a registry dictionary attached to the module the warning was *reported* from, and it is consulted **before** the filters are walked: if the key is present, `warn` returns immediately and does no further work. So the second and subsequent batches never even reach the filter list. `always` (and its `error` sibling) are the actions that skip this bookkeeping and fire every time. Two details make this sharper in production. First, the key includes the *reported* line, which `stacklevel` selects — so a library that warns with the default `stacklevel=1` gives every one of your call sites the same key, and exactly one of them is ever reported. Second, the registry is per process. A pipeline that forks or spawns a worker per batch starts each worker with an empty registry, so the same deprecation appears once per worker rather than once per run, and the count in your logs measures your process model rather than your code. ### Confirming it, not guessing The measurement you want is: how many distinct call sites in this job trigger this deprecation? Three ways to get it, in increasing intrusiveness: 1. **Re-run with the filter opened up**: `python -W always::DeprecationWarning export.py`, or `PYTHONWARNINGS=always::DeprecationWarning` in the job's environment. Every occurrence prints, with its own location. 2. **Record in process**: wrap the run in `warnings.catch_warnings(record=True)` with `warnings.simplefilter('always')`, then group the resulting `WarningMessage` objects by `.filename` and `.lineno`. This gives you a count and a list of files to fix, rather than a wall of stderr. 3. **Route them into the operational log**: `logging.captureWarnings(True)` redirects warnings into the logging system, so a staging run's occurrences land in the same place as everything else and can be aggregated. Combine it with an `always` filter, or you will aggregate the same suppression you started with. Be aware of the interaction: changing the filter list bumps an internal version counter, and any registry whose version is stale is cleared. That is a feature — it is how `catch_warnings` guarantees a clean slate on entry and restores the previous state on exit — but it also means a mid-run `simplefilter` call makes previously-suppressed warnings reappear, which looks like a new problem if you do not know why. ### Why not just make it fatal `-W error::DeprecationWarning` in the nightly job is the wrong lever. The export is measured against a 92nd-percentile completion budget; raising on the first occurrence aborts mid-run, leaves the warehouse holding a partial load, and buys nothing you could not get from a recording run in staging. Escalate to errors where a failure is cheap and someone is watching — CI, the test suite, a staging run — and keep production observing rather than dying. If the deprecated path genuinely must not run in production, that is a code change or a feature flag, not a warning filter. ### The judgment being tested The interviewer is checking whether you know that warning counts are *filtered* counts. A single log line is evidence about the filter configuration first and about the code second, and any migration plan built on "we only saw it once" is built on the `default` action's deduplication. The follow-through is to re-measure with `always`, fix the call sites the recording run names, and only then decide whether the deprecation still deserves escalation.
- What is the difference between the once and default warning actions?`default` shows a warning once per unique location — message, category, module and line — so ten distinct call sites produce ten reports. `once` shows it a single time for the whole process no matter how many places trigger it, keyed only by message and category. `module` sits between them, reporting once per module. For an audit you want `always`; for a quiet production default, `default` is the one CPython falls back to.
- Why can the same deprecation reappear after code has been running quietly for a while?Because the already-shown registries carry a version stamped from the filter list, and any mutation of `warnings.filters` — `filterwarnings`, `simplefilter`, entering or leaving `catch_warnings` — invalidates them. Code that adjusts filters at runtime, including a test helper or a plugin, resets the bookkeeping, so previously suppressed warnings print again. A new process also starts with empty registries.
- Is warnings.catch_warnings safe to use in a multi-threaded job?By default, no. It saves and restores process-wide state — the filter list, the show function and the registry version — so a block entered on one thread changes what every other thread sees, and concurrent blocks restore each other's saved state. Python 3.14 added an opt-in interpreter flag that makes those filters context-local instead, on by default only in the free-threaded build. Otherwise, keep it in single-threaded test or audit code and set filters once at startup in a concurrent job.
saying these in an interview costs you the question
- Concludes one log line means one call site
- Thinks the code path only ran on the first batch
- Cannot name an action that reports every occurrence
- Believes filters alone decide, ignoring the registry
- Proposes -W error in the production job
- Assumes the counts are comparable across worker processes