Stale .pyc bytecode caches: how would you prove they are why a redeployed catalogue importer still fails a 340-case encoding regression?
answer
- First ask what was actually imported
- Reading the source proves nothing here
- Disassemble, and check the cache header
- Two fields must both stay unchanged
- Preserved timestamps plus an equal byte count
basics
~20 sFirst confirm which files were actually imported, then decode each cached header: magic number, stored mtime and size, compared against the source on disk. A deploy that preserves timestamps and file size makes CPython trust yesterday's bytecode.
solid answer
~40 sStart by proving *what* is running, not *why*: check `module.__file__` and `module.__spec__.origin` to be sure the import resolved to the deployed tree, then disassemble a changed function with `dis` — reading the source with `inspect.getsource()` shows the new text even when old bytecode is executing, so it lies here. Next decode the cache: `importlib.util.cache_from_source()` gives the path, and the 16-byte header gives the magic number, the flags word and the stored mtime and size. If those equal the source's current mtime and size, CPython considers the cache fresh, and the usual cause is a deploy that rewrites files while preserving timestamps — an archive extraction, a timestamp-normalizing build — where an encoding fix leaves the byte count unchanged. Fix it by rebuilding with `compileall --invalidation-mode checked-hash`, or by shipping no caches and setting `PYTHONDONTWRITEBYTECODE`.
code
python · 16 linesimport importlib.util
import os
import pathlib
import py_compile
import struct
src = "catalog_demo.py"
pathlib.Path(src).write_text("ENCODING = 'utf-8'\n", encoding="utf-8")
py_compile.compile(src)
pyc = pathlib.Path(importlib.util.cache_from_source(src))
magic, flags, field1, field2 = struct.unpack("<4sIII", pyc.read_bytes()[:16])
stat = os.stat(src)
print("magic ok:", magic == importlib.util.MAGIC_NUMBER)
print("hash based:", bool(flags & 1))
print("mtime match:", field1 == (int(stat.st_mtime) & 0xFFFFFFFF))
print("size match:", field2 == (stat.st_size & 0xFFFFFFFF))go deeper
Know that a leftover compiled cache can make an interpreter skip recompiling, and that deleting the pycache directories is a safe first move. You are not expected to decode the header yet.
Explain why a deploy can defeat timestamp validation: both the recorded mtime and the recorded size must change, and a fix that preserves the file size under a restored timestamp changes neither.
Demonstrate the diagnostic order out loud: process identity and import path first, then disassembly rather than the source text, then the cache header. Finish with a durable fix, not a manual cleanup.
Frame it as artefact policy — immutable deployed trees, caches precompiled in checked-hash mode or excluded entirely, caches kept outside the source tree — so that no incident ever again spends time on whether the running code is the shipped code.
## Rule out the boring explanations first "It is running old code" has three candidate causes and only one of them is the cache. Eliminate the other two before you touch a `.pyc`. **Is it even the right file?** Print `module.__file__` and `module.__spec__.origin` for the module you changed, in the running process. A shadowing copy earlier on `sys.path`, an editable install pointing elsewhere, or a stale copy installed into the environment alongside the deployed tree explains far more incidents than bytecode ever does. **Is it even the new process?** A supervisor that failed to restart, a pre-forked worker still alive from the previous release, or a warm process pool serving the old image will reproduce every symptom of a stale cache. Only when the process is new and the path is right does the cache become the suspect. ## Prove it, do not guess The trap is that reading the source proves nothing. `inspect.getsource()` and any editor open the `.py` on disk, which contains the fix. The executing code object came from somewhere else. So compare the *bytecode* to the source: * `dis.dis(catalog.loader.decode_record)` shows what is actually loaded. If the fix changed a constant — an encoding name, a fallback value — it is visible directly in the disassembly or in the function's `__code__.co_consts`. * Run the interpreter with `-v`, which logs, per module, whether it matched a cached file or compiled the source. That is the fastest one-command answer on a running deployment. Then decode the header. `importlib.util.cache_from_source(src)` gives the cache path; the first 16 bytes give the magic number, the flags word and the two validation fields. Compare the magic against `importlib.util.MAGIC_NUMBER`, and in timestamp mode compare the stored mtime and size against `os.stat()` of the source. If they match, you have your answer: the interpreter was correct to reuse the cache, and the deploy lied to it. ## Why an honest deploy produces a lying cache Timestamp validation asks whether the source's mtime and size are exactly what they were when the cache was written. Both can be preserved while the contents change: * Unpacking an archive or copying a tree with timestamps preserved restores the *original* mtimes, which may be older than the cache that a previous release left behind in the same directory. * Build pipelines that normalize timestamps for reproducibility set every file to the same constant, so mtime carries no information whatsoever. * A one-second mtime resolution means a file rewritten within the same second as the previous compile looks unchanged. * And the size check only helps when the size moved. An encoding fix is exactly the kind of change that does not move it: swapping one codec name for another of the same length, or changing a decode argument, can leave the file byte count identical. Stack a preserved mtime on top of an unchanged size and both checks pass. Every one of the 340 regression cases then exercises the old decoding path, which is why the failure looks like the fix was never deployed. The adjacent trap is a leftover `.pyc` in the legacy location — beside the source rather than inside `__pycache__` — for a module whose `.py` was deleted in this release. That file is still importable on its own, so the removed module goes on importing and shadowing whatever replaced it. ## Fixes, in order of durability **Immediately:** delete the cache directories in the deployed tree and restart. That is always safe; the worst it costs is one recompile. **Structurally, pick one of three policies.** 1. **Precompile with hashes.** Build with `python -m compileall --invalidation-mode checked-hash <tree>`. Freshness stops depending on filesystem metadata: the cache is keyed to the source bytes, so a change of contents is always noticed and a change of timestamps never is. Use `unchecked-hash` instead only if the deployed tree is genuinely immutable and you want to skip validation entirely. 2. **Ship no caches at all.** Set `PYTHONDONTWRITEBYTECODE` (or start with `-B`) and remove any `__pycache__` from the artefact. Nothing can be stale because nothing is cached. You pay compile time on every process start, which matters for short-lived processes and barely at all for a long-running service. 3. **Move the cache out of the tree.** `PYTHONPYCACHEPREFIX`, or `sys.pycache_prefix`, mirrors caches into a separate directory. Deploying a new tree then cannot inherit a previous release's cache files, and the source tree stays clean enough to mount read-only. Whichever you choose, make the artefact itself immutable: a deployed tree that no one rewrites in place removes the entire class of bug, because rewriting a file underneath a cache is what created it. ## The habit to leave with Build the check into the pipeline rather than into your memory. A deploy step that either removes every cache directory or regenerates them in checked-hash mode costs seconds and permanently retires "is it running the new code?" as a question anyone has to ask during an incident.
- Why is reading the source file a useless check in this situation?Because every source-reading tool reads the `.py` on disk, which already contains the fix. `inspect.getsource()`, your editor and any log of the file contents will all show the new code while the interpreter executes a code object that came from the cache. The only honest comparison is against the loaded bytecode, via `dis` or the function's `__code__` attributes.
- How would checked-hash caches have prevented this specific failure?They key validation to the source bytes rather than to mtime and size, so a preserved timestamp and an unchanged byte count are irrelevant. Any content change, including a one-character codec name, produces a different hash and forces a recompile. The cost is reading and hashing each source file at import instead of one stat call.
- What would you check before blaming the bytecode cache at all?That the process actually restarted, and that the module resolved to the tree you deployed. Print `module.__file__` and `module.__spec__.origin` in the running process and inspect `sys.path` for a shadowing copy or an older installed distribution. A surviving worker or a wrong path explains far more of these incidents than caching does.
saying these in an interview costs you the question
- Says just delete __pycache__ without explaining how it went stale
- Reads the source file and concludes the fix was deployed
- Thinks CPython compares the .pyc mtime against the .py mtime
- Claims caches are content-hashed by default so this cannot happen
- Assumes a failed cache write would have raised an error
- Ignores a surviving worker process or a shadowing sys.path entry