How can a stale `.pyc` in `__pycache__` make a patched module keep running old code?
answer
- The interpreter may not be running your edit
- A cached artefact sits beside the source
- Validation is a heuristic, not a comparison
- Modification time and size, in a header
- Hash-based caching removes the guesswork
basics
~20 sCPython caches compiled bytecode beside the source and by default recompiles only when the source's recorded modification time or size has changed. A deploy that rewrites a file without changing either silently keeps running the old bytecode.
solid answer
~50 sA cached file under `__pycache__` starts with a 16-byte header: a 4-byte magic number identifying the bytecode format, a flags word, and then either the source's modification time and size (the default) or a hash of the source (PEP 552, since 3.7). On import the loader compares those against the source; a mismatch triggers a re-compile, a match means the cached bytecode is loaded as-is. The failure mode is timestamp mode believing a lie — a restored or normalised mtime with an unchanged size, a clock skew in a build image, an editor that preserves timestamps. Then a patched payroll import module keeps computing the old rounding rule and no amount of re-reading the source explains it. Diagnose by checking `importlib.util.cache_from_source` for the file the loader is actually using, delete `__pycache__`, and for deployments build with `compileall --invalidation-mode checked-hash`, which validates against a source hash instead.
code
python · 6 linesimport importlib.util
import sys
print(sys.implementation.cache_tag)
print(importlib.util.MAGIC_NUMBER)
print(importlib.util.cache_from_source("payroll_import.py"))go deeper
Know that pycache holds compiled bytecode, that Python manages it for you, and that deleting it is safe. If behaviour does not match the source you just edited, that directory is worth a look.
Explain the validation rule: the cached file's header records the source's modification time and size, and the loader compares them. Be able to name the hash-based alternative and how the interpreter tag keeps versions apart.
Diagnose it in a deployment: identify which artefact actually loaded, reproduce the stale case, then fix it at build time with checked-hash compilation or a precompiled read-only image rather than by clearing caches by hand.
Own the build policy — when caches are produced, validated and shipped — and rule out source-less deployments, which break on every interpreter upgrade and buy no real protection.
Compiling Python to bytecode is not free, so CPython caches the result. The cache is almost always invisible — until the day it is not, and a module you just patched keeps running the code you replaced. ## Where the cache lives and what is in it Importing `payroll_import.py` writes `__pycache__/payroll_import.cpython-314.pyc`. The directory and the tag in the filename come from PEP 3147: the tag is `sys.implementation.cache_tag`, so builds from different interpreters and different implementations coexist instead of clobbering each other. `importlib.util.cache_from_source(path)` computes that filename for you, which is the fastest way to see which file the loader will consult. The file begins with a 16-byte header: * **bytes 0–3** — a magic number, available as `importlib.util.MAGIC_NUMBER`. It changes whenever the bytecode format changes, which is at least once per minor release. A mismatch means "recompile"; for a source-less deployment it means `ImportError`. * **bytes 4–7** — a flags word. Bit 0 selects hash-based validation, bit 1 says the hash must actually be checked. * **bytes 8–15** — in the default *timestamp* mode, the source's modification time (whole seconds) and its size; in *hash* mode, an 8-byte hash of the source bytes. Then comes the marshalled code object. There is no machine code anywhere in the file. ## The failure In timestamp mode the loader compares the recorded mtime and size against the source file it found. If both match it loads the cached bytecode without ever reading the source. That check is a heuristic, and it is defeated by ordinary tooling: * a checkout, rsync or image build that restores or normalises timestamps; * a patch applied in place whose replacement happens to be the same size, within the same second, or with the mtime deliberately preserved; * clock skew between a build machine and a runtime, or a container layer whose mtimes are pinned for reproducibility; * a `__pycache__` directory baked into an image, then overlaid with newer sources whose timestamps look older. The symptom is maddening because the source on disk is unambiguously correct. A payroll CSV importer keeps producing the old floating-point rounding after the fix was deployed; a print statement you added never appears. The tell is that `payroll_import.__file__` points at the source you edited while the behaviour matches the previous release — at which point the question is not "what is wrong with my code" but "which artefact is the interpreter actually executing". ## Diagnosing and fixing Delete `__pycache__` (or touch the sources) and re-run; if the behaviour changes, you have your answer. Then fix it at the deployment level rather than by ritual cleaning: * Build the cache explicitly at image-build time with `python -m compileall -f --invalidation-mode checked-hash <dir>`. Checked-hash files record a hash of the source and re-verify it on every import, so a same-size, same-mtime edit is caught. The cost is reading and hashing the source, which is small next to compiling it. * Use `unchecked-hash` only when the image is immutable and produced by a trusted build step; it stores the hash but never validates, trading safety for a slightly faster import. * Or turn the cache off for the runtime with `-B` or `PYTHONDONTWRITEBYTECODE=1`, having pre-compiled at build time. This is common for read-only container filesystems, where the process would otherwise fail to write and silently recompile every module on every start — measurable for a large dependency tree and for short-lived CLI processes. * `PYTHONPYCACHEPREFIX` (or `-X pycache_prefix`) relocates the whole cache tree out of the source directory, which keeps a read-only or bind-mounted source clean. Two details round out the mental model. The optimisation level is part of the filename, not the header: running under `-O` writes `payroll_import.cpython-314.opt-1.pyc`, and `-OO` writes `.opt-2`, so an optimised build and a normal one never read each other's cache. And a cached file sitting alone inside `__pycache__` is not importable at all — the loader only considers it when the matching source is present. A source-less deployment has to put the file where the `.py` would have been, in the legacy pre-PEP-3147 layout, which is worth knowing before someone proposes shipping bytecode only. ## The related trap Bytecode is version-specific, and the magic number is what enforces it: a file compiled by 3.13 is simply not loadable by 3.14, and the cache tag in the name keeps them apart on disk anyway. That is fine when the source is present — the interpreter recompiles — and fatal when it is not. Shipping `.pyc` files without sources as a form of obfuscation therefore buys very little (they decompile readily) while guaranteeing that a Python upgrade breaks the deployment outright.
- When would you choose hash-based cache files over the default timestamp ones?Whenever source timestamps are not trustworthy: container images, reproducible builds, checkouts that reset mtimes, or bind-mounted code. `checked-hash` re-hashes the source on every import and is the safe default there; `unchecked-hash` skips validation entirely and suits an immutable image compiled by a trusted build step.
- Can a .pyc built by another Python version be loaded?No. The four-byte magic number at the head of the file changes whenever the bytecode format changes, at least every minor release, and the filename carries a tag such as `cpython-314`. With the source present the interpreter simply recompiles; without it, the import fails.
- What does disabling bytecode writing cost you?Every module is recompiled on every process start. For a large dependency tree that is noticeable, and it hurts most for short-lived CLI processes. The usual pattern for a read-only container is to compile once at build time with `compileall` and then run with `-B` or `PYTHONDONTWRITEBYTECODE=1`.
Timestamp validation is a receipt stapled to a copy: if the receipt still matches, nobody opens the original to check. Change the original without disturbing the receipt and the copy is served forever.
saying these in an interview costs you the question
- Thinks .pyc files contain native machine code
- Says deleting __pycache__ is superstition, never a real fix
- Believes a .pyc from another Python version still loads
- Assumes editing the source always invalidates the cache
- Ships .pyc without sources and calls it obfuscation
- Blames the import system before checking which file loaded