What do PEP 552 hash-based .pyc files fix that timestamp .pyc files cannot?
answer
- The timestamp is metadata, not content
- Reproducible builds want the same bytes
- The header flags word carries the mode
- Two hash sub-modes, one re-reads the source
- compileall and py_compile choose the mode
basics
~20 sA hash-based .pyc records a hash of the source bytes instead of its modification time and size, so caches survive builds that reset timestamps and stay reproducible. Checked mode re-hashes on import; unchecked trusts it.
solid answer
~50 sTimestamp validation ties a cached `.pyc` to the source's mtime and size, which are not properties of the code: a fresh clone, an unpacked archive or a build that normalizes timestamps changes them without changing a byte of source, and two builds of the same tree produce different `.pyc` bytes. PEP 552, in Python 3.7, added an alternative: set bit 0 of the header's flags word and store an 8-byte hash of the source instead. That makes the cache **reproducible** — the same source always yields the same file. Bit 1 chooses the sub-mode. **Checked** hash-based files are re-hashed on every import and rejected if the source changed, trading a stat call for reading and hashing the source. **Unchecked** files are trusted unconditionally, which suits an image where a build step owns freshness. You select the mode with `py_compile.PycInvalidationMode` or `compileall --invalidation-mode`.
code
python · 14 linesimport importlib.util
import pathlib
import py_compile
src = pathlib.Path("catalog_demo.py")
src.write_text("RECORDS = 340\n", encoding="utf-8")
pyc = pathlib.Path(py_compile.compile(
str(src),
invalidation_mode=py_compile.PycInvalidationMode.CHECKED_HASH,
))
header = pyc.read_bytes()[:16]
print(header[:4] == importlib.util.MAGIC_NUMBER)
print(int.from_bytes(header[4:8], "little"))
print(header[8:16] == importlib.util.source_hash(src.read_bytes()))go deeper
You mainly need to know the default exists and is timestamp-based, and that a second mode keys the cache to the source's contents instead. The words checked and unchecked, and roughly why, are enough at this level.
Explain the flags word and the two bits, name py_compile.PycInvalidationMode and compileall --invalidation-mode, and say plainly why an mtime is a poor key for a cache in a build pipeline.
Argue the checked-versus-unchecked choice from who owns freshness: a mutable source tree needs re-hashing, an immutable image does not, and the cost is reading every source file at import instead of one stat call.
Own it as a build-artefact policy: reproducible caches make images comparable and content-addressable, and the decision interacts with how you build, whether your trees are read-only, and how much startup latency you are willing to spend on validation.
## Why the timestamp is the wrong key The default `.pyc` header stores the source's modification time and size, and the loader accepts the cache only if both still match exactly. Both are filesystem metadata, not properties of the code, and that gap causes two distinct problems. **Reproducibility.** Build the same source tree twice and you get two different `.pyc` files, because the embedded mtime differs. Anything that wants byte-identical artefacts from identical inputs — a reproducible build, a content-addressed cache, a diff of two images to prove they contain the same code — is defeated by a field that encodes *when* rather than *what*. Build systems commonly normalize timestamps (for example by honouring `SOURCE_DATE_EPOCH`), which makes the problem worse rather than better: now every file's mtime is the same constant, and the mtime carries no information at all. **Correctness in both directions.** The check is equality, so a timestamp restored to its old value makes changed source look fresh, and a timestamp bumped by an unrelated tool makes unchanged source look stale. Version-control checkouts, archive extraction that preserves or resets times, container layer copies and network filesystems with coarse clocks all move mtimes around independently of content. ## What PEP 552 changed Python 3.7 formalized the header's 4-byte **flags** field (previously always zero) and gave two of its bits meaning: * **bit 0 — hash-based.** When set, bytes 8-15 are no longer mtime and size; they hold an 8-byte hash of the source file's bytes. `importlib.util.source_hash()` computes exactly that value, so you can verify a file by hand. * **bit 1 — check_source.** Only meaningful when bit 0 is set. When set, the loader must re-read and re-hash the source on every import and reject the cache on a mismatch. When clear, the loader trusts the cache without touching the source at all. So three modes exist in practice, and they are exactly the members of `py_compile.PycInvalidationMode`: * `TIMESTAMP` — flags `0`. The default, and what a normal import writes. * `CHECKED_HASH` — flags `0b11`. Reproducible **and** self-correcting: edit the source and the next import notices. * `UNCHECKED_HASH` — flags `0b01`. Reproducible, and never validated at runtime. The file is used as long as the magic number matches. ## Choosing between them The interesting tradeoff is checked versus unchecked, and it is a question of *who owns freshness*. Use **checked** when the source tree can still change under a running interpreter — a developer machine, a mounted source directory, anything where someone might edit a file in place. You pay for it: instead of one `stat()` per module you read the whole source and hash it, which is real work at import time for a large dependency tree, though still far cheaper than recompiling. Use **unchecked** when an immutable build produces both the source and the cache and nothing can edit them afterwards — a container image, an installed wheel, a read-only mount. Here the build step is the invalidation mechanism, and paying to re-hash thousands of files on every process start buys nothing. The risk is symmetrical: if something *does* edit the source, the interpreter will happily go on executing the old bytecode with no way to notice. Keep `TIMESTAMP` as the default for ordinary local work. It is the cheapest check, and mtimes behave well when files are only ever edited in place. ## Producing them You never get a hash-based file by accident: a plain import always writes `TIMESTAMP`. They come from an explicit compile step, which is why they belong to a build. * `py_compile.compile(path, invalidation_mode=py_compile.PycInvalidationMode.CHECKED_HASH)` for one file. * `python -m compileall --invalidation-mode checked-hash <dir>` for a tree, or `unchecked-hash`, or `timestamp`. The same choice is available programmatically through `compileall.compile_dir()`. Note that `compileall` already defaults to `checked-hash` when `SOURCE_DATE_EPOCH` is set in the environment, on the reasoning that a build which normalizes timestamps has made them useless as a cache key. A useful detail for debugging: because the mode lives in the flags word, you can tell what any `.pyc` is just by reading four bytes at offset 4. `0` is timestamp, `1` is unchecked hash, `3` is checked hash. If you are ever unsure why a deployed tree does or does not pick up a change, that word answers half the question immediately. ## What it does not fix Hash-based caching does not survive an interpreter upgrade — the magic number still governs, and it must match. It also says nothing about *which* file was imported: if `sys.path` resolves a name to a different directory than you expected, a perfectly valid cache of the wrong module is still the wrong module.
- What does an unchecked hash-based .pyc cost you if someone edits the source anyway?The interpreter keeps running the cached bytecode indefinitely, with no error and no warning, because it never looks at the source. That is the deal you accept in exchange for skipping the hash on every import, so it is only safe where the artefact is genuinely immutable — a container image or an installed, read-only tree — and the build owns invalidation.
- Is checked-hash validation more expensive than the default timestamp check?Yes. Timestamp mode needs one `stat()` per module; checked-hash mode reads the whole source file and hashes it. For a large dependency tree that is measurable at startup, though still much cheaper than recompiling. It buys correctness that does not depend on filesystem metadata, which is the reason to pay it.
- How can you tell which invalidation mode a given .pyc was built with?Read the 4-byte flags word at offset 4 of the header. `0` means timestamp, so bytes 8-15 are mtime and size; `1` means unchecked hash-based; `3` means checked hash-based, and in both hash cases bytes 8-15 hold the source hash that `importlib.util.source_hash()` reproduces.
saying these in an interview costs you the question
- Thinks hash-based .pyc files are the default since 3.7
- Confuses the source hash with PYTHONHASHSEED string hashing
- Believes unchecked mode still notices an edited source
- Says hash-based caches survive an interpreter version change
- Cannot name a way to produce one other than importing