skip to content

Bytecode Cache Invalidation

Where .pyc files land and how the interpreter decides one is stale, from the magic number and source timestamp to hash-based checking. Interviewers ask when a deploy seems to run yesterday's code.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

What is Python's __pycache__ directory and when does CPython reuse a .pyc file?

level: juniorimportance: must knowfreq 55%

answer

  1. Compilation is cached, execution is not
  2. A directory beside the source file
  3. The filename carries the interpreter tag
  4. A header is checked before the code object
  5. Magic number, then mtime and size

basics

~20 s

CPython caches each module's compiled bytecode in a pycache directory next to the source. It reuses that .pyc only when the magic number matches the interpreter and the recorded source timestamp and size still match.

solid answer

~50 s

Importing a source module compiles it to a code object and then executes it; only the compile step is cached. By default CPython writes the result to `__pycache__/<name>.<cache_tag>.pyc` beside the source, where the tag is `cpython-314` on 3.14, so several interpreters can share one source tree. A `.pyc` opens with a 16-byte header: a 4-byte magic number, a 4-byte flags field, then either the source's modification time and size or an 8-byte source hash. On the next import the loader checks the magic number first, then compares the stored mtime and size with the file on disk; any mismatch triggers a recompile. The comparison is equality, not newer-than, so a timestamp that moves backwards invalidates too. Writing is skipped silently when the directory is not writable or `sys.dont_write_bytecode` is set, and the cache never changes what the code does.

code

python · 9 lines
python
import importlib.util
import sys

print(sys.implementation.cache_tag)
print(importlib.util.MAGIC_NUMBER.hex())
print(importlib.util.cache_from_source("catalog/loader.py"))
print(importlib.util.source_from_cache(
    importlib.util.cache_from_source("catalog/loader.py")
))

go deeper

for a junior

Be ready to say what pycache holds, that it is compiled bytecode rather than machine code, and that deleting it is harmless. Knowing the cache saves compile time on import, not execution time, is the point most candidates miss.

for a middle

Explain the validation mechanically: magic number first, then the stored source mtime and size compared for equality. Mention that writing the cache is best-effort and silently skipped when the directory is not writable.

for a senior

Show you use this when diagnosing deployments: know that identical mtime and size means the cache is trusted, that the cache can be pre-warmed or relocated with sys.pycache_prefix, and that a stray sourceless .pyc can keep a deleted module importable.

for a principal

Own the policy: whether images ship precompiled caches or run with bytecode writing disabled, whether caches live in the source tree at all, and what that costs in startup time versus reproducibility and read-only-filesystem hygiene.

## Two phases, one of them cacheable Importing a Python source module always does two things: compile the text into a **code object**, then execute that code object in a fresh module namespace to produce the module's attributes. Only the first phase can be cached. Execution has to happen once per process, which is why a `.pyc` never speeds up a running program at all: it only removes parse-and-compile work from the first import. ## Where the file lands Since PEP 3147 (Python 3.2) the cache lives in a directory named `__pycache__` next to the source, and the filename carries the interpreter's **cache tag**. So `catalog/loader.py` caches to `catalog/__pycache__/loader.cpython-314.pyc` on CPython 3.14. `sys.implementation.cache_tag` holds that tag, `importlib.util.cache_from_source()` computes the cache path for a source path, and `importlib.util.source_from_cache()` inverts it. Tagging matters because one checkout is routinely imported by more than one interpreter; before PEP 3147 a 3.x and a 3.y process would overwrite each other's cache on every import. The whole cache tree can be relocated: `sys.pycache_prefix` (settable with `-X pycache_prefix=DIR` or the `PYTHONPYCACHEPREFIX` environment variable) mirrors the directory structure under `DIR` instead of writing into the source tree. That is what you reach for when the source tree must stay pristine or is mounted read-only. ## What the header holds A `.pyc` starts with a fixed 16-byte header followed by the **marshalled** code object (the `marshal` format is a CPython implementation detail, versioned with the interpreter, and is not an interchange format you should hand-craft). * Bytes 0-3: the magic number, available as `importlib.util.MAGIC_NUMBER`. It changes whenever the bytecode format or the compiler's output changes, which is often between feature releases. A `.pyc` written by another version therefore fails on the very first check. * Bytes 4-7: a flags word, added by PEP 552 in 3.7. Bit 0 says whether the file is hash-based; bit 1, on a hash-based file, says whether the hash must be verified. * Bytes 8-15: in the default timestamp mode, the source's modification time truncated to 32 bits (one-second resolution) followed by the source's size truncated to 32 bits. In hash-based mode, an 8-byte hash of the source bytes instead. ## How staleness is decided On import, the loader stats the source file and reads the header. It rejects the cache if the magic number differs. In timestamp mode it then compares the stored mtime with the source's current mtime and the stored size with the source's current size — both are **equality** checks. That is worth saying out loud in an interview, because the common mental model is *the source is newer than the cache*. It is not: any change in either direction, and any change of size at the same timestamp, invalidates. The flip side is that a source whose mtime and size are both restored to exactly what they were is considered fresh even if its contents changed, which is the root of most stale-bytecode stories. On a mismatch the module is recompiled and the cache is rewritten atomically (write to a temporary file, then rename), so two processes importing at once cannot see a half-written file. ## When nothing is written Writing the cache is best-effort and never fatal: * `sys.dont_write_bytecode` is `True` when the interpreter was started with `-B` or with `PYTHONDONTWRITEBYTECODE` set, and no `.pyc` is written at all. * If the directory cannot be created or written, the write is silently skipped and the module is recompiled on every process start. There is no error, just slower startup. * Code compiled from a string with `compile()` or run through `exec()` is never cached, and neither is `__main__`: the script you name on the command line is compiled fresh every run, only the modules it imports are cached. * You can also populate the cache ahead of time with `py_compile.compile()` for one file or the `compileall` module for a tree, which is how a build step warms the cache before the process ever serves traffic. ## Sourceless imports One footgun worth knowing: the legacy pre-3.2 layout — a `.pyc` sitting directly beside where the source used to be, not inside `__pycache__` — is still importable, via `importlib.machinery.SourcelessFileLoader`, but only when no `.py` file exists there. That mechanism is behind the classic "I deleted the module and it still imports". A `.pyc` inside `__pycache__` is never used without its source. Finally, deleting `__pycache__` is always safe. The worst it can cost is one recompile.

  • Why does the cached filename contain something like cpython-314?
    That is the interpreter's cache tag, exposed as `sys.implementation.cache_tag`. It lets one source tree be imported by several interpreters without them overwriting each other's cache files. The magic number inside the header is the second guard: even if a tag somehow collided, a file compiled by a different bytecode version is rejected on the first check.
  • What happens if the __pycache__ directory cannot be written?
    Nothing visible. The write is best-effort: CPython skips it and the module is compiled from source on every process start, which costs startup time but changes no behaviour. That is exactly what `sys.dont_write_bytecode` does on purpose, set by `-B` or by the `PYTHONDONTWRITEBYTECODE` environment variable.
  • Can a .pyc be imported when the .py file is gone?
    Only in the legacy layout, where the `.pyc` sits directly beside where the source was rather than inside `__pycache__`. Then `importlib.machinery.SourcelessFileLoader` imports it, and since there is no source there is nothing to validate against. A file inside `__pycache__` is never used without its source.

Think of it as a receipt stapled to the source file: before trusting the prepared result, the interpreter checks that the receipt was issued by this interpreter and that the file it describes still has the same date and the same length.

saying these in an interview costs you the question

  • Says .pyc files make the running program faster, not just startup
  • Describes a .pyc as machine code rather than serialized bytecode
  • Thinks deleting __pycache__ can break a working program
  • Assumes validation asks whether the source is newer than the cache
  • Believes every interpreter version overwrites the same .pyc file
  • Thinks the script named on the command line is also cached

context

open as a page

What do PEP 552 hash-based .pyc files fix that timestamp .pyc files cannot?

level: middleimportance: should knowfreq 30%

basics

~20 s

A hash-based .pyc records a hash of the source bytes instead of its modification time and size, so caches survive builds that reset timestamps and stay reproducible. Checked mode re-hashes on import; unchecked trusts it.

open as a page

Stale .pyc bytecode caches: how would you prove they are why a redeployed catalogue importer still fails a 340-case encoding regression?

level: seniorimportance: should knowfreq 35%

basics

~20 s

First confirm which files were actually imported, then decode each cached header: magic number, stored mtime and size, compared against the source on disk. A deploy that preserves timestamps and file size makes CPython trust yesterday's bytecode.

open as a page

What are frozen modules in CPython, and why does editing their source change nothing?

level: middleimportance: nice to knowfreq 12%

basics

~10 s

A frozen module has its compiled bytecode embedded in the python executable, so importing it reads no .py and no .pyc and validates nothing. Editing the matching source changes nothing until -X frozen_modules=off.

open as a page