skip to content

Why can python -m catalogue.loader leave two copies of loader.py's state in one process?

level: seniorimportance: should knowfreq 28%

answer

  1. Two copies of one file's globals
  2. Ask what the cache key is
  3. Entry module is not named for its file
  4. Classes compared by identity, not name

basics

~20 s

The file executes once as main and again under its dotted name if anything imports it, because sys.modules is keyed by module name rather than by file path. Each execution builds separate globals: two caches, two class objects, two sets of side effects.

solid answer

~50 s

The entry module is cached in `sys.modules` under the key `"__main__"`, not under its import name. So if a museum-catalogue importer is started with `python -m catalogue.loader` and anything in the process then does `import catalogue.loader`, the cache misses, the file executes a second time, and the process holds two independent module objects for one file. Everything at module level exists twice: the normalisation cache reporting an 83% hit rate in one copy is not the dict the request path reads, so callers see a stale cached value; `class Record` is two distinct classes, so `isinstance` fails across the boundary and an `except` clause can miss an exception of what looks like the same class. The fix is a thin entry module — a `__main__.py` that imports and calls — holding no state of its own.

code

python · 23 lines
python
import pathlib
import subprocess
import sys
import tempfile

SOURCE = """\
CACHE = {}
print("executed as", __name__, "cache id", id(CACHE))
if __name__ == "__main__":
    from catalogue import loader
    print("same cache object?", loader.CACHE is CACHE)
"""

with tempfile.TemporaryDirectory() as tmp:
    pkg = pathlib.Path(tmp, "catalogue")
    pkg.mkdir()
    pkg.joinpath("__init__.py").write_text("")
    pkg.joinpath("loader.py").write_text(SOURCE)
    done = subprocess.run(
        [sys.executable, "-m", "catalogue.loader"],
        cwd=tmp, capture_output=True, text=True,
    )
    print(done.stdout)

go deeper

for a junior

Know that the file you launch is called main inside the process, and that importing it by its real name is a different module as far as the import system is concerned.

for a middle

Explain the cache key: sys.modules maps names to modules, so one file under two names executes twice and produces two independent sets of module globals.

for a senior

Diagnose it from symptoms in a running service, such as a cache that never seems to warm or a handler that misses its own exception class, and fix it by making the entry module hold nothing.

for a principal

Set the entry-point convention for the organisation, so services expose a thin runnable module and all state lives in modules that are only imported, including in worker processes that rebuild it.

### `sys.modules` is keyed by name, not by file The import system caches modules in `sys.modules`, a dictionary from **module name** to module object. The entry module — whatever the interpreter was started with — is executed under the name `__main__` and cached under the key `"__main__"`. Its file path is not part of the cache key. So when you run `python -m catalogue.loader`, the file `catalogue/loader.py` executes and lands in `sys.modules["__main__"]`. If any other module in that process then does `import catalogue.loader`, the cache is asked for the key `"catalogue.loader"`, does not find it, and the import system executes **the same file a second time**, storing the result under the dotted name. One file, one process, two module objects, two independent `globals()` dictionaries. The same duplication happens with the script form, `python catalogue/loader.py`, whenever anything also imports the module by name. ### What duplication actually costs Module-level state is per module object, so everything defined at the top level exists twice: * **Caches and registries.** A museum-catalogue importer keeps a module-level `dict` mapping accession numbers to normalised records. Run as `__main__` it warms that dict and reports an **83% cache-hit rate**; the request path, which reached the module through `import catalogue.loader`, holds the *other* dict and reads a stale cached value for records the first copy had already re-normalised. Neither cache is wrong on its own — there are simply two of them, and only one is being filled. * **Classes.** `class Record` executed twice yields two distinct class objects. `isinstance(obj, Record)` is `False` across the boundary, `pickle` round-trips can fail, and equality logic that checks `type(other) is type(self)` silently reports objects unequal. * **Exceptions.** The same problem with worse symptoms: an `except CatalogueError:` clause compiled against one copy simply does not match an exception raised from the other, because `except` matches by class identity, not by name. The traceback names a class that looks exactly like the one you are catching. * **Side effects.** Anything the module body does — opening a connection pool, registering a signal handler, adding a logging handler, starting a background thread — happens twice. ### Worker processes add a third copy With the spawn and forkserver start methods, `multiprocessing` cannot inherit the parent's memory, so the child re-executes the parent's main module (under a name other than `__main__`) to rebuild the definitions a pickled callable refers to. The guard is what stops the *launch* code from running again in the child; the module-level *state* is genuinely rebuilt per child, because it must be. This is a design constraint, not a bug: state that a worker needs must be created inside the worker or passed to it explicitly, never assumed to be shared with the parent. This matters more on 3.14 than it used to. Through 3.13 the default start method on Linux was fork, which inherits the parent's memory and never re-imports anything; in **3.14** the default on Unix platforms other than macOS became forkserver, and like spawn it re-imports the main module. Code that relied on fork's inheritance now sees the main module executed again in the child. ### The fix, and how to check for it The reliable defence is to make the module that runs as `__main__` hold nothing worth duplicating: ```python # catalogue/__main__.py -- the entry module: no state, no definitions from catalogue.loader import main raise SystemExit(main()) ``` All the classes, caches and functions live in `catalogue/loader.py`, which is only ever *imported*, so exactly one copy exists. Then run it as `python -m catalogue` (or through the console command an install creates). Never import the module that is currently running as `__main__`, and never let a module import itself by its dotted name for convenience. To confirm the diagnosis in a live process, look at names rather than paths: print the `__name__` of the module holding the state, or check whether `sys.modules` contains both `"__main__"` and the dotted name pointing at objects with different `id()` values. A cheap tripwire is an assertion in the module body that the module was imported rather than executed. Reading `__file__` will *not* reveal it — both copies report the same file, which is precisely why the bug is confusing. ### How to answer it in an interview Lead with the cache key: `sys.modules` maps names to modules, and the entry file is named `__main__`, so importing it by its real name executes it again. Then name a concrete symptom — a duplicated cache, or an `except` clause that fails to catch its own exception class — and finish with the fix: a thin `__main__.py` that imports and calls, holding no state of its own.

  • How does this show up as an except clause that fails to catch its own exception class?
    `except` matches by class identity, walking the raised exception's real base classes. If the raising code holds the `__main__` copy of the module and the handler was compiled against the imported copy, the two `CatalogueError` classes are different objects with no inheritance relationship, so the clause does not match. The traceback prints a class name identical to the one you are catching, which is what makes it so confusing.
  • What shape of entry module avoids the duplication entirely?
    A `__main__.py` (or an installed console entry function) that holds no classes, no caches and no side effects — just `from catalogue.loader import main` and a call. Everything with state lives in modules that are only ever imported, so exactly one copy exists. The rule of thumb: never import the module that is currently running as `__main__`.
  • Why do worker processes rebuild module-level state rather than share it?
    With the spawn and forkserver start methods the child is a fresh interpreter that cannot inherit the parent's objects, so it re-imports the main module to reconstruct the definitions a pickled callable refers to. Module-level caches are therefore per worker by construction. State a worker needs must be created in the worker or passed explicitly; assuming a shared module-level dict across processes is a bug, not a tuning problem.

saying these in an interview costs you the question

  • Says sys.modules deduplicates by file path
  • Thinks two module objects share one globals dict
  • Claims isinstance still works across the two copies
  • Blames the cache logic rather than the double execution
  • Believes using -m by itself prevents the duplicate
  • Tries to diagnose it by comparing __file__ values

context