skip to content

Why can sys.modules hold the same source file twice under two different names?

level: seniorimportance: should knowfreq 32%

answer

  1. Ask what the cache is keyed by
  2. One file, two reachable names
  3. Parent and package both on the search path
  4. Two module objects, two copies of state
  5. isinstance fails and one buffer never drains

basics

~20 s

Because the cache is keyed by module name, not by file. A file reachable under two names — typically when a package and its parent are both on sys.path — misses the cache twice and becomes two modules.

solid answer

~50 s

`sys.modules` is keyed by fully qualified name, so identity is a *name* property, not a file property. When both a package and its parent directory are on `sys.path`, `etl.rows` and `rows` are different keys reaching the same file; each key misses the cache, so the body executes twice and you get two module objects with two copies of every module-level class, registry and cache. The symptoms are confusing rather than loud: `isinstance` fails against the twin's class, an `except SomeError` block does not catch the twin's exception class, module-level state is duplicated so half the writes are invisible, and anything holding entries — a batch buffer, a connection pool, a handler list — grows without bound because only one copy is ever drained. Diagnose by grouping `sys.modules` values by `__file__` and comparing identity; fix by giving the code exactly one import path.

code

python · 19 lines
python
import sys
import tempfile
import pathlib

root = pathlib.Path(tempfile.mkdtemp())
pkg = root / "etl"
pkg.mkdir()
(pkg / "__init__.py").write_text("")
(pkg / "rows.py").write_text("class Row:\n    pass\n\nCACHE = []\n")

sys.path[:0] = [str(root), str(pkg)]

import etl.rows
import rows

print(etl.rows is rows)
print(etl.rows.Row is rows.Row)
print(isinstance(rows.Row(), etl.rows.Row))
print(sorted(name for name in sys.modules if "rows" in name))

go deeper

for a junior

Take away the headline: the import cache is keyed by module name, so one file imported under two names becomes two separate modules with separate module-level variables.

for a middle

Explain how the second name arises — a package and its parent both on sys.path — and predict the symptoms: failed isinstance checks, an except clause that misses, and module-level state that exists twice.

for a senior

Diagnose it live. Group sys.modules by file comparing identity, read type(obj).module on the accumulating objects, tie it to the path entries at startup, and fix the layout rather than the symptom.

for a principal

Own the invariant: one canonical import path per deployable, enforced by installing the project rather than mutating sys.path, plus a startup assertion so a mis-launched process fails immediately instead of leaking for a week.

## One file, two identities The import cache is a dictionary keyed by fully qualified module name. Nothing in it is keyed by path, inode or content. So the question "is this module already loaded?" is really "has this *name* been imported?", and any arrangement that makes one file reachable under two names produces two fully independent module objects from it. The common arrangements: * **A package and its parent are both on `sys.path`.** The parent makes `etl.rows` importable; the package directory itself makes `rows` importable. Two keys, one file, two modules. This is almost always the fingerprint of runtime `sys.path` manipulation or of an entry inherited from `PYTHONPATH`. * **The same tree is reachable through two path entries** — a checkout and an installed copy, or a symlinked and a real directory. * **A file that is both executed as the program and imported by name** carries the same hazard, since the executed copy lives under a different key than the imported one. Note that duplicated *names* pointing at the same object are fine and even deliberate — the standard library does it for a few aliases. The defect is two distinct objects backed by one file. ## Why the symptoms are so indirect Nothing fails at import time. The failures come later, from the fact that the two modules share no state: * **Type checks stop working.** Classes are compared by identity, so `isinstance(row, etl.rows.Row)` is `False` for a `Row` built from the other copy, and `except etl.rows.ExportError` will not catch the twin's `ExportError` even though both were written once, in one file. * **Module-level state is duplicated.** Registries populated by decorators end up half-full in each copy, so a lookup "randomly" misses depending on which copy the caller imported. * **Initialisation happens twice.** Anything the body does — installing a logging handler, opening a pool, spawning a thread, registering an exit hook — happens once per copy, which is how you get duplicated log lines and double-registered hooks. * **Unbounded growth.** This is the one that reaches production. Take an ETL export to a warehouse whose staging module keeps a module-level list of rows and flushes it when it crosses a threshold. Under duplicate identity there are two lists: the writer appends to the copy it imported, the flusher drains the copy *it* imported, and the untouched list only ever grows. Resident memory climbs steadily under exactly the workload the code was written to bound, and no single-batch test reproduces it because a short run never gets far enough for either list to matter. In the same codebase a 27-minute suite can stay green throughout, because the runner imports everything under one consistent name and never creates the second copy at all. ## Diagnosing it The direct check is to group loaded modules by file and look for two *different objects* sharing one path: ```python import sys seen = {} for name, module in list(sys.modules.items()): path = getattr(module, "__file__", None) if path is None: continue first_name, first_module = seen.setdefault(path, (name, module)) if first_module is not module: print("loaded twice:", path, first_name, name) ``` Comparing objects rather than names matters: grouping by name alone reports harmless aliases such as `os.path` and `posixpath`, which are two keys for one object. Other signals worth knowing: * `type(obj).__module__` on a value that failed a type check tells you which copy created it — seeing `rows` where you expected `etl.rows` is the whole diagnosis. * `python -X importtime` shows the same file being executed under two names. * Printing `sys.path[:5]` at startup usually reveals the offending pair of entries immediately. ## Fixing it — and preventing it The fix is never to patch around the symptom. Do not normalise types with duck-typing, do not delete one of the `sys.modules` entries, and do not try to reconcile the two copies of the state. Give the code exactly one import path: * **Install the project** into the environment — editable during development — so the package is importable from one place and nobody needs to append anything at runtime. * **Never put both a package directory and its parent on `sys.path`.** If you must add an entry programmatically, add the parent only. * **Run entry points as installed commands or with `python -m`**, so the interpreter does not silently add a directory that turns an internal module into a top-level one. * **Harden the process** where it matters: `-P` or `PYTHONSAFEPATH=1` (3.11 and later) stops CPython prepending the script or working directory at all, which removes the most common accidental second entry. * **Assert it in a smoke test.** The duplicate-detector above is a handful of lines and, run once at startup or as a test, converts a class of week-long memory investigations into an immediate, obvious failure. The underlying lesson generalises: in Python, *module identity is a name, resolved through `sys.path`*. Two names are two modules, however identical their source, and every bit of module-level state you rely on being singular depends on that name being unambiguous.

  • How would you prove that duplicate module identity is the cause of climbing memory rather than a plain leak?
    Look for two module objects sharing one `__file__` in `sys.modules`, and check `type(obj).__module__` on the objects that are accumulating. If the growing container lives at module level and the code that drains it imported the twin, the two names will differ — `rows` versus `etl.rows` — and that is conclusive. A plain leak shows one module and growth traceable to referrers instead.
  • Why is deleting one of the two sys.modules entries the wrong fix?
    Because it removes the cache entry, not the module. Every object already created from that copy — classes, instances, registry entries, the accumulating container — keeps pointing at it, so the mismatch survives while the next import silently creates a third module. The only durable fix is to make the file reachable under one name: install the project, keep the package's parent alone on the path, and stop appending directories at runtime.
  • Why can a full test run stay green while production shows the problem?
    Because the runner controls the import path. It typically imports everything under one consistent package name and never adds the package directory itself, so the second copy is never created. The duplicate identity is a property of how the process was launched, not of the code, which is why it should be asserted at startup — a short duplicate check run in the real entry point catches what the suite structurally cannot.

It is the same document filed in two drawers under two labels. Both are real, both accept edits, and the one nobody empties keeps filling up.

saying these in an interview costs you the question

  • Thinks sys.modules is keyed by file path or inode
  • Says two identical classes from one file must be the same object
  • Proposes deleting a sys.modules entry as the fix
  • Adds both a package and its parent to sys.path deliberately
  • Blames the garbage collector for the growing module-level container
  • Assumes a green test suite rules out an import-layout problem

context