skip to content

What does Python do with sys.modules when a module is imported a second time?

level: middleimportance: must knowfreq 58%

answer

  1. Import checks something before searching
  2. A module body executes only once
  3. The lookup is a dictionary keyed by name
  4. The entry exists before the body runs
  5. Deleting the entry yields a second module object

basics

~20 s

It finds the module already there and stops. The import statement checks sys.modules by name first, and on a hit it only binds the name in the importing namespace; the module body runs exactly once per interpreter.

solid answer

~50 s

Every `import` starts by looking the fully qualified name up in `sys.modules`, a plain dictionary from module name to module object. On a hit, the import is done — the machinery binds the name locally and no file is searched, read or executed. Only on a miss does the search along `sys.path` happen; the resulting module object is then inserted into `sys.modules` *before* its body runs, so a module's top-level code executes exactly once per interpreter no matter how many places import it. Two consequences matter in practice. First, `from mod import thing` copies a reference into the importing namespace, so replacing the `sys.modules` entry later does not rebind names that were already imported. Second, deleting an entry makes the next import re-execute the file and produce a *new, distinct* module object while every existing reference still points at the old one.

code

python · 15 lines
python
import sys
import tempfile
import pathlib

tmp = pathlib.Path(tempfile.mkdtemp())
(tmp / "cfg.py").write_text("print('cfg body running')\nCOUNT = 1\n")
sys.path.insert(0, str(tmp))

import cfg
import cfg
print(cfg is sys.modules["cfg"])

del sys.modules["cfg"]
import cfg as cfg_again
print(cfg is cfg_again)

go deeper

for a junior

Recall that Python imports a module once per process and caches it, so top-level code in that module does not run again on later imports. Knowing sys.modules is the cache is enough here.

for a middle

Explain the lookup order — cache first, path search only on a miss — that the entry is created before the body runs, and why a from-import binding does not follow later changes to the cache.

for a senior

Show what this means in a running system: why import-time side effects are hard to operate, why deleting cache entries to pick up edits produces two module objects, and when restarting is the correct answer.

for a principal

Own the design consequence: module-level state is a per-interpreter singleton, so where initialisation lives, how it is configured and whether it survives a subinterpreter or worker boundary are architecture decisions, not import trivia.

## The cache is the first step of import, not an optimisation bolted on `sys.modules` is a dictionary whose keys are fully qualified module names — `'json'`, `'etl.rows'`, `'concurrent.futures'` — and whose values are module objects. The import statement consults it first, every time. If the name is there, the import ends immediately: the machinery binds the name (or the requested attribute) in the importing namespace and returns. No finder runs, no file is opened, no code executes. That single rule produces most of the behaviour people describe as "Python imports are weird". ### A module body runs exactly once A hundred modules can `import config`; the file executes on the first one and never again for the life of that interpreter. This is why module-level code is effectively initialisation-once code, and why a module-level object is the simplest singleton Python offers: everyone importing the name receives the same object because everyone receives the same module. It is also why import-time side effects are risky. Opening a connection or reading a file at module level happens once, at an unpredictable moment — whenever some import chain first reaches that name — and never again, so it cannot be retried or reconfigured. ### The entry is created *before* the body runs The machinery inserts the new, still-empty module object into `sys.modules` and only then executes its body. If the body imports something that imports back, the second import sees the partially populated module rather than starting a fresh execution — which is what stops mutual imports from recursing forever, though the half-initialised state it exposes is its own well-known failure mode. A second consequence of insert-before-execute: if the body raises, the machinery removes the entry again, so a failed import does not leave a broken module cached for later importers to find. ### Binding versus caching — the distinction that catches people The cache maps *names to module objects*. Your namespace holds *bindings to whatever you asked for*. `import mod` binds the module object; `from mod import thing` binds `thing` itself and forgets where it came from. So replacing `sys.modules['mod']` afterwards changes what *future* imports get, and changes nothing about names that were already bound. Any code holding `from mod import thing` keeps the original object. This is exactly why patching by target string has rules about *where* to patch: you must replace the attribute on the object the calling code looks it up on, not the one it was originally defined on. ### Deleting an entry does not "unload" anything `del sys.modules['mod']` removes the cache entry, nothing more. The module object stays alive as long as anything references it — other modules, class definitions, instances, closures. The next `import mod` misses the cache, searches the path again, executes the file again, and creates a *second* module object. You now have two modules from one file, with two copies of every module-level variable, class and registry, and code holding references to the first will not agree with code that imported the second. Classes are compared by identity, so `isinstance` across the pair fails; module-level caches, registries and connection pools exist twice. That is why "clear it from `sys.modules` and import again" is a poor way to pick up a source edit in a long-lived process. Restarting the process is the honest answer; anything else leaves you reasoning about which copy each object came from. ### Submodules and packages Importing `pkg.sub` puts three things in play: an entry `'pkg'`, an entry `'pkg.sub'`, and — importantly — an attribute `sub` set on the `pkg` module object. So `sys.modules` mirrors the dotted namespace rather than describing a flat set of files, and a package's `__init__` body has run before any of its submodules do. It is also perfectly normal for two *names* to map to the same module object: aliases such as `os.path` do it deliberately. Two names sharing one object is harmless; two objects sharing one file is the bug. ### Practical uses * **Inspect what is loaded.** `sorted(sys.modules)` after startup shows exactly what your import graph pulled in, and how much of it you did not expect. * **Detect whether something is already imported** without importing it: `'mod' in sys.modules`. * **Understand startup cost.** Because the body runs once, import cost is a startup property, not a per-call one; `python -X importtime` attributes it per module. The mental model to carry into an interview is one sentence: *`sys.modules` is the identity map for modules, keyed by name, populated before execution, and everything odd about repeated imports follows from that.*

  • Why does replacing sys.modules['config'] not change a name that was imported with from config import SETTINGS?
    Because `from ... import ...` copies a reference into the importing namespace at import time. `SETTINGS` is now an independent binding to the original object and no longer consults the module. Swapping the `sys.modules` entry only affects imports that happen afterwards. To change what existing code sees you must set the attribute on the module object that code actually looks it up on.
  • Why is the module object put into sys.modules before its body executes?
    So that an import cycle terminates. If the body imports a module that imports back, the second import finds the entry already present and returns the partially initialised module instead of executing the file again, which would recurse. The trade-off is that the module can be observed half-built. If the body raises, the machinery removes the entry so a failed import is not left cached.
  • What does the presence of a name in sys.modules actually guarantee?
    Only that an import of that name has begun in this interpreter and produced a module object. It does not guarantee the body finished — during a cycle the module can be present but incomplete — and it says nothing about which file it came from. When that matters, check the object's `__file__` rather than the key.

saying these in an interview costs you the question

  • Says every import statement re-reads and re-executes the file
  • Thinks deleting from sys.modules unloads the old module
  • Believes the module object is cached only after its body finishes
  • Claims from-imports keep tracking the module they came from
  • Assumes two names in sys.modules always mean two modules

context