skip to content

sys.path and sys.modules

The search path Python walks to find a module and the cache that guarantees it loads once. Most 'importing the wrong thing' bugs are a local file shadowing a stdlib name; most 'my edit did nothing' is the cache.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why does a file named random.py in your project break import random?

level: juniorimportance: must knowfreq 72%

answer

  1. Which one does the interpreter find first?
  2. The search path is ordered
  3. Your script's directory leads the list
  4. sys.path[0] precedes the standard library
  5. Print the module's __file__ to confirm

basics

~10 s

Python searches sys.path in order, and the running script's own directory sits at the front, ahead of the standard library. Your random.py is found first, so import random binds your file instead.

solid answer

~40 s

`import random` means "bind the first thing named `random` found along `sys.path`", not "load the standard library". When you run `python app.py`, CPython puts the directory containing `app.py` at `sys.path[0]` — an absolute path since 3.11 — and the standard-library directories and site-packages come after it. A sibling `random.py` therefore wins, and the import succeeds silently; the failure surfaces later as `AttributeError: module 'random' has no attribute 'randint'`, sometimes raised from inside library code that imported the same name. Confirm it by printing `random.__file__` and checking `sys.stdlib_module_names`. The fix is to rename the file, and to restart the process, since the wrong module object stays in `sys.modules`. Prevention: keep application code inside a package and never give a top-level file a standard-library name. Modules compiled into the binary, listed in `sys.builtin_module_names`, are immune.

code

python · 13 lines
python
import sys
import tempfile
import pathlib

tmp = pathlib.Path(tempfile.mkdtemp())
(tmp / "json.py").write_text("VALUE = 'this is the local file, not the stdlib'\n")
sys.path.insert(0, str(tmp))

import json

print(json.__file__)
print(json.VALUE)
print("json" in sys.stdlib_module_names)

go deeper

for a junior

Recall the one-line cause: imports search sys.path in order and your own directory is first, so a same-named file wins. Be ready to say you would print the module's file to prove which file was loaded.

for a middle

Explain the mechanics: what sits at sys.path[0] for a script versus an interactive session, why the import succeeds silently and fails later as an AttributeError, and why sys.builtin_module_names is the exception.

for a senior

Show the diagnosis under pressure — file, sys.path[:3], loading traces — and give the structural prevention: code in a package, project installed, entry points run properly, -P or PYTHONSAFEPATH where you want it guaranteed.

for a principal

Own the policy angle: how repository layout and the choice to make the working directory importable decide whether this class of bug is possible at all, and what you standardise so it cannot recur across many services.

## The import statement does not know what "standard library" means `import random` is not a request for a specific, blessed module. It is a request to resolve the *name* `random` through the import machinery, and for anything that is not compiled into the interpreter binary that resolution walks `sys.path` — an ordinary Python list of directory strings — and takes **the first match**. Nothing in that walk gives the standard library priority. The standard library is simply a set of directories that happen to sit somewhere in that list, and by default they sit *after* your own code. ## Where your file gets its head start When you run `python app.py`, CPython prepends the directory containing `app.py` to `sys.path`. Since Python 3.11 that entry is an absolute path (earlier versions inserted it as written, which made it change meaning if the process chdir'd). For an interactive session or `python -c`, the leading entry is `''`, which means "the current working directory, resolved at each import". Either way, entry zero is *your* directory. Everything the standard library ships — `random`, `json`, `types`, `select`, `logging`, `email`, `code`, `test` — is reachable only from later entries, so a same-named file next to your script outranks it. The one class of exception is modules compiled into the interpreter executable itself. Those names are listed in `sys.builtin_module_names` and are resolved before the path search happens at all, which is why a local `sys.py` is harmless while a local `random.py` is not. ## Why it is so confusing in practice The shadowing import does not fail. There is no warning, no collision error, no hint that anything unusual happened; the name binds and execution continues. The symptom appears later and somewhere else: * `AttributeError: module 'random' has no attribute 'randint'` — your file simply does not define what the caller expected. * A traceback that points *inside* library code. Absolute imports in third-party or standard-library modules go through the same `sys.path`, so a module you never touched can end up importing your file. * An import-time side effect firing at a bizarre moment, because your file's top-level code runs the first time anyone imports that name. * A test suite that passes under one runner and fails under another, because the two runners put different directories at the front of the path. The worst version is a file whose name matches a module the standard library imports internally during startup or during the first use of some feature — the failure then looks like the interpreter itself is broken. ## Diagnosing it in ten seconds Ask the module where it came from: ```pycon >>> import random >>> random.__file__ '/home/me/project/random.py' ``` A `__file__` under your project instead of under the interpreter's `lib` directory settles it. Two supporting checks: `'random' in sys.stdlib_module_names` tells you the name belongs to the standard library on this version (that attribute exists from 3.10), and `python -v` prints where each module was loaded from as it loads. Printing `sys.path[:3]` shows you which directory is winning. ## Fixing and preventing it The fix is to rename the file — and to *restart the process*. Renaming does not help a running interpreter, because the wrong module object is already sitting in `sys.modules` and every subsequent `import random` will keep returning it. Prevention is structural rather than vigilance: * **Put application code in a package.** If your modules live in `myapp/`, the top-level names on the path are `myapp` and nothing else, so the collision surface shrinks to one name you chose. * **Do not develop from a directory full of loose top-level modules**, especially one that is also the current working directory of your tooling. * **Use `-P`, or `PYTHONSAFEPATH=1`** (both 3.11 and later) when you want CPython to stop prepending the script directory or the current directory entirely; imports then resolve only from installed locations. * **Check the name before you create the file.** `python -c "import sys; print('email' in sys.stdlib_module_names)"` costs nothing. The same mechanism bites one level up. A *directory* on the path named like an installed distribution's importable package shadows that installed package just as completely, and because installed distributions and importable packages often have different names, the message you get can name something you never installed. The rule that resolves every one of these cases is the same: the import machinery searches `sys.path` in order and stops at the first hit, so whoever is earlier on the path wins.

  • Why does renaming the offending file sometimes appear not to fix anything?
    Because the process is still running. The first import stored the wrong module object under that name in `sys.modules`, and every later `import` of the same name returns the cached object without touching the filesystem. Renaming the file only changes what a *new* interpreter would find, so the fix takes effect on restart. In a REPL or a notebook kernel, restart the kernel rather than trying to un-import.
  • Why does a local sys.py not shadow the sys module the way random.py shadows random?
    Because `sys` is compiled into the interpreter binary. Names in `sys.builtin_module_names` are resolved before the path-based search runs, so no file on `sys.path` can intercept them. Everything implemented as a `.py` file in the standard library — which is most of it — has no such protection and is shadowable by any earlier path entry.
  • How would you stop this class of bug for a whole project rather than one file at a time?
    Structurally. Keep application modules inside one package directory so only that single top-level name is exposed, install the project into the environment rather than relying on the current directory being importable, and run entry points as installed commands or with `python -m`. Where you want the guarantee enforced, run with `-P` or `PYTHONSAFEPATH=1` so CPython never prepends the script or working directory at all.

sys.path is a stack of phone books searched top to bottom, and the interpreter dials the first matching name it finds. Dropping your own one-page phone book on top means every call for "random" reaches you, not the person everyone meant.

saying these in an interview costs you the question

  • Claims the standard library always takes priority over local files
  • Says Python raises an error or warning on a name collision
  • Thinks site-packages is searched before the script's own directory
  • Treats it as a missing dependency and reinstalls something
  • Believes renaming the file fixes an already-running process
  • Confuses it with a circular import failure

context

open as a page

What does Python do with sys.modules when a module is imported a second time?

level: middleimportance: must knowfreq 58%

basics

~20 s

It finds the module already there and stops. The import statement checks sys.modules by name first, and on a hit it only binds the name in the importing namespace; the module body runs exactly once per interpreter.

open as a page

How does CPython assemble sys.path when the interpreter starts?

level: middleimportance: should knowfreq 48%

basics

~10 s

In order: an entry for the script's directory or the current directory, then PYTHONPATH entries, then the interpreter's own standard-library directories, then the site-packages directories that the site module appends while processing .pth files.

open as a page

Why can sys.modules hold the same source file twice under two different names?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Because the cache is keyed by module name, not by file. A file reachable under two names — typically when a package and its parent are both on sys.path — misses the cache twice and becomes two modules.

open as a page