How does Python assemble a namespace package's __path__ from several sys.path entries?
answer
- The walk does not stop at a directory
- Portions collected in sys.path order
- Anything concrete ends the search
- __path__ recomputes when sys.path changes
- First portion wins a duplicate submodule
basics
~20 sPython walks every sys.path entry in order, records each same-named directory that has no init.py as a portion, and puts them all in path. A module file or a directory with init.py stops the walk instead.
solid answer
~40 sFor `import ingest`, each `sys.path` entry is checked in order. A directory `ingest/` with no `__init__.py` is *recorded* and the walk continues; a directory with `__init__.py`, or an `ingest.py`, is *used* and the walk stops. If the walk ends with only recorded directories, they become the namespace package's `__path__`, **in `sys.path` order**. Submodule lookups then search those portions in that order, so if two portions both contain `parser.py`, the earlier one wins and the other is unreachable. `__path__` is not a plain list: it is a dynamic object that recomputes from `sys.path`, so appending a new path entry makes that portion's submodules importable immediately, with no reload. This is what lets `ingest.parser` and `ingest.shipper` arrive from two separately installed distributions.
code
python · 16 linesimport pathlib
import sys
import tempfile
root = pathlib.Path(tempfile.mkdtemp())
for dist, submodule in (("wheel_a", "parser"), ("wheel_b", "shipper")):
portion = root / dist / "ingest"
portion.mkdir(parents=True)
(portion / f"{submodule}.py").write_text(f"OWNER = {dist!r}\n")
sys.path.append(str(root / dist))
import ingest
from ingest import parser, shipper
print(len(list(ingest.__path__)), parser.OWNER, shipper.OWNER)
print(ingest.__spec__.origin, ingest.__file__)go deeper
Know that a namespace package can be made of more than one directory, and that you can see them with list(pkg.path). You are not expected to recite the scan rules yet.
This is your level's core: explain the three checks per sys.path entry, why a portion does not stop the walk, and how the collected directories end up in path in path order.
Reason about the consequences — duplicate submodule names silently shadowing, one distribution's init.py collapsing the whole namespace, and the extra startup cost of a scan that can never short-circuit.
Treat a shared namespace as governed surface: decide who may own which submodule names, whether portions may be added at runtime at all, and how installs are verified so path order is never load-bearing by accident.
### The walk, step by step When a name is not already in `sys.modules`, the import system asks each finder on the meta path, and the path-based finder walks `sys.path` entry by entry. For a top-level name `ingest`, each entry gets three checks in a fixed priority: 1. `ingest/__init__.py` present → a regular package. **Terminate**: this is the answer. 2. `ingest.py` or a matching extension module → a module. **Terminate**. 3. A directory `ingest/` with no `__init__.py` → **record it as a portion and keep walking.** The same three checks apply within a single entry, which is why a `ingest.py` file next to a bare `ingest/` directory in the *same* directory wins: the namespace portion never terminates the search, so anything concrete beats it. Only when the walk exhausts `sys.path` with a non-empty list of recorded portions is the namespace package created, its `__path__` holding the portions in the order the entries appeared on `sys.path`. For a *sub*package the mechanics are identical, except the search runs over the parent's `__path__` rather than `sys.path`. So a nested namespace inside a namespace works the same way, one level down. ### __path__ is dynamic, and that matters A regular package's `__path__` is a plain list containing one directory. A namespace package's is a special object that remembers the parent path it was computed from and **recomputes the portions whenever that parent path changes**. Practically: import `ingest`, then append another directory containing `ingest/shipper.py` to `sys.path`, and `import ingest.shipper` works straight away — no `importlib.reload`, no cache busting. The same dynamism is why a `.pth` file or a path manipulation that runs at startup can add portions the first import never saw. It also means `__path__` is genuinely mutable state that other code can disturb. Mutating `sys.path` at runtime is already a smell; with namespace packages, it silently changes *which* files a package is made of. ### Ordering and collisions Submodule resolution walks `__path__` in order and takes the first hit, so portions must own **disjoint** submodule names. If two portions both ship `ingest/parser.py`, whichever came earlier on `sys.path` wins and the other is dead code — no warning, no error. That is the strongest argument for treating a shared namespace as a governed thing: one owner per submodule name, enforced by review or by CI rather than by hope. A subtler collision is the one that removes the namespace entirely. If any single distribution installs `ingest/__init__.py`, rule 1 fires at that entry and the whole namespace collapses to that one regular package: every other portion becomes unimportable, and the error surfaces as `ModuleNotFoundError: No module named 'ingest.shipper'` in code that has worked for months. Adding an `__init__.py` to a shared namespace is therefore a breaking change to every other distribution in it. ### Inspecting it `list(pkg.__path__)` is the direct read — the number of entries tells you how many portions merged, and the paths tell you which installed trees they came from. `pkg.__spec__.submodule_search_locations` is the same information on the spec, and `importlib.util.find_spec("pkg")` gets it without importing. When a submodule is missing, the useful question is not "is it installed?" but "is its directory in `__path__`?" — usually it is not, because a regular package terminated the walk early or the distribution landed on a `sys.path` entry that is not actually on the path in that environment. ### Cost Because a portion never terminates the search, resolving a namespace package touches **every remaining `sys.path` entry**, and each entry means a directory listing (cached, but still). On a short path this is noise. On a long path — many `.pth` additions, network or archive entries — it is a real startup cost, and it is a reason not to make top-level namespaces out of habit. A regular package, by contrast, short-circuits at the first entry that has it. ### The mental model Think of the walk as looking for one concrete answer and collecting consolation prizes on the way. Any concrete answer ends it. If there is no concrete answer, the consolation prizes are glued together, in path order, into one package whose contents are the union of the directories — and the union is recomputed, not frozen, every time the path underneath it changes.
- What happens if two portions of the same namespace both contain parser.py?The portion that came earlier in `sys.path` order wins, and the other file is simply unreachable — no warning is issued. Submodule lookup walks `__path__` in order and takes the first hit, so portions must own disjoint submodule names, and that ownership has to be enforced by convention or CI because the import system will not.
- Why can appending to sys.path make a new submodule importable without a reload?A namespace package's `__path__` is not a plain list; it is a dynamic object that recomputes its portions from the parent path whenever that path changes. So the next `import ingest.shipper` re-scans and finds the directory you just added. A regular package's `__path__` is a fixed one-element list and behaves nothing like this.
- Does the same scan apply to a namespace subpackage?Yes, one level down: the search for `ingest.subpkg` runs over `ingest.__path__` instead of `sys.path`, with the same three checks and the same priority. So a namespace nested inside a namespace merges its own portions across all the parent's directories.
saying these in an interview costs you the question
- Says the first matching directory wins and the scan stops
- Thinks __path__ is a plain list fixed at import time
- Believes duplicate submodules across portions raise an error
- Claims sys.path order does not affect submodule resolution
- Assumes adding __init__.py to one portion is harmless