How does Python import modules straight out of a .zip file on sys.path?
answer
- It is not a special case in import
- A path entry hook, not meta_path
- Each entry gets its own finder, cached
- One member table instead of many stat calls
- No compiled extensions from an archive
basics
~20 sA .zip entry on sys.path is handled by sys.path_hooks: the zipimport hook accepts the entry and returns a zipimport.zipimporter for it, cached in sys.path_importer_cache. That finder reads .py and .pyc members out of the archive, so no file is ever unpacked to disk.
solid answer
~40 s`PathFinder` does not read directories itself. For each `sys.path` entry it asks the callables in `sys.path_hooks`, in order, to produce a *path entry finder* for that string; a hook that cannot handle the entry raises `ImportError` and the next is tried, and the winner is memoised in `sys.path_importer_cache`. CPython ships two hooks: one yields a `zipimport.zipimporter` for an archive, the other a `FileFinder` for a directory. So `sys.path.insert(0, "lib.zip")` is all it takes — imports then resolve against the archive's directory table, and `__file__` becomes a path *inside* the archive. The consequences are what interviewers are after: `open(__file__)` fails, package data must be read through `importlib.resources`, compiled extension modules cannot be loaded from an archive at all, and changing the archive needs `importlib.invalidate_caches()` because the finder is cached.
code
python · 14 linesimport pathlib
import sys
import zipfile
pathlib.Path("bidmath.py").write_text("RATE = 42\n")
with zipfile.ZipFile("lib.zip", "w") as archive:
archive.write("bidmath.py")
sys.path.insert(0, "lib.zip")
import bidmath
print(bidmath.RATE)
print(bidmath.__file__)
print(type(bidmath.__spec__.loader).__name__)go deeper
Know that a sys.path entry does not have to be a directory: an archive works too, and imports from it behave normally. You are not expected to know the hook that makes it happen.
Explain the two layers — PathFinder on sys.meta_path, then sys.path_hooks producing a per-entry finder — and name at least one thing that does not work from an archive.
Show the consequences you would hit in production: package data through importlib.resources, no compiled extensions, no bytecode written back, and cache invalidation when the archive changes under a running process.
Own the packaging decision. Argue when a single archive is the right shipping artefact for pure-Python code versus a wheel and a virtual environment, and what it costs the team in debuggability and extension support.
## Two layers, not one Importing from an archive is not a special case bolted onto the import statement; it falls out of the ordinary machinery. `importlib.machinery.PathFinder` sits on `sys.meta_path` and is responsible for everything reachable through `sys.path`. For each entry on that list it needs an object that knows how to search *that kind of* location, and it builds one by calling each callable in `sys.path_hooks` with the entry string until one accepts it. A hook that does not recognise the entry raises `ImportError`, and `PathFinder` moves to the next. The accepted result is stored in `sys.path_importer_cache` keyed by the entry, so the hook runs once per location, not once per import. A stock CPython 3.14 has exactly two hooks: `zipimport.zipimporter`, which accepts a path naming a zip archive, and the hook that produces a `FileFinder` for a directory. That is the whole of the mechanism — the archive support is a path hook and nothing more, which is also the template for adding a location type of your own. ## What zipimport can and cannot do A `zipimport.zipimporter` reads the archive's central directory once and serves `.py` and `.pyc` members from it, including packages: a member `pkg/__init__.py` makes `pkg` importable, and a nested archive path works as a `sys.path` entry too (`lib.zip/subdir` is valid). What it cannot do is just as important: * **No compiled extension modules.** A `.so` or `.pyd` has to be handed to the operating system's dynamic loader, which needs a real file on the filesystem. Archives can only carry pure Python. Tools that ship a single archive work around this by unpacking extensions to a temporary directory first. * **No bytecode caching.** Nothing can be written back into the archive, so a `.py` member is recompiled on every process start unless a matching `.pyc` was placed in the archive when it was built. For a startup-sensitive process, pre-compiling into the archive is the difference that matters. * **No filesystem access to the source.** `__file__` is set to something like `lib.zip/tools.py`. It reads well in a traceback, and tracebacks do show source lines because the loader can supply them, but it is not a path any filesystem call can open. ## The data-file trap The usual way a working package breaks when it is moved into an archive is package data. Code that does `open(os.path.join(os.path.dirname(__file__), "table.json"))` works from a directory and fails from an archive, because that path does not exist. The portable form is `importlib.resources`: `importlib.resources.files("pkg").joinpath("table.json").read_text()` works for both, and when a library genuinely insists on a real filesystem path, `importlib.resources.as_file()` materialises one for the duration of a context manager. Writing package access this way from the start is what keeps the archive option open. ## Caching and invalidation Because the finder for a path entry lives in `sys.path_importer_cache`, and because the archive's directory table is read once, adding a member to an archive that is already on `sys.path` does not make it importable in the running process. `importlib.invalidate_caches()` tells every finder that caches state to drop it, which is the supported way to make a newly written module visible; dropping the specific entry from `sys.path_importer_cache` is the blunter version. This matters for anything that generates code at runtime and for test suites that build fixtures on the fly. ## When it is worth it A single archive gives you one artefact to copy and verify instead of a directory tree, and imports from it are often *faster* than from a directory because one directory table replaces many filesystem lookups. It costs you extension modules, writable bytecode caching and naive data-file access. That trade is usually the right one for a pure-Python internal tool, and the wrong one for an application with compiled dependencies. ## Version notes This mechanism is stable across Python 3.10 to 3.14: two default path hooks, `zipimport.zipimporter` for archives and a directory finder for everything else.
- How do you read a data file that ships inside a package imported from an archive?Through `importlib.resources`, not `open()`. `importlib.resources.files("pkg").joinpath("table.json").read_text()` works whether the package lives in a directory or an archive, because the traversable it returns is supplied by the loader. When a third-party library demands a real filesystem path, wrap the resource in `importlib.resources.as_file()`, which materialises a temporary file for the duration of the context manager.
- You add a module to an archive that is already on sys.path and the import still fails. Why?Two layers of caching. The path entry finder for that archive is memoised in `sys.path_importer_cache`, and the finder itself read the archive's member table once. Call `importlib.invalidate_caches()` so every cached finder drops its state, or remove the entry from `sys.path_importer_cache` outright, and the archive is re-read on the next import.
- Why can a compiled extension module not be imported directly from an archive?Loading one means asking the operating system's dynamic loader to map a shared library, and that API takes a filesystem path — there is no supported way to hand it bytes from an archive. Distributions that must ship extensions extract them to a temporary directory first and put that directory on the search path, which reintroduces the unpacking step archives were meant to avoid.
saying these in an interview costs you the question
- Thinks Python unpacks the archive to a temporary directory
- Expects open(__file__) to work for a module in an archive
- Believes compiled extension modules load from an archive
- Confuses sys.path_hooks with sys.meta_path
- Assumes bytecode is cached back into the archive
- Forgets the path entry finder is cached per entry