skip to content

What does an importlib.machinery.ModuleSpec carry, and how does a loader use it?

level: middleimportance: should knowfreq 22%

answer

  1. The value that crosses the two phases
  2. Name plus loader, origin, search locations
  3. Create the module, then execute it
  4. Registered before the body finishes
  5. create_module returning None means default

basics

~20 s

A ModuleSpec carries the module's name, the loader that can execute it, its origin, and for a package its submodule_search_locations. The machinery builds a module object from the spec, registers it under its name, then calls loader.exec_module(module) to run the body.

solid answer

~40 s

`find_spec` answers *where and by whom*, not *what*. The `importlib.machinery.ModuleSpec` it returns holds `name`, `loader`, `origin` (the filename or other source label), `submodule_search_locations` (non-`None` only for a package, and it becomes the package's `__path__`), `parent`, `cached` and `has_location`. The machinery then hands the spec to `importlib.util.module_from_spec`, which calls `loader.create_module(spec)` — returning `None` there means *give me a default module object* — and pre-sets `__name__`, `__spec__`, `__file__` and `__path__` on it. Only then is the module registered under its name and `loader.exec_module(module)` called; `exec_module` returns nothing and works purely by mutating the module's namespace. If the body raises, the machinery removes the registration so a later import retries rather than handing out a half-built module.

code

pycon · 10 lines
pycon
>>> import importlib.util
>>> spec = importlib.util.find_spec("json")
>>> spec.name
'json'
>>> type(spec.loader).__name__
'SourceFileLoader'
>>> spec.submodule_search_locations is not None
True
>>> spec.parent, spec.has_location
('json', True)

go deeper

for a junior

Recall the shape: something locates the module and describes it, something else runs its code. Knowing that every imported module exposes a __spec__ you can print is enough at this level.

for a middle

Explain the fields you would actually use — name, loader, origin, submodule_search_locations — and the create-then-execute sequence, including what create_module returning None means.

for a senior

Demonstrate the consequences: the module is registered before its body finishes, a failed body is unregistered again, and a wrapper loader can intercept exec_module without touching how the module was found.

for a principal

Frame the split as an interface boundary. Argue when your team should implement a loader at all, versus generating real source files, and what the debuggability cost of a synthesized module is for everyone who reads a traceback later.

## The spec is the contract between the two phases CPython's import machinery is split into finding and loading, and `importlib.machinery.ModuleSpec` is the value that crosses the boundary. A finder never creates a module; it fills in a spec and hands it back. Everything the loading phase needs is on that object. ## What is on a spec * `name` — the fully qualified dotted name, which becomes the module's `__name__`. * `loader` — the object that will actually execute the module. This is the only mandatory partner of `name`. * `origin` — a human-readable label for where the module came from: a filename for source on disk, a path inside an archive, or something like `built-in` for a module compiled into the interpreter. When the spec has a real location, `origin` becomes `__file__`. * `submodule_search_locations` — `None` for a plain module; a (possibly empty) sequence for a package. It becomes the package's `__path__`, and it is exactly what gets passed as the `path` argument when a submodule of that package is looked up. * `parent` — the containing package's name, derived from `name` and whether the spec is a package. * `cached` — where compiled bytecode lives, surfaced as `__cached__`. * `has_location` — whether `origin` refers to something loadable from a real location. * `loader_state` — a free slot a finder may use to pass private information to its own loader, for example a byte range or an already-open handle. You rarely build one by hand. `importlib.util.spec_from_loader(name, loader)` and `importlib.util.spec_from_file_location(name, path)` cover most cases, and `importlib.util.find_spec(name)` runs the normal search and gives you the spec the machinery would have used. ## From spec to module The loading phase is deliberately two-step. First the module object is created: `importlib.util.module_from_spec(spec)` calls `spec.loader.create_module(spec)`. A loader that has no special needs returns `None` there, which means *use the ordinary module type*; a loader for an extension module or an exotic module subclass returns its own object instead. `module_from_spec` then wires the attributes the module needs before any of its code runs — `__name__`, `__spec__`, `__file__` when the spec has a location, and `__path__` for a package. Second, the module is registered under `spec.name` and only then is `spec.loader.exec_module(module)` called. `exec_module` returns nothing useful; its whole job is to populate `module.__dict__`, which for a source loader means compiling the source and executing the resulting code object with the module's namespace as globals. The ordering matters and is worth stating in an interview: the module is visible under its name *before* its body has finished, which is what allows a module that imports itself recursively to see the in-progress object instead of executing twice. It is also why a module can be observed half-initialized. If `exec_module` raises, the machinery deletes the registration, so the failed module does not linger and the next import genuinely retries. ## Why the split exists Separating create from execute is what makes several features possible at all. A loader can hand back a specially prepared module object and let the standard machinery run the body. A wrapper loader can accept `exec_module` and choose to do nothing yet, deferring the real execution — that is precisely how deferred-execution loaders in `importlib.util` are built. Reloading reuses the same module object and re-runs only `exec_module`, so existing references to the module keep working. And a loader that produces no Python source at all — synthesizing functions, wrapping an interface definition — plugs in at exactly one method. ## Version notes The spec-based protocol has been the only protocol since **Python 3.12** removed the legacy `find_module`/`load_module` pair; on 3.14 a loader that defines only the old methods is not used. `ModuleSpec` itself is stable across 3.10 to 3.14. Every imported module exposes the spec it was built from as `__spec__`, which is the quickest way to see all of this from a REPL.

  • What does importlib.util.module_from_spec do that constructing a bare module object does not?
    It gives the loader its say and wires the module before any code runs. It calls `spec.loader.create_module(spec)` so a loader can return its own object, then sets `__name__`, `__spec__`, `__file__` when the spec has a location, and `__path__` for a package. A hand-built module object has none of that, so its body would execute against an unwired namespace.
  • Why is the module registered under its name before exec_module runs rather than after?
    So the name resolves to the in-progress object while the body is still executing. Without it, a module reached again during its own execution would be executed a second time and you would end up with two module objects and duplicated top-level state. The trade-off is that the module can be observed partially initialized; if the body raises, the machinery removes the registration again.
  • What marks a spec as describing a package rather than a plain module?
    A non-`None` `submodule_search_locations`. The machinery copies it onto the module as `__path__`, and that value is what gets passed as the `path` argument when a submodule of the package is searched for. An empty list is still a package — namespace-style packages simply start with nothing in it.

saying these in an interview costs you the question

  • Thinks find_spec itself imports or executes the module
  • Says exec_module returns the module object
  • Believes create_module must always build a module
  • Cannot say what makes a spec a package
  • Assumes the module is registered only after a successful body

context