skip to content

What does importlib.import_module do that a plain import statement cannot?

level: juniorimportance: should knowfreq 42%

answer

  1. The statement bakes the name in
  2. A function that accepts a string
  3. Top-level package versus leaf module
  4. importlib.import_module, not __import__
  5. fromlist is why __import__ surprises you

basics

~20 s

importlib.import_module takes the module name as a string computed at runtime and returns that exact module object. The import statement needs the name spelled out in source code, and the low-level import builtin returns the top-level package instead of the submodule you asked for.

solid answer

~40 s

The `import` statement is syntax: the name is baked into the bytecode at compile time, so you cannot import a module chosen by a config file or a CLI argument with it. `importlib.import_module("xml.etree.ElementTree")` takes that dotted name as an ordinary string and returns the **leaf** module object. The older `__import__` builtin is the hook the statement itself compiles down to; called directly with one argument it returns the *top-level* package (`xml`), not the submodule, which is a classic source of confusion — you only get the leaf if you pass a non-empty `fromlist`. The docs say to prefer `import_module`. Both go through the normal machinery, so the module is cached in `sys.modules` and its top-level code runs exactly once; a second call is just a dictionary hit.

code

pycon · 7 lines
pycon
>>> import importlib
>>> importlib.import_module("xml.etree.ElementTree").__name__
'xml.etree.ElementTree'
>>> __import__("xml.etree.ElementTree").__name__
'xml'
>>> __import__("xml.etree.ElementTree", fromlist=["*"]).__name__
'xml.etree.ElementTree'

go deeper

for a junior

Be ready to say that the import statement needs a literal name and that importlib.import_module accepts a string instead. Knowing the function name and that it returns a module object is enough at this level.

for a middle

Explain the mechanics: the compiler bakes the name into bytecode, import_module returns the leaf while a bare import returns the top-level package, and the result is cached in sys.modules so the body runs once.

for a senior

Show the production instincts around it — allowlist names that come from config or requests, distinguish ModuleNotFoundError from an ImportError raised inside the module, and remember invalidate_caches after generating files.

for a principal

Own the policy question: which names a service is allowed to import at all, whether dynamic imports happen at startup or lazily on first use, and what that choice costs you in startup time versus late-failure risk.

### The statement is syntax, not a call When you write `import xml.etree.ElementTree`, the compiler turns the dotted name into a *constant* in the bytecode. That is fine when you know the name while writing the file, and useless the moment the name arrives from somewhere else: a `[plugins]` section of a config file, a `--parser` command-line flag, a column in a database, an environment variable. For that you need a function that accepts a `str`. `importlib.import_module(name)` is that function. It performs the same work the statement performs — consult `sys.modules`, otherwise search the finders on `sys.meta_path`, create the module, execute its top-level code — and then it **returns the module named by the last component**. `importlib.import_module("xml.etree.ElementTree")` hands you the `ElementTree` module object itself. ### It returns, it does not bind One more difference trips people up in the REPL. The statement `import csv` does two things: it imports, and it *binds* the name `csv` in the current namespace. `importlib.import_module("csv")` only does the first — it returns the module object and binds nothing. If you want a name, you assign one yourself, and if you want to expose the module under a computed name you assign into a namespace explicitly. That is a feature for a loader: a plugin host usually wants the object in a registry keyed by a logical name, not a module-level global nobody wrote. ### Why not `__import__`? `__import__` is a builtin and is the actual hook the `import` statement uses; its signature is `__import__(name, globals=None, locals=None, fromlist=(), level=0)`. Because it exists to *implement* the statement rather than to be pleasant, it mirrors the statement's semantics: `import xml.etree.ElementTree` binds the name `xml` in your namespace, so `__import__("xml.etree.ElementTree")` returns `xml`. Getting the leaf requires the un-obvious `__import__("xml.etree.ElementTree", fromlist=["*"])`, which mirrors `from ... import ...`. The standard library documentation explicitly recommends `importlib.import_module` over calling `__import__` directly; the only real reason to touch `__import__` is that you are *replacing* it to hook the import system globally. ### What the call actually does Three consequences are worth internalising. **It is cached.** The first call executes the module body; every later call for the same name is effectively a `sys.modules` lookup. So `import_module` in a hot loop is cheap, but it is *not* free the first time and it is not idempotent in the sense of side effects: whatever the module does at import time (opening a connection, registering itself, spawning a thread) happens once, when the first caller asks for it. **It executes arbitrary code.** Importing a module runs that module's top level. If the name comes from user-controlled input, you have handed the user code execution over anything importable on `sys.path`. Validate against an allowlist of permitted names, or at minimum require a fixed prefix; never pass a raw request field into `import_module`. **It raises `ModuleNotFoundError`.** That is a subclass of `ImportError`, so `except ImportError` catches both. Distinguish the two cases carefully: `ModuleNotFoundError` means the *name* could not be located, while a plain `ImportError` raised from inside the module body means the module was found and its own imports failed. Swallowing both under one `except ImportError: pass` is how a genuine bug inside an optional dependency turns into a silent "plugin not installed". ### The supporting cast `import_module` takes a second argument, `package`, used as the anchor when `name` begins with a dot; that relative form is rarely what you want when the name came from configuration, because configuration should carry fully-qualified names. If you want to know whether a module exists *without* executing it, use `importlib.util.find_spec(name)`, which returns `None` when nothing can supply the module and raises `ModuleNotFoundError` if an intermediate parent package is missing. That is the right check for a `--help` output that lists available backends, or for a feature flag that should not pay the import cost. If your program writes a `.py` file and then imports it in the same run, call `importlib.invalidate_caches()` first: the import machinery caches directory listings for speed, and a file created after that cache was populated is invisible until the caches are dropped. ### Practical shape A loader usually looks like a thin wrapper that adds the two things `import_module` deliberately does not do — validation and a good error message: ```python import importlib ALLOWED = {"json", "csv", "pickle"} def load_codec(name): if name not in ALLOWED: raise ValueError(f"unknown codec {name!r}; expected one of {sorted(ALLOWED)}") return importlib.import_module(name) ``` The interviewer is usually checking two things: that you know the `import` statement is compile-time syntax, and that you reach for `importlib.import_module` rather than `__import__` or, worse, `eval("import " + name)` — which is not even valid, since `import` is a statement and `eval` takes expressions.

  • Why does calling __import__ with a dotted name give you the wrong module object?
    Because `__import__` mirrors the statement it implements. `import a.b.c` binds only `a` in your namespace, so `__import__("a.b.c")` returns `a`. Passing a non-empty `fromlist` switches it to the `from a.b import c` behaviour and returns the leaf. `importlib.import_module` always returns the leaf, which is why it is the recommended API.
  • How would you check that a module is importable without paying the cost of importing it?
    `importlib.util.find_spec(name)` runs the finders and returns a spec — or `None` if nothing can supply the module — without executing the module body. It still imports parent packages, and it raises `ModuleNotFoundError` if an intermediate parent is missing, so wrap it when the name is fully arbitrary.
  • What happens the second time your code calls importlib.import_module with the same name?
    It hits the `sys.modules` cache and returns the existing module object without re-executing anything. That makes repeat calls cheap, but it also means import-time side effects happen exactly once for the life of the process, no matter how many callers ask for the module.

The import statement is a street address printed on a form; import_module is asking a receptionist for whichever office you name aloud.

saying these in an interview costs you the question

  • Claims eval or exec is the normal way to import by name
  • Thinks __import__('a.b') returns the b submodule
  • Believes each import_module call re-executes the module body
  • Passes an unvalidated user string straight into import_module
  • Cannot name any importlib function at all

context