skip to content

Why does importlib.import_module on a user-supplied name amount to arbitrary code execution?

level: middleimportance: must knowfreq 55%

answer

  1. An import is not a lookup
  2. The module body runs top to bottom
  3. sys.path decides what the name resolves to
  4. First import executes, later ones are cached
  5. Map request names to fixed module names

basics

~20 s

Importing a module executes that module's top-level code. If the caller names the module, the caller chooses which code runs, and search order plus every installed distribution decide what that name resolves to. Import from a fixed allowlist instead.

solid answer

~50 s

An import is not a lookup, it is an execution: the first time a name resolves, Python finds the source, compiles it and **runs the whole module body**, then caches the result in `sys.modules`. So `importlib.import_module(user_string)` hands the caller a choice of which code runs in your process, with your privileges. The reachable set is not just your own modules — it is every module on `sys.path`, meaning the whole standard library and every installed distribution, plus anything an attacker can write into a directory that is on the path. Dotted names walk into packages, and importing a submodule runs the parent package's `__init__` first. The builtin `__import__` has the same execution semantics with more awkward arguments. Resolve the request name against an explicit mapping of permitted plugin names to module names, and import only those.

code

python · 12 lines
python
import importlib
import pathlib
import sys
import tempfile

tmp = tempfile.mkdtemp()
pathlib.Path(tmp, "plugin_demo.py").write_text("print('module body ran during import')\n")
sys.path.insert(0, tmp)

importlib.import_module("plugin_demo")     # the body runs here
importlib.import_module("plugin_demo")     # cached in sys.modules, body does not rerun
print("plugin_demo" in sys.modules)

go deeper

for a junior

Recall the one-line fact: importing a module runs its top-level code the first time, and the result is cached. That is why letting a request choose the module name is letting it choose code.

for a middle

Explain the search order and the sys.modules cache, and why the reachable set is everything on sys.path rather than your own package. Show the mapping from permitted request names to fixed module names.

for a senior

Demonstrate the operating judgement: keep writable and attacker-influenced directories off sys.path, pin what is installed, and know that error messages from a failed import leak an inventory of the environment.

for a principal

Own where the plugin trust boundary sits. Entry-point discovery over whatever is installed is reasonable for a developer tool and wrong for a network-facing service; decide that explicitly and write it down.

## Import is execution The mental model that makes this dangerous is "importing a module looks something up". It does not. On the first import of a given name, CPython's import machinery locates a source or bytecode file, compiles it if needed, creates a module object, and **executes the module body top to bottom** in that module's namespace. Class statements run, decorators run, module-level function calls run, and any side effect the file's author wrote — opening a socket, writing a file, spawning a process — happens right there inside your interpreter, in your process, with your user's privileges. Only afterwards is the module object stored in `sys.modules`, which is why a second `importlib.import_module` of the same name is cheap and does *not* re-execute the body. That cache is also why "I imported it and nothing bad happened, so it is safe" is not a test: the dangerous moment already passed on the first import. ## What a user-chosen name can resolve to `importlib.import_module("thing")` searches the entries of `sys.path` in order. So the reachable set is much larger than the plugin package you had in mind: * **The whole standard library**, including modules whose import-time or subsequent behaviour an attacker can chain into something useful. * **Every installed distribution** in the environment, including transitive dependencies nobody on the team has read. Note that the *installed distribution* name and the *importable* name often differ, so an inventory of what is installed is not an inventory of what is importable. * **Anything in a directory that happens to be on the path.** If the process ever has a writable or attacker-influenced directory on `sys.path` — an upload directory, a working directory the process was launched from, a path assembled from configuration — then a file dropped there becomes an importable module. Dotted names make it worse in a subtle way: `import_module("pkg.sub")` imports `pkg` first, running its `__init__` body, then `pkg.sub`. A relative import (`import_module(".sub", package="pkg")`) still resolves to a module whose body executes. And a name that resolves to nothing raises `ModuleNotFoundError` — a probing oracle that tells an attacker exactly what your environment has installed. The builtin `__import__` is the older, lower-level entry point with the same execution semantics and confusing return value (it returns the *top-level* package for a dotted name unless you pass `fromlist`). Preferring `importlib.import_module` is a readability improvement, never a security one. ## The safe shape The fix mirrors the attribute case: the *set* of importable targets is fixed by your code, and only the *choice* among them comes from the request. ```python PLUGINS = { "csv_export": "myapp.plugins.csv_export", "pdf_export": "myapp.plugins.pdf_export", } def load(requested: str): target = PLUGINS.get(requested) if target is None: raise ValueError(f"unknown plugin: {requested!r}") return importlib.import_module(target) ``` Things that are *not* fixes, and are worth being able to reject out loud in an interview: * **Validating the string is an identifier.** `str.isidentifier` says the name is well-formed. It says nothing about which code that well-formed name runs. * **Prefixing a package name.** `import_module("myapp.plugins." + name)` looks constrained, but a name containing dots walks deeper, and if the plugins directory is writable the constraint is cosmetic. * **Checking the module exists first.** `importlib.util.find_spec` tells you a name resolves and *where* it resolves to without executing the body — genuinely useful for a clean error message, and worth knowing — but if you then import whatever it found, you have only moved the decision. * **Catching `ImportError`.** Error handling is not authorisation. ## Related operational points Entry-point-style plugin discovery, where installed distributions advertise importable names, moves the trust boundary to "whatever is installed in this environment" — a real and reasonable model for a developer tool, and a bad one for a server that accepts a plugin name over the network. Keep the two straight, pin what is installed, and keep writable directories off `sys.path` for any process that handles untrusted input. Deleting an entry from `sys.modules` to force a re-import re-runs the body, so cache invalidation is another way to trigger execution of code you thought had already settled. ## A note on where the name comes from Two sources deserve different treatment. A module name from **configuration** — a settings file or an environment variable set at deploy time — is trusted to the same degree as the deployment itself, and importing it is ordinary plugin loading; the question there is only whether the config source is genuinely privileged. A module name from a **request, an uploaded document, a queue message or a database row that a user can influence** is untrusted, and the allowlist is mandatory. The most common real incident is the third case: a name that was configuration when the code was written, and became user-reachable later when an admin screen started writing that same config row. That is why the review question is about the name's provenance chain, not about the call site in isolation.

  • Does validating the name with str.isidentifier make the import safe?
    No. It only proves the string is a syntactically valid name. A valid name still selects among every module reachable on sys.path — the standard library, every installed distribution, and anything in a writable directory on the path — and importing it runs that module's body. Well-formedness is not authorisation; only a fixed mapping from request names to module names is.
  • How would you check that a plugin name resolves without running its code?
    importlib.util.find_spec returns a module spec, or None, without executing the module body, so it is a good way to produce a clean 'unknown plugin' error and to see which file a name would load. It is a diagnostic, not a control: if you then import whatever it located, the execution decision is unchanged. The allowlist still has to come first.
  • Why does re-importing the same module name not re-run its top-level code?
    The first import stores the module object in sys.modules, and subsequent imports of that name return the cached object. That is what makes module-level state a singleton per interpreter. It also means the security-relevant moment is the first import, so 'we import it all the time and it is fine' proves nothing — and deleting the sys.modules entry to force a reload runs the body again.
  • Is the builtin __import__ any different from importlib.import_module here?
    Not in the way that matters: both run the module body. __import__ is the lower-level hook with a clumsier signature — for a dotted name it returns the top-level package unless fromlist is supplied — so importlib.import_module is preferred for readability and correctness. Neither adds any validation of the name.

saying these in an interview costs you the question

  • Thinks importing only defines names without running code
  • Believes a valid identifier check makes the import safe
  • Says only your own package is reachable by name
  • Treats catching ImportError as a security control
  • Assumes find_spec succeeding means the module is trusted
  • Ignores that dotted names walk into parent packages first

context