Why does `import xml` leave `xml.etree` an AttributeError while `import xml.etree` works?
answer
- A parent import does not pull children
- Dotted names resolve one step at a time
- The child is bound onto the parent object
- Someone else's import hides the bug
- sys.modules key versus parent attribute
basics
~20 sImporting a package does not import its submodules. Plain import xml runs only xml/init.py; the name etree becomes an attribute of the xml module object only once xml.etree is itself imported, by you or by someone else.
solid answer
~40 sA dotted import is resolved left to right: `import xml.etree.ElementTree` imports `xml`, then `xml.etree`, then `xml.etree.ElementTree`, and after each step the child module is cached in `sys.modules` under its full dotted name **and bound as an attribute on its parent module object**. A bare `import xml` performs only the first step, so `xml.etree` simply is not an attribute yet — unless `xml/__init__.py` imports it itself. The nasty part is that the attribute binding is process-global and permanent: once any module anywhere imports `xml.etree`, every other module sees `xml.etree` resolve fine. Code that relies on that works until import order changes. Import the submodule you actually use.
code
pycon · 8 lines>>> import xml
>>> xml.etree
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
AttributeError: module 'xml' has no attribute 'etree'
>>> import xml.etree.ElementTree
>>> xml.etree.ElementTree.__name__
'xml.etree.ElementTree'go deeper
Recall the rule and the fix: importing a package does not import what is inside it, so import the specific module you call, such as the submodule itself rather than only its parent package.
Explain the two bindings each import step performs — the sys.modules entry under the full dotted name and the attribute set on the parent module object — and why from-import behaves differently from attribute access.
Be ready to diagnose the field version of this: code that works in the service and fails in an isolated test because it depends on an import performed elsewhere. Show how you find and remove that ordering assumption.
Own it as a convention and a check: explicit imports in the module that uses them, no reliance on a parent package exposing submodules it does not import, and awareness that re-imports in spawned worker processes expose the assumption.
## What the import statement actually does `import a.b.c` is not one operation. The import system walks the dotted name left to right and imports each prefix in turn: `a`, then `a.b`, then `a.b.c`. For each step it checks `sys.modules` for the full dotted key, and on a miss it finds, loads and executes that module. Two bindings then happen for each child: 1. the module object is stored in `sys.modules["a.b"]`, and 2. it is **set as an attribute named `b` on the parent module object `a`**. That second binding is the one candidates forget, and it is what makes `a.b` work as an expression afterwards. Note also what the statement binds locally: `import a.b.c` binds only the top-level name `a` in the importing namespace — you reach `c` by attribute traversal. `import a.b.c as name` binds only `name`. ## Why a bare parent import is not enough `import xml` runs step one and stops. Nothing has created `xml.etree`, so attribute lookup on the `xml` module object raises `AttributeError: module 'xml' has no attribute 'etree'`. A package's `__path__` tells the import system where its submodules *live*; it does not import them. Eager loading would mean importing a package always dragged in every module beneath it, which for a large package is exactly the cost nobody wants. There are two legitimate ways `pkg.sub` becomes usable: the package's own `__init__.py` imports the submodule (making it part of what the package offers), or the consumer imports it explicitly with `import pkg.sub` or `from pkg import sub`. ## from-import is subtly different `from pkg import name` first imports `pkg`, then looks for the attribute `name` on it. If that lookup fails, the import system tries to import `pkg.name` as a submodule and looks again; only if that also fails does it raise `ImportError`. So `from pkg import sub` works on a submodule that no one has imported, whereas `import pkg` followed by `pkg.sub` does not. This is also why a from-import failure message so often reads "cannot import name" — the attribute was missing and no submodule of that name existed either. ## The action-at-a-distance trap Because the child-on-parent attribute binding is stored on a module object shared by the whole process, it is permanent and global. Suppose a scheduled payment-reconciliation job has a module that writes `xml.etree.ElementTree.fromstring(...)` after importing only `xml`. It works in production, because some other module imported earlier in start-up happened to pull in `xml.etree.ElementTree`. It then fails in a unit test that imports the module alone, or on the day someone deletes that unrelated import, or when the entry point changes and the import order shifts. The failure is an `AttributeError` in code that has not been touched for a year, which is why it costs so much to diagnose: the module that broke is not the module that changed. The same shape bites under `multiprocessing` with a `spawn` or `forkserver` start method, where the child process re-imports the entry module from scratch and the convenient import ordering of the parent process does not exist. ## The rule that removes the class of bug Import what you use, by its own name, in the module that uses it: `import xml.etree.ElementTree` or `from xml.etree import ElementTree`. Never rely on a parent package exposing a submodule that its `__init__.py` does not import, and never rely on a third module having imported something first. A static checker will flag an undefined name but cannot see through an attribute chain rooted in a real, imported parent package — which is precisely why this bug survives review. ## The special case people over-generalize `import os` does give you `os.path`, and that misleads. `os` is a hand-written special case: its body selects the right path module for the platform, binds it as `os.path`, and also registers it in `sys.modules` so `import os.path` works. That is a module doing the work explicitly in its own body, not a general rule of the import system. Treat any package that exposes its submodules as having chosen to import them in `__init__.py`, and do not assume the next package made the same choice.
- What name does `import a.b.c` bind in the importing module's namespace?Only `a`, the top-level name. `a.b` and `a.b.c` are reached by attribute traversal from it, having been set as attributes on their parents during the import. The `as` form is different: `import a.b.c as thing` binds only `thing` and leaves `a` unbound locally.
- How does `from pkg import name` decide whether name is an attribute or a submodule?It imports `pkg`, then looks for the attribute `name`. On failure it attempts to import `pkg.name` as a submodule and retries the lookup, raising `ImportError` only if both fail. That fallback is why a from-import reaches a submodule nobody imported, while attribute access after a bare parent import does not.
- So why does `import os` give you a working `os.path`?Because `os` does it deliberately in its own body: it picks the platform's path module, binds it as `os.path`, and registers it in `sys.modules` so `import os.path` also works. It is a hand-written special case inside one stdlib module, not general import behaviour, and generalising from it is exactly how the bug gets written.
saying these in an interview costs you the question
- Thinks importing a package imports all its submodules
- Blames a broken install for the AttributeError
- Relies on another module having imported the submodule first
- Believes import a.b binds the name b locally
- Cannot distinguish a sys.modules key from a parent attribute