skip to content

How does a module-level `__getattr__` defer a package's heavy submodule imports?

level: seniorimportance: should knowfreq 22%

answer

  1. Fires only when a lookup misses
  2. Defined at module scope, not on a class
  3. Import on demand, then cache the result
  4. Unknown names must still raise
  5. Pair it with a companion listing hook

basics

~20 s

A __getattr__ function defined at module scope is called when an attribute lookup on that module misses. A package's __init__ can drop its eager submodule imports and import each one on first access instead, caching the result into the module globals.

solid answer

~50 s

PEP 562 (Python 3.7) lets a module define a module-level `def __getattr__(name)`, invoked only when normal attribute lookup on the module object fails. A package whose `__init__` eagerly does `from .heavy import Thing` pays for `heavy` in every process that touches the package; replacing those lines with a `__getattr__` that calls `importlib.import_module` on demand and then writes the result into `globals()` moves the cost to the first process that actually uses `Thing` — and after that write, subsequent accesses skip `__getattr__` entirely and hit the module dict. Pair it with a module-level `__dir__` so `dir()` and completion still list the lazy names, and always `raise AttributeError` for names you do not recognise, because the interpreter and libraries probe modules for optional dunders. The trade-off is that typos and import failures surface at attribute-access time, and static tooling loses the symbol unless you also declare it under `typing.TYPE_CHECKING`.

code

python · 22 lines
python
import importlib
import sys

_LAZY = {"rows": "csv", "money": "decimal"}


def __getattr__(name):
    if name in _LAZY:
        module = importlib.import_module(_LAZY[name])
        globals()[name] = module
        return module
    raise AttributeError(f"module {__name__!r} has no attribute {name!r}")


def __dir__():
    return sorted(set(globals()) | set(_LAZY))


here = sys.modules[__name__]
print("decimal" in sys.modules)
print(here.money.Decimal("0.10"))
print("decimal" in sys.modules)

go deeper

for a junior

Recall that a module can define a __getattr__ function at module scope and that it runs only when an attribute is not found on the module. Knowing it exists as the hook behind lazy package APIs is enough at this level.

for a middle

Explain the lookup order — module dict first, hook only on a miss — and why the hook caches its result into globals(). Be able to say why an unknown name must raise AttributeError and what a companion __dir__ is for.

for a senior

Show when the pattern earns its keep: a package whose measured startup is dominated by eager submodules and whose public API must not change. Cover the costs you accept — deferred failures, weaker static analysis, an extra frame in tracebacks — and how you mitigate them with a TYPE_CHECKING block and tests on the lazy paths.

for a principal

Decide whether a whole library adopts this shape. Lazy public surfaces buy startup for users in short-lived processes and cost debuggability and tooling accuracy for everyone; the call depends on how your users invoke the code, and it is a commitment the whole package must keep consistently.

## The mechanism Since **Python 3.7** (PEP 562), if a module's namespace contains a callable named `__getattr__`, the module object's attribute lookup falls back to it. The order is: look the name up in the module's `__dict__`; if it is there, return it; if not, call `__getattr__(name)` and return whatever it returns; if there is no `__getattr__`, raise `AttributeError`. Two things follow immediately. First, it is a **miss-only** hook — it costs nothing for names that already exist, unlike an instance `__getattribute__`, which intercepts every access. Second, it fires for attribute access **on the module object**, which includes `package.Thing`, `getattr(package, "Thing")` and the fallback path of `from package import Thing`. It does **not** fire for a bare name inside the module's own code, because that is a global lookup, not an attribute access. ## Why it is the natural lazy-import tool for a package A package's `__init__.py` is where a convenient public API is usually assembled: `from .parsing import RowParser`, `from .money import Amount`, and so on. Every one of those lines runs the corresponding submodule's body the moment anything imports the package — including a process that only wants a version string. Rewriting the `__init__` around `__getattr__` keeps the public API identical while making each submodule's cost conditional: * the name maps to a module (or an attribute of one) to import on demand; * `importlib.import_module` performs the import when the name is first touched; * the result is written into `globals()`, so the second access finds it in the module dict and never reaches `__getattr__` again; * an unknown name raises `AttributeError` with the standard message shape. That last point is not cosmetic. The interpreter and the standard library probe modules for optional attributes — `__all__`, `__path__`, `__wrapped__` and various protocol dunders among them. A `__getattr__` that returns something truthy for every name, or that raises `ImportError` instead of `AttributeError`, breaks `hasattr` checks, `copy`, `pickle` and introspection in ways that are painful to trace back to their cause. ## Keeping discoverability `dir(module)` reads the module dict, so lazily exposed names are invisible until they have been touched. PEP 562 also allows a module-level `__dir__` returning the names you want listed; define it alongside `__getattr__` and REPL completion, `dir()` and documentation tools behave as before. For static analysis, the lazy names do not exist at all — a type checker or IDE sees a `__getattr__` and, at best, gives up on precision. The usual remedy is to declare the real imports under `if typing.TYPE_CHECKING:`, which is `False` at runtime, so tools see the symbols and the process never imports them. ## Concurrency and repeated access Two threads can hit `__getattr__` for the same name at once. The import system takes a per-module lock, so the submodule's body still runs exactly once and both callers receive the finished module; the subsequent `globals()` assignment is idempotent — both threads store the same object. This holds on the free-threaded build as well, and it is why the pattern is safe in practice. What is *not* safe is a `__getattr__` that does something non-idempotent — appending to a registry, opening a file, mutating shared state — because it can run more than once before the caching write lands. ## The alternatives, and when each fits * **Move the import into the function that needs it.** Simplest, and correct when the dependency is used in one place. `__getattr__` is for the case where the dependency must remain part of the package's public surface. * **`importlib.util.LazyLoader`.** Builds a module object whose body is executed on first attribute access, for when you want a whole module lazy rather than a set of names. It is fiddlier — you construct the spec and loader yourself — and it does not help you keep a curated public API. ## What you give up Errors move. A typo in a lazy name is an `AttributeError` at the moment of use, not an `ImportError` at startup. A submodule that fails to import fails later, in a less obvious place, and inside whatever thread first touched it. Debuggers and tracebacks get one extra frame. And a reader of `__init__.py` can no longer see the package's real shape at a glance. That is a reasonable trade for a library whose users import it in short-lived processes, and a poor one for a small internal package that a long-running service imports once. Reach for it when a measurement — not an instinct — says a package's eager submodules dominate startup.

  • Why must a module-level `__getattr__` raise `AttributeError` for names it does not know?
    Because attribute lookup on a module is how the interpreter and the standard library probe for optional attributes — `__all__`, `__path__`, `__wrapped__` and assorted protocol dunders. Those probes are written as `hasattr` or a caught `AttributeError`. Returning a value for every name makes the module claim to support things it does not; raising the wrong exception type turns a probe into a crash in unrelated code such as copying, pickling or introspection.
  • Why does the function assign the imported object into `globals()` rather than just returning it?
    So the hook runs once. After the assignment the name is in the module dict, so ordinary attribute lookup succeeds and `__getattr__` is never consulted for it again. Without the write, every access pays a `sys.modules` lookup through `importlib.import_module` plus a Python-level function call — small, but pure waste on a name that may be read in a loop.
  • How do you keep a type checker working with a lazily exposed package API?
    Declare the real imports inside `if typing.TYPE_CHECKING:` in the same module. That block is never executed at runtime, so no cost is added, but static tools read it and see the true symbols and their types. Adding a module-level `__dir__` covers the runtime half of the same problem for `dir()` and REPL completion.

saying these in an interview costs you the question

  • Thinks module `__getattr__` intercepts every attribute access
  • Returns None for unknown names instead of raising
  • Raises ImportError where AttributeError is required
  • Forgets to cache, so the hook runs on every access
  • Expects it to fire for bare names inside the same module
  • Puts non-idempotent side effects inside the hook

context