A Python translation-memory CLI takes 1.4 s to print --help. How do you cut its import cost?
answer
- Measure first, against a baseline
- Only the help path must be fast
- Push heavy imports into their functions
- LazyLoader defers body execution, not binding
- Errors move to the first call
basics
~20 sMeasure first with python -X importtime against a -c pass baseline, then keep only what the help path truly needs at module level and push the heavy imports down into the functions that use them, re-measuring after each move.
solid answer
~40 sStart by measuring: `python -X importtime cli.py --help` on a warm run, minus the `python -X importtime -c "pass"` baseline, tells you which top-level imports own the 1.4 s. Then ask what `--help` actually needs — usually just the argument parser — and move everything else into the function that uses it, so the exact-arithmetic and file-format machinery a 6,800-row batch needs is paid only when a batch actually runs. For a package that must keep its public names, a module-level `__getattr__` or `importlib.util.LazyLoader` defers a module's body until an attribute is first touched. Guard annotation-only imports behind `if typing.TYPE_CHECKING:`. The tradeoff is real: deferred imports move `ImportError` and import-time side effects from startup to first call, so keep them few, deliberate and covered by a smoke test that exercises each subcommand.
code
python · 11 linesimport sys
def score_batch(rows):
from decimal import Decimal # paid on first call, not at import
return sum(Decimal(r) for r in rows) / len(rows)
print("decimal" in sys.modules)
print(score_batch(["0.1", "0.2", "0.3"]))
print("decimal" in sys.modules)go deeper
Know that an import statement can live inside a function, that it runs on first call and is cached afterwards, and that this is the usual fix for a slow-starting command-line tool.
Walk the method: measure with -X importtime against a baseline, identify the top-level imports the fast path does not need, move them into functions, then re-measure to prove the saving.
Show the judgement — which mechanism fits (function-level import, module __getattr__, LazyLoader), what deferral costs in moved failures and first-call latency, and the smoke tests that keep import errors surfacing in CI.
Frame it by deployment shape and enforce it: import cost matters per process creation, so set a startup budget that a build check defends, or dependency creep quietly restores the latency within a few releases.
### Measure before you cut Startup work is invisible in an ordinary profile of the program's logic, so start with the instrument built for it: 1. `python -X importtime -c "pass"` — the interpreter's own floor. You cannot optimise below it. 2. `python -X importtime cli.py --help 2> before.txt` — run it **twice** and keep the second: the first run in a fresh checkout compiles source to bytecode and writes `__pycache__`, which inflates the numbers. 3. Sum the cumulative column at the outermost indentation level. That is your import budget, and it is the number every change gets judged against. Now the crucial question for a CLI: **what does the `--help` path actually need?** Almost always just the argument parser and the strings that describe the subcommands. A translation-memory updater might pull in exact-decimal arithmetic to keep a per-segment score from accumulating floating-point rounding drift across a 6,800-row batch, plus file-format readers, a diffing library and a logging configuration — none of which `--help` touches, and all of which the module body imports unconditionally. ### The remedies, roughly in order of leverage **Move heavy imports into the function that needs them.** The single highest-leverage change. An `import` inside a function body executes the first time that function runs and is a `sys.modules` cache hit on every call after, so the cost is paid once per process, and only by processes that take that path. **Do not do work in module bodies.** A 6,800-entry lookup table, a compiled pattern set, a config file read at column zero: build them inside a function and cache the result with `functools.cache` so the first caller pays and nobody else does. **Keep the package's public surface lazy.** If users write `from tmtool import Updater`, a package `__init__` that eagerly imports every submodule reintroduces the whole tree. A module-level `__getattr__` (PEP 562, Python 3.7) lets the package resolve `Updater` on first attribute access and import only that submodule. **`importlib.util.LazyLoader` for whole modules.** Install it on a module's spec and the module object is created immediately while its body runs only when an attribute is first accessed. It is the right tool when you want the module bound at the top of the file yet not executed. Its caveats matter: it works with `import x`, not with `from x import y` (which reads an attribute at once and so forces execution); errors in the body surface at the point of first attribute access, in an unrelated part of the program; and it does not help for a module whose attributes are touched immediately anyway. Do not reach for it as the default — a deferred import inside a function is simpler and easier to reason about. **Guard annotation-only imports.** Imports needed solely for type annotations belong under `if typing.TYPE_CHECKING:`, which is `False` at runtime. Since Python 3.14 (PEP 649/749) annotations are evaluated lazily, so the annotation can name the type directly without quoting it — but anything that resolves annotations at runtime will still need the real import at that moment. **Ship warm bytecode.** In a container or a read-only deployment, compile the source tree at build time so no process pays compilation and no process fails to write a cache it cannot write. ### The tradeoffs a senior is expected to state Deferred imports are not free. They **move failure**: a missing dependency or a broken module body that used to blow up on startup now blows up on the first invocation of one subcommand, possibly in production, possibly after the operator has already typed something destructive. They also move latency into the first call, which matters if that call is on a latency-sensitive path rather than a batch. Mitigate with a smoke test that actually executes every subcommand, so import errors still surface in CI. They can also mislead in review: a function-level import reads as a mistake unless it carries a one-line comment saying why, and some lint configurations flag it by default. Finally, the discipline only holds if it is measured. Import budgets rot as dependencies are added; without a check that fails when total import time crosses a threshold, the 1.4 s comes back. ### Knowing when not to bother All of this matters in proportion to how often the process starts. A long-lived service pays imports once at boot and gains nothing from deferral — there, lazy imports are pure downside. A CLI invoked in a loop, a pre-commit hook, or a cold-started worker pays on every invocation, and the same 1.4 s is the dominant cost of the whole system. Decide from the deployment shape, not from taste.
- What breaks when you move an import inside a function?Failure moves in time. A missing dependency, a broken module body or an import-time side effect that used to fail at startup now fails on the first call of that function — potentially in production, on one rarely-used subcommand. Cover it with a smoke test that executes every code path, and keep deferred imports few and commented so reviewers know they are deliberate.
- Why does `from x import y` defeat `importlib.util.LazyLoader`?Because the `from` form reads an attribute off the module as part of the import, and reading any attribute is exactly what triggers the deferred body execution. You get the binding you asked for and none of the deferral. LazyLoader only helps when the module is bound with a plain `import x` and attributes are touched later.
- When is optimising import time not worth doing?When the process is long-lived. A service that boots once and serves for days pays imports once, so deferral adds risk and reading difficulty for no gain. The payoff scales with process creations per unit of work: a CLI in a loop, a pre-commit hook, or a cold-started worker justifies the effort; a daemon does not.
saying these in an interview costs you the question
- Optimises imports without measuring a baseline first
- Reaches for lazy loading everywhere instead of deferring what is heavy
- Ignores that deferred imports move ImportError to first call
- Claims deleting __pycache__ or precompiling makes the body free
- Applies CLI startup tuning to a long-lived service
- Uses LazyLoader with from-import and expects deferral