skip to content

Startup and Import Cost

Where the first second goes before your code runs: every import executes module-level code, and a deep dependency tree can dominate a CLI or a cold serverless start. Measured with -X importtime.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

Why does a Python `import` statement execute code rather than just bind a name?

level: juniorimportance: must knowfreq 50%

answer

  1. Import runs code, not a declaration
  2. The body executes top to bottom
  3. Only the first time per process
  4. sys.modules holds the finished module
  5. Top-level tables are startup cost

basics

~20 s

The first import of a module runs its whole body top to bottom, then stores the finished module object so later imports are just a cache lookup. Every top-level statement in it is startup cost.

solid answer

~40 s

A module body is ordinary code, and `import` is the statement that runs it. The first time a process imports a module, CPython finds it, compiles it to bytecode if no valid cached bytecode exists, creates a module object, and executes the body in that module's namespace; only then does it land in `sys.modules` and get bound to a name. Any later `import` of the same module in the same process finds the `sys.modules` entry and binds the name without re-running anything. So `def`, `class`, decorators, default argument values, module-level constants and — recursively — the module's own top-level imports are all work done before your entry point runs. That is why an import tree, not your code, usually dominates the first second of a CLI or a cold start.

code

python · 14 lines
python
import sys
import types

body = """
print("tm_index body is running")
ROWS = 6800
TABLE = [i / 3 for i in range(ROWS)]
"""
module = types.ModuleType("tm_index")
exec(body, module.__dict__)
sys.modules["tm_index"] = module

import tm_index          # already in sys.modules: the body does not run again
print(len(tm_index.TABLE), tm_index.TABLE[1])

go deeper

for a junior

Be ready to say plainly that importing runs the module's code the first time and caches the result, so anything written at the top level of a module happens before your program starts.

for a middle

Explain the mechanics: body executed in the module's namespace, object stored in sys.modules, later imports reduced to a name binding, and the transitive cost of the module's own top-level imports.

for a senior

Show you connect it to production: the same import graph is harmless in a long-lived service and fatal in a per-invocation CLI or a cold-started worker, and import-time side effects are untestable and unconfigurable.

for a principal

Own the policy angle — a convention that module bodies stay cheap and side-effect-free is what keeps startup from rotting as a codebase grows, and it needs a measurement in CI, not good intentions.

### An import is a statement, not a declaration In some languages an import is a compile-time instruction that makes names visible. In Python it is an executable statement with a side effect: *run this module, once, and give me the object that results.* The first time a process executes `import tm_index`, CPython locates the module, compiles its source to bytecode unless valid cached bytecode already exists, creates an empty module object, and then **executes the module's body from top to bottom** with that module's `__dict__` as the global namespace. Only after the body finishes does the module object settle into `sys.modules` and a name get bound in the importing namespace. Every later `import tm_index` anywhere in the process is a `sys.modules` lookup plus a name binding — the body does **not** run again. (`importlib.reload` is the explicit escape hatch; it re-executes the body into the same module object.) ### What "the body runs" actually costs Everything at column zero of a module is code that executes at import: * `def` and `class` statements build function and class objects. Decorators are *called*. Default argument values are *evaluated* — `def f(rows=list(range(6800)))` builds that list at import time. * Module-level constants build real objects. A 6,800-entry lookup table written as a dict literal or a comprehension is 6,800 insertions performed before your `main()` is reached. * `import` statements at the top of a module recursively execute *those* modules' bodies. This is the big one: a single convenience import can drag a tree of dozens of modules behind it, and you pay for the whole transitive closure. * Genuine side effects happen: reading a config file, compiling regular expressions, registering plugins in a global table, building a logger, even opening a socket. All of it before the first line of your program logic. Because imports nest, import cost is a *tree*, and `python -X importtime` is the tool that prints that tree with per-module timings. ### The consequences you are expected to name **Cost is per process, not per call.** A long-lived service pays the import bill once at boot and never thinks about it again. A CLI pays it on every invocation — run it once per record over a 6,800-row batch and a 900 ms import tree is fifteen minutes of pure startup. A serverless or short-lived worker pays it on every cold start. The same import graph is free in one deployment shape and unacceptable in another, which is why "is this import slow?" is always really "how often is this process created?". **Partially-initialised modules are visible.** Because the module object is placed in `sys.modules` *before* the body finishes, a module that is re-entered while still executing (a cycle) hands out a half-built namespace, which is why circular imports fail with a name that "should" exist. That is a direct consequence of body execution, not a quirk of the syntax. **Import-time side effects are hard to test and hard to undo.** Work that runs at import cannot be skipped, configured or mocked by a caller who imports the module — it has already happened. Prefer functions and lazily-built values over module-level work whenever the work is expensive or environment-dependent. **Imports are cached, not memoised per importer.** Ten modules importing the same dependency cost one execution. The corollary matters when reading a profile: whoever imports a shared dependency *first* is charged for it, and everyone after sees a free import. ### Two costs, not one Import cost splits into **compilation** (source → bytecode, done once and saved in `__pycache__`) and **execution** (running the body, done once per process, every process). Only the first is cached across runs. Deleting a `.pyc` makes the *first* run slower; it never makes the body free. ### A 3.14 note on annotations Historically, annotations on module-level functions and classes were evaluated at definition time, so an annotation naming a type forced its module to be imported and its expression to be computed during import. Python 3.14 changed this: under PEP 649/749 annotations are evaluated lazily, only when something actually asks for them, so annotations no longer contribute to import time and `from __future__ import annotations` is no longer needed to avoid that cost. The rest of the module body still runs exactly as before.

  • If two modules both import the same dependency, how many times does its body run?
    Once per process. The first import executes the body and stores the module object in `sys.modules`; the second import finds that entry and only binds a name. This is also why a profile can mislead — the shared dependency's cost is attributed entirely to whoever imported it first, and every later importer looks free.
  • What is the practical difference between work in a module body and work in a function?
    Body work is unconditional and happens at import, before the caller can influence it — it cannot be skipped, configured or deferred, and it is paid even by a process that never uses the feature. Function work is paid only when called, can be cached, and is easy to test. Anything expensive or environment-dependent belongs in a function or behind a lazily-built value.
  • Does deleting `__pycache__` change how often a module's body runs?
    No. `__pycache__` caches compilation, not execution. Removing it means the source is parsed and compiled again on the next run, which slows that one run down; the body still executes exactly once per process either way.

An import is less like reading a chapter's title and more like performing the whole chapter aloud: the first reader performs every line, and everyone who comes later is simply handed the finished result.

saying these in an interview costs you the question

  • Thinks import only makes names visible, runs nothing
  • Believes the module body runs on every import statement
  • Assumes a .pyc file means the body is not executed
  • Cannot explain why a top-level constant costs startup time
  • Thinks unused imports are free because nothing calls them

context

open as a page

What does CPython store in `__pycache__`, and why does it not make startup free?

level: middleimportance: should knowfreq 28%

basics

~20 s

It stores compiled bytecode for the source files beside it, named like mod.cpython-314.pyc, so later runs skip parsing and compiling. It caches compilation only — every process still executes each module body, which is usually the larger cost.

open as a page

How do you read the self and cumulative columns of `python -X importtime` output?

level: middleimportance: should knowfreq 30%

basics

~20 s

-X importtime prints one line per module to stderr with two microsecond columns: self is time in that module's own body, cumulative is self plus everything it imported. Indentation shows the tree, and children print before their parent.

open as a page

A Python translation-memory CLI takes 1.4 s to print --help. How do you cut its import cost?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Measure first with python -X importtime against a -c pass baseline, then keep only what the help path truly needs at module level and push the heavy imports down into the functions that use them, re-measuring after each move.

open as a page