skip to content

What does CPython's `-X importtime` print, and how do its two time columns differ?

level: middleimportance: should knowfreq 32%

answer

  1. One line per import, on stderr
  2. Two microsecond columns per module
  3. One excludes children, one includes them
  4. Children print above their parent
  5. Already-imported modules never appear

basics

~20 s

It writes one line per module imported during the run to stderr, showing that module's own execution time and the cumulative time including everything it triggered. Cumulative tells you which branch to cut; self tells you which module is doing the work.

solid answer

~40 s

Running `python -X importtime yourscript.py` makes the import system report every import it performs, as `import time: self [us] | cumulative | imported package` on **stderr**, so stdout stays clean. **Self** is the microseconds spent executing that module's own body, excluding the imports it triggered. **Cumulative** is self plus everything imported beneath it. Children are printed before their parent and indented one level deeper, so the report reads as a tree completed innermost-first. Only real imports appear: anything already in `sys.modules` costs nothing and is not listed. Read cumulative to pick which top-level import to defer or delete, then read self to find the module actually spending the time — a façade package often has a tiny self and a huge cumulative. The environment variable `PYTHONPROFILEIMPORTTIME=1` does the same thing.

code

console · 1 line
console
python -X importtime -c "import csv" 2>&1 | tail -5

go deeper

for a junior

Know that the flag exists and what it is for: it shows what a program imports at startup and how long each import took, printed to stderr. Being able to run it and point at the biggest line is enough here.

for a middle

Explain the two columns precisely — one module's own body versus that body plus everything under it — and the tree layout where children print above their parent. Say which column you use to choose what to cut and which one tells you where the work really is.

for a senior

Demonstrate the measurement discipline: profile a representative cold invocation, run it twice so bytecode caching does not skew the first read, measure on the deployment interpreter, and re-measure after each change because imports are charged to whichever branch reached them first.

for a principal

Frame it as evidence, not a report. Decide what invocation defines your startup number, what the acceptable figure is, and how the measurement becomes a standing check rather than a one-off investigation someone did before a release.

## What the flag does `-X importtime` instruments CPython's import machinery so that every module it actually imports during the process prints a timing line. The report goes to **stderr**, which matters: you can measure a tool whose stdout is piped somewhere without corrupting the pipe. The header line is literally: ``` import time: self [us] | cumulative | imported package ``` and each following line carries two integer microsecond figures and a module name whose indentation shows how deep in the import tree it sits. ## Self versus cumulative **Self** is the time spent executing *that module's own body* — its statements, its class and function definitions, whatever tables it builds, whatever regexes it compiles — plus its own find-and-load overhead, but **excluding** the modules its body imported. **Cumulative** is self plus the cumulative time of every module it was the first to trigger. The two answer different questions: * Cumulative answers **"which import should I stop doing?"** Deferring or deleting a top-level import buys you its cumulative number, not its self number. * Self answers **"who is actually spending the time?"** If a package's `__init__` has a self of 200 microseconds and a cumulative of 300 milliseconds, the package is a façade — it re-exports names from heavy children, and the fix is somewhere below it. If a module has a large *self*, its own body is doing real work at import time, and that is usually code you control and can move into a function. ## Reading the tree The report is emitted as each import **completes**, so a child always appears **above** its parent, indented one level further. Reading it top-down feels backwards; read it as a call tree printed post-order. The last unindented lines are the top-level imports your own code performed, and their cumulative figures are the budget lines that matter. Two consequences follow from "each module is charged to whoever imported it first": * A module imported by two different branches is charged entirely to whichever branch reached it first. Deferring only that branch may move the cost rather than remove it — you have to defer both. * Nothing already in `sys.modules` is listed at all. The interpreter's own startup imports are visible near the top, but a module you imported earlier in the same process shows up once and only once. ## Practical use ```console python -X importtime -c "import csv" 2>&1 | tail -5 ``` Run it more than once. The first run after an edit compiles source to cached bytecode and the numbers are inflated; the steady-state figure is what a user experiences. Instrumentation is not free either — measuring adds a small constant per import, so a tree with a thousand tiny modules reads slightly worse than it is. Compare runs rather than trusting absolute numbers, and compare on the machine and interpreter you deploy on, not on a fast laptop. `PYTHONPROFILEIMPORTTIME=1` is the environment-variable equivalent, which is how you get the report out of a process you cannot re-invoke with different flags — a wrapper script, a container entrypoint, a scheduled job. ## What it does not tell you It measures **import**, not execution. Time your program spends after startup — in the work itself — is invisible here, and steady-state profiling is a different tool. It also does not tell you *why* a module is heavy; for that, read its body. Common answers are: a large table or dataclass hierarchy constructed at import, regexes compiled at module level, a plugin scan walking a directory, a configuration file read, or simply a deep dependency tree where no single module is at fault. Finally, it does not tell you what to do. The report only ranks candidates. The decisions — which imports are part of the eager core, which move into the functions that need them, which module bodies should stop doing work at all — are yours, and they need a target number to be judged against. ## The typical workflow Measure a cold, representative invocation (often `--help`, because it is the cheapest thing your tool can be asked to do and therefore the purest measure of startup). Take the largest cumulative branch at the top level. Ask whether every invocation truly needs it. Move it, re-measure, and keep the before/after numbers — a startup fix that nobody recorded is a startup fix that will be undone.

  • A package shows a self time of 90 microseconds and a cumulative of 300 milliseconds. What do you conclude?
    That the package's own body does almost nothing and its cost is entirely in the children it imports — a façade `__init__` that re-exports names from heavy submodules. Deferring the import of the package buys the whole 300 milliseconds, but fixing it from the inside means making those submodules lazy, for example behind a module-level `__getattr__`, rather than editing the `__init__` body.
  • Why might deferring one import fail to reduce total startup time at all?
    Because each module is charged to whoever imported it first. If two top-level branches both reach the same heavy subtree, the report attributes all of it to the branch that got there first; deferring only that branch just re-attributes the cost to the other one. Re-measure after every change and defer every path that reaches the subtree, or the number will not move.
  • Why is the report written to stderr rather than stdout?
    So it can be collected from a tool whose stdout is data. A command whose output is piped into another process, redirected to a file, or parsed by a caller would be corrupted by timing lines on stdout; on stderr you can profile a real invocation in place, and redirect with `2>` when you want the report in a file of its own.

saying these in an interview costs you the question

  • Says the report goes to stdout
  • Treats self time as including nested imports
  • Reads the report top-down as execution order
  • Expects already-imported modules to appear again
  • Trusts a single cold run's absolute numbers
  • Confuses it with profiling the program's actual work

context