skip to content

How do you read the self and cumulative columns of `python -X importtime` output?

level: middleimportance: should knowfreq 30%

answer

  1. Two numbers and a tree
  2. Written to stderr, not stdout
  3. Self excludes the imports it triggered
  4. Cumulative covers the whole subtree
  5. Cache hits print no line at all

basics

~20 s

-X importtime prints one line per module to stderr with two microsecond columns: self is time in that module's own body, cumulative is self plus everything it imported. Indentation shows the tree, and children print before their parent.

solid answer

~50 s

Run `python -X importtime yourscript.py` (or set `PYTHONPROFILEIMPORTTIME`) and CPython writes a table to **stderr**, headed `import time: self [us] | cumulative | imported package`. Each line is a module: **self** is the microseconds spent executing that module's own body excluding the imports it triggered, **cumulative** is self plus the whole subtree it pulled in. Indentation encodes nesting and leaves print before their parents, so read bottom-up: the last line at the outermost level is the top-level import and its cumulative is what that one statement cost you. Sort your attention by self time to find a module whose own body is slow, and by cumulative on top-level imports to decide what is worth deferring. The catch: a module already imported by something earlier produces no line at all, so shared dependencies are charged entirely to whoever imported them first.

code

console · 1 line
console
python3 -X importtime -c "import decimal"

go deeper

for a junior

Know the flag exists and what it is for: python -X importtime tells you which imports are making startup slow, and the table comes out on stderr so you need 2>&1 to pipe it.

for a middle

Be able to read the table unaided — self versus cumulative, indentation as nesting, children printed above parents — and say which column answers 'make it cheaper' versus 'delete the import'.

for a senior

Demonstrate scepticism about the numbers: first-importer attribution for shared dependencies, a cold first run that includes bytecode compilation, and the -c pass baseline you must subtract before drawing conclusions.

for a principal

Treat total import time as a number worth tracking, with a budget that a build check enforces; without one, dependency creep silently returns startup to where it was within a couple of releases.

### What the flag does `python -X importtime script.py` makes CPython time every import performed in the process and print a table when each import completes. The equivalent environment variable is `PYTHONPROFILEIMPORTTIME`, useful when you cannot edit the command line — a container entrypoint, a supervisor unit, a serverless runtime. Output goes to **stderr**, not stdout. That is deliberate: the program's own output stays clean and pipeable, and it means you capture the table with `2>` or `2>&1`, not with a plain pipe. The header is: ``` import time: self [us] | cumulative | imported package ``` ### The two columns * **self** — microseconds spent executing *that module's own body*, with the time of any imports it triggered subtracted out. High self time means this module's own top-level code is slow: building big tables, compiling patterns, reading files, defining hundreds of classes. * **cumulative** — self plus the entire subtree of modules this one caused to be imported. High cumulative with low self means the module is cheap itself but drags in an expensive dependency tree. The distinction is the whole point of the tool. Self answers *"which module should I make cheaper?"*; cumulative answers *"what would I save if this import went away?"*. ### Reading the tree Indentation shows nesting, and because a line is printed when an import *finishes*, children appear **above** their parent. The output therefore reads bottom-up: a line at the leftmost indentation level is a module imported directly by your program, and its cumulative number is the true cost of that one `import` statement. ``` import time: 213 | 213 | numbers import time: 972 | 1184 | _decimal import time: 94 | 1278 | decimal ``` Here `import decimal` cost 1,278 microseconds in total; `decimal`'s own body accounted for 94 of them, and nearly all of the rest came from the extension module and `numbers` beneath it. ### The trap: shared dependencies are charged to the first importer An import that hits the `sys.modules` cache does no work, so it produces **no line in the table**. If two of your top-level imports both need the same expensive dependency, whichever is imported first is billed for all of it and the second looks nearly free. Deleting the "expensive" import then saves far less than the profile promised, because the cost simply moves to the other importer. The practical defence is to compare *totals*, not lines: measure the whole startup with and without a candidate import, rather than trusting one cumulative figure. Re-ordering imports and re-running is also a quick way to reveal a shared subtree. ### Establishing a baseline `python -X importtime -c "pass"` shows what the interpreter imports before your code exists at all — the codec machinery, the `site` module and what it pulls in. On a normal installation that is already a few dozen lines. Subtract that from your program's total before congratulating or blaming yourself; you cannot optimise away the interpreter's own boot. ### Measurement hygiene * **Run it twice.** The first run in a fresh checkout or a fresh container compiles source to bytecode and writes `__pycache__`; the second run reads the cache. Only the second is representative of steady state. * **Cold filesystem caches inflate everything.** A first import after boot pays disk latency that a warm run does not. * **It measures imports only.** Time spent in your `main()`, in argument parsing, or waiting on the network is invisible here. This is a startup tool, not a general profiler. * **Microseconds are noisy at the leaves.** Chase the four-figure lines; ignore the ones costing 50 microseconds however numerous they are. * **Sum the top level.** Adding the cumulative values of the outermost lines gives you total import cost, which is the number to track over time. ### Reading it in awkward places Lines are emitted as each import *completes*, so even a program that crashes during startup leaves a usable partial table — often the fastest way to see which import was in flight when it died. In a container or a managed runtime where you cannot change the command line, `PYTHONPROFILEIMPORTTIME` in the environment gives the identical output, and redirecting stderr to a file keeps it out of the log stream you are trying to read. Expect volume: a moderately sized program produces hundreds of lines. Skim for the outermost indentation level first to get the shape of the tree and the total, and only then go hunting inside the one subtree that owns most of the time. ### Where it fits `-X importtime` is the right first instrument for "the program is slow before it does anything" and for cold starts, exactly because it needs no code change, no library and no instrumentation. Once the process is running and doing work, the question stops being an import question and calls for a different instrument entirely.

  • A module shows 900 ms cumulative and 2 ms self. What do you conclude?
    Its own body is trivial and essentially all the cost is the dependency subtree it pulls in. Making that module faster is pointless; the questions are whether you need it on this code path at all, whether the import can be deferred to the function that uses it, and which specific child in its subtree is actually expensive.
  • Why can removing the import with the largest cumulative time save almost nothing?
    Because a `sys.modules` cache hit prints no line, so a dependency shared by two importers is charged entirely to whichever ran first. Remove that importer and the cost reappears under the other one. Verify by measuring total startup before and after, not by reading a single cumulative figure.
  • Why should you ignore the first run's numbers in a fresh container?
    The first run compiles source to bytecode and writes `__pycache__`, and it also pays cold filesystem-cache latency. Both inflate import times relative to steady state. Run it at least twice and use the warm numbers — unless the cold run *is* the thing you are optimising, as with a read-only image that never gets a warm run.

saying these in an interview costs you the question

  • Expects the table on stdout and loses it to a pipe
  • Reads cumulative as time spent in that module's own body
  • Reads the tree top-down instead of bottom-up
  • Trusts one run in a fresh checkout as steady state
  • Assumes a missing module means it was never imported
  • Thinks it profiles the whole program, not just imports

context