skip to content

Profiling and Benchmarking

Measuring instead of guessing: time a snippet with timeit, profile a whole run with cProfile, attribute memory with tracemalloc. 'This job is slow, what first?' wants a method, not a hunch.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

23

Why does a Python `import` statement execute code rather than just bind a name?

level: juniorimportance: must knowfreq 50%

answer

  1. Import runs code, not a declaration
  2. The body executes top to bottom
  3. Only the first time per process
  4. sys.modules holds the finished module
  5. Top-level tables are startup cost

basics

~20 s

The first import of a module runs its whole body top to bottom, then stores the finished module object so later imports are just a cache lookup. Every top-level statement in it is startup cost.

solid answer

~40 s

A module body is ordinary code, and `import` is the statement that runs it. The first time a process imports a module, CPython finds it, compiles it to bytecode if no valid cached bytecode exists, creates a module object, and executes the body in that module's namespace; only then does it land in `sys.modules` and get bound to a name. Any later `import` of the same module in the same process finds the `sys.modules` entry and binds the name without re-running anything. So `def`, `class`, decorators, default argument values, module-level constants and — recursively — the module's own top-level imports are all work done before your entry point runs. That is why an import tree, not your code, usually dominates the first second of a CLI or a cold start.

code

python · 14 lines
python
import sys
import types

body = """
print("tm_index body is running")
ROWS = 6800
TABLE = [i / 3 for i in range(ROWS)]
"""
module = types.ModuleType("tm_index")
exec(body, module.__dict__)
sys.modules["tm_index"] = module

import tm_index          # already in sys.modules: the body does not run again
print(len(tm_index.TABLE), tm_index.TABLE[1])

go deeper

for a junior

Be ready to say plainly that importing runs the module's code the first time and caches the result, so anything written at the top level of a module happens before your program starts.

for a middle

Explain the mechanics: body executed in the module's namespace, object stored in sys.modules, later imports reduced to a name binding, and the transitive cost of the module's own top-level imports.

for a senior

Show you connect it to production: the same import graph is harmless in a long-lived service and fatal in a per-invocation CLI or a cold-started worker, and import-time side effects are untestable and unconfigurable.

for a principal

Own the policy angle — a convention that module bodies stay cheap and side-effect-free is what keeps startup from rotting as a codebase grows, and it needs a measurement in CI, not good intentions.

### An import is a statement, not a declaration In some languages an import is a compile-time instruction that makes names visible. In Python it is an executable statement with a side effect: *run this module, once, and give me the object that results.* The first time a process executes `import tm_index`, CPython locates the module, compiles its source to bytecode unless valid cached bytecode already exists, creates an empty module object, and then **executes the module's body from top to bottom** with that module's `__dict__` as the global namespace. Only after the body finishes does the module object settle into `sys.modules` and a name get bound in the importing namespace. Every later `import tm_index` anywhere in the process is a `sys.modules` lookup plus a name binding — the body does **not** run again. (`importlib.reload` is the explicit escape hatch; it re-executes the body into the same module object.) ### What "the body runs" actually costs Everything at column zero of a module is code that executes at import: * `def` and `class` statements build function and class objects. Decorators are *called*. Default argument values are *evaluated* — `def f(rows=list(range(6800)))` builds that list at import time. * Module-level constants build real objects. A 6,800-entry lookup table written as a dict literal or a comprehension is 6,800 insertions performed before your `main()` is reached. * `import` statements at the top of a module recursively execute *those* modules' bodies. This is the big one: a single convenience import can drag a tree of dozens of modules behind it, and you pay for the whole transitive closure. * Genuine side effects happen: reading a config file, compiling regular expressions, registering plugins in a global table, building a logger, even opening a socket. All of it before the first line of your program logic. Because imports nest, import cost is a *tree*, and `python -X importtime` is the tool that prints that tree with per-module timings. ### The consequences you are expected to name **Cost is per process, not per call.** A long-lived service pays the import bill once at boot and never thinks about it again. A CLI pays it on every invocation — run it once per record over a 6,800-row batch and a 900 ms import tree is fifteen minutes of pure startup. A serverless or short-lived worker pays it on every cold start. The same import graph is free in one deployment shape and unacceptable in another, which is why "is this import slow?" is always really "how often is this process created?". **Partially-initialised modules are visible.** Because the module object is placed in `sys.modules` *before* the body finishes, a module that is re-entered while still executing (a cycle) hands out a half-built namespace, which is why circular imports fail with a name that "should" exist. That is a direct consequence of body execution, not a quirk of the syntax. **Import-time side effects are hard to test and hard to undo.** Work that runs at import cannot be skipped, configured or mocked by a caller who imports the module — it has already happened. Prefer functions and lazily-built values over module-level work whenever the work is expensive or environment-dependent. **Imports are cached, not memoised per importer.** Ten modules importing the same dependency cost one execution. The corollary matters when reading a profile: whoever imports a shared dependency *first* is charged for it, and everyone after sees a free import. ### Two costs, not one Import cost splits into **compilation** (source → bytecode, done once and saved in `__pycache__`) and **execution** (running the body, done once per process, every process). Only the first is cached across runs. Deleting a `.pyc` makes the *first* run slower; it never makes the body free. ### A 3.14 note on annotations Historically, annotations on module-level functions and classes were evaluated at definition time, so an annotation naming a type forced its module to be imported and its expression to be computed during import. Python 3.14 changed this: under PEP 649/749 annotations are evaluated lazily, only when something actually asks for them, so annotations no longer contribute to import time and `from __future__ import annotations` is no longer needed to avoid that cost. The rest of the module body still runs exactly as before.

  • If two modules both import the same dependency, how many times does its body run?
    Once per process. The first import executes the body and stores the module object in `sys.modules`; the second import finds that entry and only binds a name. This is also why a profile can mislead — the shared dependency's cost is attributed entirely to whoever imported it first, and every later importer looks free.
  • What is the practical difference between work in a module body and work in a function?
    Body work is unconditional and happens at import, before the caller can influence it — it cannot be skipped, configured or deferred, and it is paid even by a process that never uses the feature. Function work is paid only when called, can be cached, and is easy to test. Anything expensive or environment-dependent belongs in a function or behind a lazily-built value.
  • Does deleting `__pycache__` change how often a module's body runs?
    No. `__pycache__` caches compilation, not execution. Removing it means the source is parsed and compiled again on the next run, which slows that one run down; the body still executes exactly once per process either way.

An import is less like reading a chapter's title and more like performing the whole chapter aloud: the first reader performs every line, and everyone who comes later is simply handed the finished result.

saying these in an interview costs you the question

  • Thinks import only makes names visible, runs nothing
  • Believes the module body runs on every import statement
  • Assumes a .pyc file means the body is not executed
  • Cannot explain why a top-level constant costs startup time
  • Thinks unused imports are free because nothing calls them

context

open as a page

Why do objects appended to a module-level list never get freed in Python?

level: juniorimportance: must knowfreq 55%

basics

~20 s

A module-level list stays reachable from sys.modules for the whole process, so every object it holds keeps a live reference and can never be freed. Python reclaims an object only when nothing reachable still refers to it.

open as a page

What is the difference between wall-clock time and CPU time when timing Python code?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Wall-clock time is elapsed real time, read with time.perf_counter(). CPU time is the time the processor actually spent executing the process, read with time.process_time(). Sleeping, waiting on a socket or blocking on a lock adds wall time but no CPU time.

open as a page

Why does timeit.timeit take a separate setup argument instead of one block of code?

level: juniorimportance: must knowfreq 55%

basics

~20 s

timeit runs the setup string once, before the clock starts, and only the stmt string inside the timed loop. Imports and test-data construction are therefore excluded, so you measure the operation itself rather than its fixtures.

open as a page

In a cProfile report, what is the difference between tottime and cumtime?

level: middleimportance: must knowfreq 55%

basics

~20 s

tottime is the time spent inside a function's own body, excluding the calls it makes. cumtime adds everything it called. Sort by tottime to find hot code, by cumtime to find the expensive call path.

open as a page

How does tracemalloc's Snapshot.compare_to help you find a slow leak in a long-running service?

level: middleimportance: must knowfreq 50%

basics

~20 s

Take one snapshot after warm-up and another after many work cycles, then call second.compare_to(first, 'lineno'). The result is sorted by size_diff, so the lines that grew between the two snapshots rise to the top and steady-state memory cancels out.

open as a page

How do you profile a script with cProfile and inspect the saved profile using pstats?

level: juniorimportance: should knowfreq 45%

basics

~10 s

Run python -m cProfile -o out.prof script.py to capture a profile, then load it with pstats.Stats("out.prof") and call strip_dirs, sort_stats and print_stats to read the top rows.

open as a page

How do you use tracemalloc to find which lines of a script allocated the most memory?

level: juniorimportance: should knowfreq 25%

basics

~10 s

Call tracemalloc.start() before the work, tracemalloc.take_snapshot() after it, then sort the snapshot with Snapshot.statistics('lineno'). Each entry names a file and line and reports the bytes and block count still allocated there.

open as a page

What does CPython store in `__pycache__`, and why does it not make startup free?

level: middleimportance: should knowfreq 28%

basics

~20 s

It stores compiled bytecode for the source files beside it, named like mod.cpython-314.pyc, so later runs skip parsing and compiling. It caches compilation only — every process still executes each module body, which is usually the larger cost.

open as a page

How do you read the self and cumulative columns of `python -X importtime` output?

level: middleimportance: should knowfreq 30%

basics

~20 s

-X importtime prints one line per module to stderr with two microsecond columns: self is time in that module's own body, cumulative is self plus everything it imported. Indentation shows the tree, and children print before their parent.

open as a page

How does gc.get_objects() with collections.Counter give a census of live objects?

level: middleimportance: should knowfreq 35%

basics

~20 s

gc.get_objects() returns a list of every container object the cyclic collector currently tracks. Feeding the type name of each one into collections.Counter yields counts per class; taking two censuses around a workload and diffing them shows which type is accumulating.

open as a page

Why does timeit report the best of several repeats instead of the average?

level: middleimportance: should knowfreq 40%

basics

~20 s

Timing noise on a real machine is one-sided: other processes, interrupts and cache evictions can only make a run slower. The minimum repetition is therefore the least disturbed sample and the most reproducible estimate of the code's cost.

open as a page

A cProfile run on an invoice-PDF renderer blames a tiny helper called millions of times; how far do you trust that number?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Trust the ranking, not the seconds. cProfile pays a fixed cost per call and return event, so a function called millions of times absorbs enormous instrumentation cost. Compare the profiled run against an unprofiled one before acting.

open as a page

A Python translation-memory CLI takes 1.4 s to print --help. How do you cut its import cost?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Measure first with python -X importtime against a -c pass baseline, then keep only what the help path truly needs at module level and push the heavy imports down into the functions that use them, re-measuring after each move.

open as a page

How do you use gc.get_referrers() to find what still holds a leaking object?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Take one instance of the accumulating type and call gc.get_referrers() on it to get the tracked objects pointing at it, then repeat on each result to climb towards a named holder. Most first-level hits are anonymous dicts.

open as a page

A route-optimisation job runs 40 minutes, yet its Python call profile accounts for two — where did the wall-clock time go?

level: seniorimportance: should knowfreq 45%

basics

~10 s

A deterministic call profile records calls in one thread and counts time inside them, not time spent waiting. Compare time.perf_counter() with time.process_time(), then bracket each phase with wall-clock timers and sample every thread.

open as a page

Why does sys.setprofile installed in the main thread miss work done in worker threads?

level: seniorimportance: should knowfreq 30%

basics

~20 s

The profiling hook set by sys.setprofile is per-thread interpreter state, so it only fires for the thread that installed it. Use threading.setprofile for threads started later, threading.setprofile_all_threads for existing ones, or a profiler that samples every thread's stack.

open as a page

timeit says a rewritten bid-scoring function is 3x faster, but the ad-auction service's p99 did not move. What do you check?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Check, in order: the function's share of a request, whether the benchmark's inputs and warm state resemble production, and whether the timed snippet did the work at all. Then confirm end to end, not in a loop.

open as a page

Why can tracemalloc report flat Python memory while the process's RSS keeps climbing?

level: seniorimportance: should knowfreq 45%

basics

~20 s

tracemalloc only accounts for blocks requested through CPython's own allocators after tracing started. Resident set size covers the whole process: the interpreter itself, thread stacks, and memory that compiled extensions or linked C libraries obtained straight from the system allocator.

open as a page

Why does storing a caught exception keep its traceback frames and their locals alive?

level: middleimportance: nice to knowfreq 25%

basics

~10 s

An exception holds its traceback in traceback, each traceback entry holds the frame it was raised from, and every frame holds its locals. Keeping the exception keeps the whole call chain's locals alive.

open as a page

When does time.thread_time() answer a question that time.process_time() cannot?

level: middleimportance: nice to knowfreq 15%

basics

~20 s

time.process_time() sums CPU time across every thread in the process, so one busy thread is hidden among the others. time.thread_time() reports CPU time for the calling thread only, which is how you charge processor cost to one worker.

open as a page

Why does timeit.timeit turn off garbage collection during the timed loop?

level: middleimportance: nice to knowfreq 20%

basics

~20 s

timeit disables the cyclic garbage collector around the timed loop and restores it afterwards, so a collection triggered by unrelated earlier allocations cannot land inside your measurement. It buys comparability at the cost of flattering allocation-heavy code.

open as a page

What does running with tracemalloc enabled cost, and how do you keep that cost bounded?

level: seniorimportance: nice to knowfreq 18%

basics

~20 s

Tracing adds bookkeeping to every allocation and stores a traceback per live block, so expect a substantial slowdown on allocation-heavy code plus real memory overhead. Keep nframe small, filter traces, and enable it on one instance rather than the fleet.

open as a page