skip to content

How do you profile a script with cProfile and inspect the saved profile using pstats?

level: juniorimportance: should knowfreq 45%

answer

  1. One flag saves, another prints
  2. The saved file is re-readable forever
  3. A reader class wraps the saved table
  4. Shorten paths, sort, then cap rows
  5. Stats accepts a file or a live profiler

basics

~10 s

Run python -m cProfile -o out.prof script.py to capture a profile, then load it with pstats.Stats("out.prof") and call strip_dirs, sort_stats and print_stats to read the top rows.

solid answer

~40 s

For a whole program, `python -m cProfile -o out.prof script.py` writes a binary profile; `-s cumtime` instead prints a sorted report directly, and `-m` profiles a module run as `__main__`. Inside code, `cProfile.run("expr")` profiles a statement string, `runctx` supplies its globals and locals, and `cProfile.Profile` gives an object you can `enable()`/`disable()`, pass to `runcall`, or use as a context manager (supported since Python 3.8). To read a saved file, `pstats.Stats("out.prof")` then `.strip_dirs()` to shorten paths, `.sort_stats("cumtime")` to order rows, and `.print_stats(20)` — an int caps rows, a float takes a fraction, a string restricts by regex. `print_callers` attributes a hot function to its callers, `add` merges runs, and `python -m pstats out.prof` opens an interactive browser.

code

console · 3 lines
console
python -c 'open("s.py","w").write("print(sum(i*i for i in range(200000)))\n")'
python -m cProfile -o render.prof s.py
python -m cProfile -s cumtime s.py

go deeper

for a junior

Be able to run one command that profiles a script and one that reads the result. Knowing -o saves, sort_stats orders and print_stats(20) caps the output is the working minimum here.

for a middle

Explain the three entry points and when each fits: the CLI for a whole program, run/runctx for a statement, and Profile for a region of live code. Know what strip_dirs, the restriction arguments and print_callers actually do.

for a senior

Show the workflow around the tool: profile a realistic workload, keep the .prof files as baselines, merge runs before drawing conclusions, and confirm any win with a separate unprofiled measurement rather than a second profile.

for a principal

Argue for profiling as a repeatable practice rather than a one-off rescue: saved baselines, a standard capture recipe, and a stated policy that no performance claim ships on a profile alone.

### Three ways in, one way out Everything `cProfile` produces is the same object: a stats table that `pstats` reads. What differs is how you start the profiler. **The command line** wraps a whole program without editing it: ```console python -m cProfile -o render.prof render_invoices.py python -m cProfile -s cumtime render_invoices.py python -m cProfile -o render.prof -m mypackage.render ``` `-o` writes the binary profile to a file for later analysis; `-s` skips the file and prints a report sorted by the named key straight to stdout; `-m` profiles a module run as `__main__`, the same way `python -m` normally would. Use `-o` for anything you will look at more than once — the printed report is a one-shot summary, while the `.prof` file can be re-sorted, filtered and merged forever. **`cProfile.run(...)`** profiles a statement given as a string, executed in a fresh `__main__` namespace, and either prints the report or writes it to the filename you pass as the second argument. `cProfile.runctx(...)` is the same with explicit `globals` and `locals` dicts, which is what you need when the statement has to see objects you already built. **`cProfile.Profile`** is the object form, and the one that scales to real code: `enable()` / `disable()` around a region, `runcall(func, *args)` for one call, or — since **Python 3.8** — the `with` statement, because `Profile` supports the context-manager protocol. This is how you profile a slice of a long-lived process rather than a whole script. ```python import cProfile, pstats def render(pages): return sum(len(f"invoice page {p}".encode("utf-8")) for p in range(pages)) with cProfile.Profile() as prof: render(20000) prof.dump_stats("render.prof") pstats.Stats("render.prof").strip_dirs().sort_stats("cumtime").print_stats(5) ``` ### Reading the file back `pstats.Stats` is the reader. Its constructor takes a filename, several filenames, or a live `Profile` object, and an optional `stream=` to send output somewhere other than stdout. The methods you will actually use: * **`strip_dirs()`** — collapses the absolute file paths to bare module names. Purely cosmetic, but a report full of long interpreter paths is unreadable, so it is nearly always the first call. It is irreversible on that `Stats` object; reload the file if you need the paths back. * **`sort_stats(key)`** — orders the rows. `"tottime"` (self time), `"cumtime"` (inclusive time), `"ncalls"` and `"nfl"` (name/file/line) are the ones worth remembering. Multiple keys break ties. * **`print_stats(...)`** — prints the table. With no argument it prints every row, which for a real program is thousands of lines. Pass an integer to cap the row count, a float between 0 and 1 to print that fraction of rows, or a regex string to restrict to matching function names — and you can pass several restrictions, applied in order. * **`print_callers(...)` / `print_callees(...)`** — the caller/callee attribution for the rows that survive the same kind of restriction, which is how you find *who* is driving a hot shared helper. * **`add(*filenames)`** — merges further profiles into this one, summing counts and times. Useful for aggregating several runs or several inputs before you draw conclusions from one sample. * **`dump_stats(filename)`** — writes the (possibly merged) table back out. There is also an interactive browser: `python -m pstats render.prof` opens a small prompt where `sort cumtime`, `stats 20` and `callers <regex>` do the same work without writing a script. It is handy when you are exploring rather than automating. Because the `.prof` file is just a marshalled table, third-party visualizers can render it as a flame-graph or a call graph, and that is often the fastest way to see the shape of a large profile. Nothing about the capture step changes; you still produce the file with `-o` or `dump_stats`. ### The habits that make the output worth reading Profile the **realistic** workload, not a toy one: an invoice renderer profiled over three pages will put its fixed setup at the top of the report and hide the per-page cost that dominates a real run. Profile a **whole** unit of work rather than sprinkling `enable()`/`disable()` around fragments — the overlapping regions are hard to reason about, and you lose the entry-point rows that make `cumtime` useful. Keep the `.prof` files: the interesting comparison is almost always before-and-after, and a saved baseline is what turns "it feels faster" into a number. Finally, remember what the report is measuring. It is CPU-time-shaped, per-call attribution of a single deterministic run, not a benchmark. Use it to decide *what* to optimise; confirm the win with a separate, unprofiled measurement of the same work.

  • When would you reach for cProfile.Profile instead of the module-level cProfile.run function?
    Whenever you need to profile a slice of a live program rather than a whole script. `cProfile.run` takes a statement as a **string**, executed in a fresh namespace, which is awkward when the code depends on objects you already built. `Profile` lets you wrap exactly the region with `enable()`/`disable()`, `runcall(func, *args)`, or a `with` block, and hand the object straight to `pstats.Stats`.
  • How do you combine profiles captured from several separate runs?
    `pstats.Stats` accepts multiple filenames in its constructor, and `Stats.add(*filenames)` merges more into an existing object; call counts and times sum. That is the normal way to aggregate several inputs or several worker runs into one report before sorting, and `dump_stats` writes the merged table back out.
  • What does passing a string to Stats.print_stats do?
    It is treated as a regular expression restricting which rows print, matched against the `filename:lineno(function)` text. You can combine restrictions — for example an int and a regex — and they are applied in order, so `print_stats("render", 10)` narrows to matching rows and then caps the output at ten of them.

saying these in an interview costs you the question

  • Thinks -o prints the report instead of saving it
  • Calls print_stats with no argument on a large program
  • Believes cProfile.run can profile arbitrary objects, not a string
  • Cannot name a way to sort the report at all
  • Assumes the .prof file is human-readable text
  • Confuses the profiler module with the report reader module

context