skip to content

tracemalloc and Memory Profiling

Finding where memory was allocated rather than how much is in use: tracemalloc attributes allocations to source lines and diffs two snapshots. The tool to name when chasing a slow leak in a service.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How does tracemalloc's Snapshot.compare_to help you find a slow leak in a long-running service?

level: middleimportance: must knowfreq 50%

answer

  1. One snapshot hides the baseline
  2. Measure change, not the total
  3. Warm up, snapshot, cycle, snapshot
  4. compare_to sorts by size_diff
  5. count_diff separates many-objects from bigger-buffers

basics

~20 s

Take one snapshot after warm-up and another after many work cycles, then call second.compare_to(first, 'lineno'). The result is sorted by size_diff, so the lines that grew between the two snapshots rise to the top and steady-state memory cancels out.

solid answer

~40 s

A single snapshot is dominated by memory that is *supposed* to be there — caches, loaded modules, the connection pool — so a leak hides inside it. Diffing removes that baseline. Start tracing, let the service reach steady state, take a first snapshot, run a few hundred identical cycles, then take a second and call `second.compare_to(first, 'lineno')`. That returns `StatisticDiff` objects sorted by `size_diff` descending; each also carries `count_diff`, plus the absolute `size` and `count` and the `traceback`. Anything genuinely per-cycle nets out to roughly zero, so a line whose `size_diff` grows in proportion to the number of cycles is your suspect. Repeat with more cycles to confirm the growth is linear rather than a warm-up artefact; `Snapshot.filter_traces()` trims noise from library code.

code

python · 21 lines
python
import tracemalloc

registry = {}

def import_records(batch):
    rows = [bytes(512) for _ in range(200)]
    registry[batch] = rows          # accidentally kept forever
    return len(rows)

tracemalloc.start()
import_records(0)
first = tracemalloc.take_snapshot()

for batch in range(1, 21):
    import_records(batch)

second = tracemalloc.take_snapshot()
for stat in second.compare_to(first, "lineno")[:3]:
    print(stat.size_diff, "bytes", stat.count_diff, "blocks")
    print("   ", stat.traceback.format()[-1].strip())
tracemalloc.stop()

go deeper

for a junior

Know that two snapshots can be subtracted: take one, do the work repeatedly, take another, and compare_to shows what grew. The point to remember is that a diff removes the memory that was always there.

for a middle

Explain the mechanics end to end: warm-up first, many identical cycles, second.compare_to(first, 'lineno'), and how to read size_diff against count_diff to tell an unbounded collection from a buffer that keeps being extended.

for a senior

Demonstrate the discipline around the measurement — proving growth is proportional to cycles, filtering library noise, dumping snapshots from a live service on a schedule, and knowing that a located line is the start of the investigation rather than the end.

for a principal

Own how leak evidence is produced and read across teams: what is captured routinely versus on demand, the overhead you are willing to run with, and when a slow leak is worth chasing rather than absorbing with a restart policy.

### Why one snapshot is not enough Ask a healthy Python service what is holding memory and the honest answer is "almost everything you would expect": imported modules, compiled regexes, a connection pool, a deliberate cache, the framework's own structures. A leak of a few megabytes an hour is invisible inside that. The signal you actually want is not *what is allocated* but *what is more allocated than it was an hour ago*, and that is exactly what `tracemalloc.Snapshot.compare_to()` computes. ### The procedure ```python import tracemalloc tracemalloc.start(10) warm_up() # let caches and pools fill first = tracemalloc.take_snapshot() for _ in range(500): handle_one_cycle() second = tracemalloc.take_snapshot() for stat in second.compare_to(first, "lineno")[:10]: print(stat) ``` Three details in that sketch carry the method. **Warm up before the first snapshot**, or every lazily-built cache in the process shows up as growth. **Run many identical cycles**, so a per-cycle leak is multiplied into something obvious while one-off costs stay one-off. **Compare, do not read.** `compare_to` returns a list of `tracemalloc.StatisticDiff` objects sorted by `size_diff` descending by default. ### Reading a StatisticDiff A `StatisticDiff` carries both the delta and the absolute figures: * `size_diff` — bytes gained (or lost, if negative) at that grouping since the old snapshot; * `count_diff` — the change in the number of live blocks; * `size` and `count` — the current absolute totals, so you can see whether a 4 MB gain sits on top of 4 MB or of 400 MB; * `traceback` — where those blocks were allocated. The pairing of `size_diff` with `count_diff` is what makes the diagnosis. Growth in both, in step, is the classic unbounded collection: more objects, each about the same size, accumulating somewhere. `size_diff` climbing while `count_diff` stays flat is a small number of buffers being reallocated ever larger — a list that keeps being extended, a `bytearray` that keeps growing. A large positive `count_diff` with a modest `size_diff` points at many tiny objects, where per-object overhead rather than payload is the story. ### A worked shape Imagine a museum-catalogue importer that ingests one institution's records per run and has been drifting upward across a three-week release train — nothing dramatic, a few megabytes per import, until the container is restarted weekly to stay under its limit. A single snapshot of that process shows a large parsed-record structure at the top, which is exactly what an importer is supposed to hold, and tells you nothing. Snapshot the process after the first import, run twenty more, snapshot again: now the top `StatisticDiff` is a line that had no business growing at all — a module-level dictionary keyed by record id, a list of parsed rows appended to for logging, a memoisation decorator applied to a method so the cache outlives every instance. The diff turned an ambient number into a line of code. ### Making the signal cleaner `Snapshot.filter_traces()` takes a sequence of `tracemalloc.Filter` objects and returns a new snapshot with only the traces you care about — typically excluding the standard library and third-party site-packages so your own modules stand out. Filtering both snapshots the same way before comparing keeps the diff honest. For a service you cannot keep a Python session attached to, `Snapshot.dump()` writes a snapshot to a file and the `Snapshot.load()` classmethod reads it back, so you can capture on the box hourly and diff any two of them later on your laptop. Keep tracing on with a small `nframe` for that; snapshots taken with different settings still compare, but the traces they carry will not line up in a `'traceback'` grouping. ### Confirming before you act Growth is evidence, not a verdict. A bounded cache filling toward its limit grows between your two snapshots and then stops. The check is proportionality: double the number of cycles and the true leak's `size_diff` roughly doubles, while the warm-up artefact does not. Take three snapshots rather than two if you can afford it, and compare the second pair as well. Once a line is identified, the remaining question — which object still refers to those blocks, so that they are never collected — is a different investigation with different tools. `compare_to` reliably answers *where the memory was born*; it does not answer *who is still holding it*.

  • In a tracemalloc diff, what does growing size_diff with a flat count_diff tell you?
    That the same small number of blocks is getting bigger rather than more blocks appearing — a container being repeatedly extended and reallocated, a `bytearray` or buffer accumulating, a string being rebuilt ever longer. Growth in both figures together is the opposite shape: an unbounded collection gathering more objects of similar size. The pair narrows the fix before you read any code.
  • How do you rule out that the growth between two snapshots is just a cache warming up?
    Test proportionality. Run twice as many cycles between the snapshots; a real per-cycle leak roughly doubles its `size_diff` while a bounded cache filling toward its ceiling does not. Taking three snapshots and diffing the later pair as well is the same check in one run — if growth continues at the same rate after the cache should be full, it is not the cache.
  • How would you diff snapshots from a service you cannot attach an interactive session to?
    Write them to disk with `Snapshot.dump()` on a schedule or on a signal, then load any two later with the `Snapshot.load()` classmethod and call `compare_to` offline. Keep `nframe` small so the traces stay cheap, and dump to a path with room for them. This also gives you a history, so you can compare a snapshot from before a release with one from after.

It is stock-taking rather than a shelf survey: counting everything in the warehouse tells you little, but counting the same shelves twice a week apart tells you exactly what nobody is shipping out.

saying these in an interview costs you the question

  • Hunts a leak from a single snapshot's biggest entry
  • Snapshots before the service has warmed up
  • Treats any growth between snapshots as a leak
  • Confuses size_diff with the process's memory change
  • Ignores count_diff and so misreads the leak's shape

context

open as a page

How do you use tracemalloc to find which lines of a script allocated the most memory?

level: juniorimportance: should knowfreq 25%

basics

~10 s

Call tracemalloc.start() before the work, tracemalloc.take_snapshot() after it, then sort the snapshot with Snapshot.statistics('lineno'). Each entry names a file and line and reports the bytes and block count still allocated there.

open as a page

Why can tracemalloc report flat Python memory while the process's RSS keeps climbing?

level: seniorimportance: should knowfreq 45%

basics

~20 s

tracemalloc only accounts for blocks requested through CPython's own allocators after tracing started. Resident set size covers the whole process: the interpreter itself, thread stacks, and memory that compiled extensions or linked C libraries obtained straight from the system allocator.

open as a page

What does running with tracemalloc enabled cost, and how do you keep that cost bounded?

level: seniorimportance: nice to knowfreq 18%

basics

~20 s

Tracing adds bookkeeping to every allocation and stores a traceback per live block, so expect a substantial slowdown on allocation-heavy code plus real memory overhead. Keep nframe small, filter traces, and enable it on one instance rather than the fleet.

open as a page