skip to content

Why can tracemalloc report flat Python memory while the process's RSS keeps climbing?

level: seniorimportance: should knowfreq 45%

answer

  1. Two tools, two definitions of memory
  2. One counts blocks, one counts pages
  3. Only what CPython's allocators handed out
  4. Compiled code can bypass the traced heap
  5. Log both series, compare the shapes

basics

~20 s

tracemalloc only accounts for blocks requested through CPython's own allocators after tracing started. Resident set size covers the whole process: the interpreter itself, thread stacks, and memory that compiled extensions or linked C libraries obtained straight from the system allocator.

solid answer

~40 s

The two numbers measure different things. `tracemalloc` instruments CPython's memory allocators and reports the bytes of traced blocks currently alive — a Python-heap figure, and only for allocations made after `tracemalloc.start()`. RSS is what the operating system says the process has resident: the interpreter binary and its data, thread stacks, memory-mapped files, allocations made by compiled extension modules or linked C libraries that call the system allocator directly, and memory the process has freed internally but not handed back. So a divergence usually means the growth is happening outside the traced heap. The move is to measure both over time — log `tracemalloc.get_traced_memory()` beside an OS-level resident figure — and if only RSS rises, the investigation shifts to native code and allocator behaviour rather than to Python objects.

code

python · 10 lines
python
import tracemalloc

tracemalloc.start()
keep = [bytearray(1_000_000) for _ in range(20)]
current, peak = tracemalloc.get_traced_memory()
print(f"traced current {current / 1e6:.1f} MB, peak {peak / 1e6:.1f} MB")
del keep
current, peak = tracemalloc.get_traced_memory()
print(f"after del: current {current / 1e6:.1f} MB, peak {peak / 1e6:.1f} MB")
tracemalloc.stop()

go deeper

for a junior

Remember that tracemalloc measures memory Python itself allocated, not the number your process monitor shows. The two will not match, and that is expected rather than a bug in either tool.

for a middle

Be able to list what sits in the gap: the interpreter's own footprint, thread stacks, memory-mapped files, buffers allocated by compiled extensions, and anything created before tracing started.

for a senior

Show the diagnostic routine — log both series over time and read their shapes together, then say which class of cause each combination points at, and report 'no growth in the traced heap' rather than 'no leak'.

for a principal

Own the observability contract: which memory series are collected for every service, what the restart and limit policy assumes about them, and how an investigation escalates from Python-heap evidence to native tooling.

### Two numbers, two definitions The single most common confusion in a Python memory investigation is treating `tracemalloc`'s totals and the process's resident set size as the same quantity measured two ways. They are not. `tracemalloc` hooks CPython's memory-allocator layer. Every block requested through it — which is to say every Python object, every dict table, every list slot array — is recorded with the source location that asked for it, and `tracemalloc.get_traced_memory()` returns the `(current, peak)` byte totals of blocks it traced. It is a *Python heap* figure, scoped to allocations made after `tracemalloc.start()`. RSS is an operating-system accounting of pages the process currently has resident in physical memory. It includes the traced heap, and then a great deal more. ### What lives in the gap **The interpreter itself.** The binary, its static data, imported extension modules, compiled bytecode objects created before tracing began. A bare interpreter has tens of megabytes of RSS before your program does anything. **Allocations that never touched CPython's allocators.** A compiled extension module is free to call the system allocator directly for its own working buffers, and many do for anything large. Those bytes are entirely invisible to tracing, no matter how big. If a library that wraps a native codec, a compression engine or a numeric kernel grows a buffer per request, `tracemalloc` will show a flat Python heap while RSS marches upward. The same applies to memory a linked C library allocates internally, and to memory-mapped files, whose pages count toward RSS while nothing was allocated at all. **Thread and interpreter machinery.** Each OS thread gets a stack; each additional interpreter created through the standard library's multiple-interpreter support carries its own state. None of it is a traced Python object. **Memory freed inside the process but not returned to the operating system.** Freeing a Python object reduces the traced total immediately, but the underlying pages are not necessarily unmapped, so RSS can stay at its high-water mark long after a spike is over. This is normal allocator behaviour, not a leak. **Everything allocated before tracing started.** Even for pure Python objects, if the accumulation began at import time and you started tracing in a request handler, the existing bulk is not in the traced total. Only its subsequent growth is. ### Diagnosing the divergence The useful discipline is to record both series and look at their shapes rather than a single reading of either. * If **RSS and traced memory rise together**, the growth is in Python objects. Snapshot diffing will name the line, and the investigation proceeds inside your own code. * If **RSS rises while traced memory is flat**, the growth is outside the traced heap. Suspect native allocations from an extension, an unbounded native cache, file descriptors and mappings, or thread stacks from threads that are never joined. Tools that read the operating system's own accounting are the right instrument here, not a deeper Python snapshot. * If **traced memory rises while RSS is flat**, you are usually looking at reuse: the process already had the pages, so a growing Python heap fits inside memory it holds. The traced figure is the leading indicator and worth acting on before RSS confirms it. * If **both are flat but the container is being killed**, look outside the heap entirely — a child process, a temporary file on a memory-backed filesystem, a limit counted differently from RSS. ### Practical consequences Two consequences follow for how you report a memory problem. First, never claim "there is no leak" on the strength of a flat `tracemalloc` total; the honest statement is "there is no growth in the traced Python heap", which is a much narrower claim and points the next person at native code rather than sending them round the same loop. Second, start tracing as early as possible — `PYTHONTRACEMALLOC` or the equivalent `-X tracemalloc` option enables it before the first import — so that "allocated before tracing began" is a small category rather than the bulk of the process. The complementary tools split along the same line. Line-oriented memory profilers that sample the process's resident usage as a function executes see the native allocations `tracemalloc` misses, at the cost of attributing them only to the Python frame that was running. Object-graph inspectors count live instances by type, which is the right lens when the traced heap *is* growing and you want to know what kind of object is accumulating. `tracemalloc` sits between them: precise about where a Python-heap block was born, and silent about everything the Python heap does not own.

  • Traced memory is flat but RSS climbs steadily. What do you investigate next?
    Anything that allocates outside CPython's allocators: compiled extension modules holding native buffers, a linked C library with its own cache, memory-mapped files, thread stacks from threads that are started and never joined, and child processes counted against the same limit. Operating-system level accounting and native heap tooling are the instruments; taking deeper Python snapshots will keep telling you the same non-answer.
  • Why can tracemalloc's current total drop sharply while the process's RSS does not?
    Freeing a Python object immediately reduces the traced total because the block is no longer alive, but the pages behind it stay mapped in the process for reuse rather than being returned to the operating system. RSS therefore tends to track the high-water mark. That is ordinary allocator behaviour, and reading it as a leak sends an investigation in the wrong direction.
  • How do you make sure tracing covers allocations made during import?
    Enable it before the program's first line rather than in code: the `PYTHONTRACEMALLOC` environment variable or the equivalent `-X tracemalloc` interpreter option starts tracing at interpreter startup and takes a frame count. Without that, anything built at import time is absent from every snapshot, and a module-level structure that grows once at startup will never appear.

tracemalloc is the till receipt for one department; RSS is the building's electricity bill. A flat receipt does not mean nothing else in the building is drawing power.

saying these in an interview costs you the question

  • Declares no leak because traced memory is flat
  • Assumes tracemalloc sees native extension allocations
  • Treats RSS and Python heap size as one number
  • Expects RSS to fall the instant objects are freed
  • Starts tracing late, then trusts absolute totals

context