skip to content

pymalloc and Arena Behaviour

CPython's small-object allocator sits in front of malloc, grouping allocations under 512 bytes into pools and arenas. It explains why resident memory stays high after the objects are gone.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

A Python inventory-sync worker's RSS climbs each batch and never falls: how do you tell a real leak from allocator retention?

level: seniorimportance: must knowfreq 50%

answer

  1. Three explanations, not two
  2. Measure the live count per cycle
  3. Flat count, rising RSS: not a leak
  4. Arena count against bytes in use
  5. Recycle on schedule, not reactively

basics

~20 s

Log sys.getallocatedblocks() once per batch. A live block count that returns to baseline while RSS does not means the objects are gone and pymalloc is holding fragmented arenas; a climbing count is a real leak, localised with tracemalloc.

solid answer

~40 s

Measure before theorising. `sys.getallocatedblocks()` logged once per batch is the decisive number: a count that returns to baseline while RSS plateaus means the objects really are deallocated and the memory is retained by the allocator, because pymalloc frees a 1 MiB arena only when every pool inside it is empty. A count that climbs batch over batch is a genuine leak, and `tracemalloc` snapshots compared with `Snapshot.compare_to` will name the allocation site. `sys._debugmallocstats()` confirms fragmentation directly: many arenas allocated against few bytes in allocated blocks. The fixes differ — a leak means finding the container still holding references; retention means lowering the high-water mark, keeping bulk data in one large buffer above the 512-byte threshold, or recycling the worker on a schedule that respects its 45-second cold start.

code

python · 10 lines
python
import gc, sys

records = [{"sku": i} for i in range(200_000)]
print("blocks at peak:", sys.getallocatedblocks())

survivors = records[::1000]   # keep 0.1% as a batch index
del records
gc.collect()
print("blocks after drop:", sys.getallocatedblocks())
sys._debugmallocstats()       # arenas still allocated: the survivors pin them

go deeper

for a junior

Know that memory staying high after work finishes is not automatically a leak, and that Python can tell you how many objects are actually still allocated rather than leaving you to guess from a process monitor.

for a middle

Be able to run the check: log the live block count each cycle, take two tracemalloc snapshots and diff them, and say which of the two curves — live objects or resident bytes — is the one that is actually rising.

for a senior

Demonstrate the full triage: separate live-object growth, allocator retention and non-Python memory; produce evidence for the one you name; and pick a remedy sized to the cause rather than sprinkling collector calls.

for a principal

Own the operating tradeoff: whether a service is designed to plateau at a known high-water mark or to return memory, what capacity headroom that implies, and whether scheduled recycling with warm replacements is acceptable given the cold-start cost.

### Separate the three claims before touching the code "Memory grows" can mean three different things, and the fix differs for each: 1. **Live Python objects are accumulating** — a genuine leak: a cache, a module-level list, a logging handler holding records, a closure capturing batch state. 2. **Objects are freed but pymalloc keeps the arenas** — allocator retention/fragmentation. Nothing is leaking; RSS simply ratchets to the high-water mark. 3. **Non-Python memory is growing** — a C extension's own allocations, thread stacks, or the system allocator's heap, none of which pymalloc reports on. The whole diagnosis is about telling these apart with evidence rather than guessing. ### Step 1 — is the block count flat? Instrument the worker to log `sys.getallocatedblocks()` at the same point in every batch — right after the batch is dropped, so you are comparing like with like. * Count returns to roughly the same baseline each cycle, while RSS keeps rising to a plateau → case 2. The objects genuinely are being deallocated. * Count climbs batch over batch → case 1. Something is holding Python objects. * Count flat *and* RSS rising without bound → case 3; the growth is outside the Python heap. This one number is the cheapest, most decisive measurement available, and it costs nothing in production. ### Step 2 — if it is a real leak, name the objects Start `tracemalloc` (optionally with `tracemalloc.start(25)` for deeper frames), take a snapshot at the end of an early batch and another many batches later, and diff them: ```python import tracemalloc tracemalloc.start() snap1 = tracemalloc.take_snapshot() # ... run several sync batches ... snap2 = tracemalloc.take_snapshot() for stat in snap2.compare_to(snap1, "lineno")[:10]: print(stat) ``` The diff ranks allocation sites by growth, which usually names the offending line directly. Pair it with `gc.get_objects()` counted by `type` if you want a type histogram rather than an allocation site. ### Step 3 — if it is fragmentation, confirm the shape `sys._debugmallocstats()` prints "# arenas allocated current" and "# bytes in allocated blocks". Fragmentation has an unmistakable signature: a large arena count multiplied by 1 MiB against a small live-bytes figure. A tiny fraction of survivors is enough to cause it, because an arena is returned to the OS only when every pool inside it is empty: ```python import gc, sys records = [{"sku": i} for i in range(200_000)] survivors = records[::1000] # keep 0.1% — an index of the batch del records gc.collect() sys._debugmallocstats() # arenas still allocated: the survivors pin them ``` In a run of exactly that shape the arena count went from 4 to 47 while the batch was live, stayed at 47 with only 200 survivors held, and dropped back to 5 once the survivors were dropped too — 0.1% of the objects pinning ~47 MB. ### Fixing case 2 * **Lower the high-water mark.** Stream the sync in smaller pages so the peak is never reached; an inventory feed rarely has to be materialised whole. * **Do not interleave survivors with the spike.** If each batch keeps a small index, build the batch, copy the survivors into a fresh structure, and drop the originals — the copies land in fresh pools instead of being sprinkled across every arena the batch touched. * **Hold bulk data above the 512-byte line.** One `bytes`/`bytearray`/`array.array` buffer bypasses pymalloc and is released as a unit rather than pinning arenas. * **Recycle the worker.** Restarting after N batches, or doing the heavy phase in a child process, returns the address space unconditionally. Weigh it against startup cost: a 45-second cold start means you recycle on a schedule you control, with a replacement warm before the old one exits — never reactively at the moment memory looks bad. And say plainly what does *not* help: `gc.collect()` frees cycles but never shrinks the process, and `del` on an already-dropped name changes nothing. Claiming a fix without a before/after block count is how these investigations go around twice.

  • Why can keeping 0.1% of a batch alive hold tens of megabytes resident?
    Because release is arena-granular. Survivors picked from across the batch land in blocks scattered over every arena the batch touched, and an arena goes back to the OS only when it is completely empty. Two hundred surviving dicts spread over 47 arenas keep all 47 megabytes. Copying survivors into a fresh structure before dropping the batch concentrates them and lets the rest go.
  • What if the block count is flat and RSS still grows without bound?
    Then the growth is not in the Python object heap. Look at allocations above the 512-byte threshold that go to the system allocator, at a C extension's own buffers, and at per-thread stacks if the worker keeps spawning threads. `sys.getallocatedblocks()` and `tracemalloc` will both look innocent, which is itself the diagnostic: you have ruled out Python objects and must measure at the process level instead.
  • How do you decide between periodic worker restarts and re-shaping the batch?
    By cost and blast radius. Restarts are unconditional and cheap to implement but pay the 45-second cold start and need a warm replacement before the old worker exits, which means overlap capacity. Re-shaping — smaller pages, bulk data in one buffer, survivors copied out — removes the pressure permanently and keeps the deploy simple. Restarts are the stopgap; re-shaping is the fix.

saying these in an interview costs you the question

  • Declares a leak from an RSS graph alone
  • Adds gc.collect() calls and calls the problem fixed
  • Never measures live objects across cycles
  • Assumes any flat RSS plateau is unbounded growth
  • Blames the C library without any allocator evidence
  • Restarts workers reactively despite the cold start

context

open as a page

Does deleting a large Python list return its memory to the operating system?

level: juniorimportance: should knowfreq 40%

basics

~20 s

Usually not right away. CPython deallocates the objects immediately, but its small-object allocator keeps the freed blocks for reuse and hands a 1 MiB arena back to the OS only once that whole arena is empty.

open as a page

How does CPython's pymalloc allocator handle small objects differently from malloc?

level: middleimportance: should knowfreq 35%

basics

~20 s

pymalloc groups requests of 512 bytes or less into size-class pools carved from 1 MiB arenas, so most Python objects are served by a free-list pop instead of a system call path. Larger requests fall through to malloc.

open as a page

What does the PYTHONMALLOC environment variable control in CPython?

level: middleimportance: nice to knowfreq 15%

basics

~20 s

PYTHONMALLOC picks which allocator CPython installs and whether debug hooks wrap it: pymalloc by default, malloc to route every request to the C library, and the debug variants to add guard bytes and fill patterns.

open as a page