skip to content

Does deleting a large Python list return its memory to the operating system?

level: juniorimportance: should knowfreq 40%

answer

  1. Freed inside the interpreter, not the kernel
  2. Reuse comes before release
  3. Blocks go back to a pool free list
  4. An arena leaves only when fully empty
  5. 512 bytes decides which allocator

basics

~20 s

Usually not right away. CPython deallocates the objects immediately, but its small-object allocator keeps the freed blocks for reuse and hands a 1 MiB arena back to the OS only once that whole arena is empty.

solid answer

~50 s

`del` drops a reference, and when the last reference goes the objects are deallocated at once, so the memory is free **inside** the interpreter. What the process returns to the kernel is a separate question. Requests of 512 bytes or less go through CPython's small-object allocator, pymalloc, which carves 16 KiB pools out of 1 MiB arenas; a freed block goes back on its pool's free list, and an arena is released to the OS only when every pool in it is empty. One surviving small object pins the whole megabyte. Anything over 512 bytes goes to the system allocator, which has its own retention rules and often does return a big block promptly. So resident memory typically plateaus at the high-water mark and later objects reuse that space. Flat-but-high RSS after a spike is normal, not a leak.

code

python · 9 lines
python
import gc, sys

base = sys.getallocatedblocks()
records = [{"sku": i} for i in range(200_000)]
peak = sys.getallocatedblocks()

del records
gc.collect()
print(base, peak, sys.getallocatedblocks())

go deeper

for a junior

Recall the two-layer answer: the object is freed immediately by reference counting, but the process keeps the pages for reuse. Saying 'the memory is reusable inside Python even though the OS still counts it' is enough at this level.

for a middle

Explain the mechanism: requests up to 512 bytes come from pools inside 1 MiB arenas, and an arena is released only when every pool in it is empty. Be ready to contrast that with a single large buffer, which bypasses pymalloc entirely.

for a senior

Show how you would prove it rather than assert it: log the live block count per cycle, compare against RSS, and conclude retention versus leak from the two curves. Then talk about lowering the high-water mark instead of fighting the allocator.

for a principal

Own the design consequence: if a service must return memory predictably, that is an architecture decision — bounded batch sizes, bulk data in single large buffers, or short-lived worker processes — not something you tune after the fact with allocator settings.

### Two different questions hide inside "is the memory freed?" When you write `del records` (or rebind the name, or let a local go out of scope), CPython drops one reference. If that was the last reference, the object is deallocated **immediately** — CPython's reference counting is deterministic, so there is no pause and no waiting for a collector. At that instant the memory the object occupied is free *inside the interpreter*. Whether the **process** gives those bytes back to the kernel is a separate question, and the answer is usually "not yet". That is what people are really asking when `del` on a huge list does not move the resident set size (RSS) they see in a process monitor. ### The layer that decides: pymalloc CPython does not call the C library's `malloc` once per Python object — an object graph of a few million dicts would spend most of its time in the system allocator. Instead, every request of **512 bytes or less** goes to CPython's own small-object allocator, *pymalloc*: * pymalloc requests memory from the OS in **arenas**, 1 MiB each on CPython 3.14 (they were 256 KiB before 3.10). * Each arena is cut into **pools** of 16 KiB. A pool is dedicated to one *size class* — 16, 32, 48 … up to 512 bytes, 32 classes in all, each rounded to a 16-byte boundary. * A pool hands out fixed-size **blocks** from a free list. Allocating a small object is a pointer pop; freeing it is a pointer push back onto that pool's free list. `sys._debugmallocstats()` prints exactly this structure, starting with the line `Small block threshold = 512, in 32 size classes.` Requests **larger than 512 bytes** skip pymalloc entirely and go to the system allocator. A 50 MB `bytes` object, or a list's internal pointer array once it grows past that threshold, is a plain `malloc`/`free` pair. ### Why the memory stays with the process Freeing a small object returns its block to a pool free list — memory the interpreter will reuse for the next object of that size class, but memory the OS still counts as yours. pymalloc releases an arena back to the OS **only when every pool in that arena is completely empty**. One surviving 50-byte object anywhere in a 1 MiB arena pins the entire megabyte. That is why RSS after a spike behaves like a ratchet: it climbs to the high-water mark of your workload and then sits there. It is not a leak — the space is reused. If you run the same batch again, RSS typically does not grow further, and that flatness is the signal that nothing is actually leaking. The system allocator adds its own retention on the >512-byte path. A glibc-style `malloc` satisfies large requests with `mmap` and can `munmap` them on `free` (so a single huge `bytes` object often *does* return promptly), but it keeps smaller chunks on its own free lists and only trims the heap when the top of it is free. ### Measuring it honestly `sys.getallocatedblocks()` reports how many pymalloc blocks are currently allocated. Watching it across a build-and-drop cycle separates the two questions cleanly: the block count returning to its baseline proves the objects really were deallocated, even while RSS stays flat at the peak. ```python import gc, sys base = sys.getallocatedblocks() records = [{"sku": i} for i in range(200_000)] peak = sys.getallocatedblocks() del records gc.collect() print(base, peak, sys.getallocatedblocks()) # e.g. 24236 623615 23504 ``` ### What to do about it If a plateau at the peak is unacceptable, you do not fight the allocator, you change the shape of the work: process the data in smaller chunks so the high-water mark is lower; hold bulk data in one large buffer object (`bytes`, `bytearray`, `array.array`) that bypasses pymalloc and is released in one piece; or run the memory-hungry phase in a short-lived child process so the whole address space goes away when it exits. Calling `gc.collect()` does not shrink the process — it only breaks reference cycles so the objects in them can be deallocated in the first place.

  • Does the same answer hold for a single 50 MB bytes object?
    Often not. Anything over 512 bytes bypasses pymalloc and goes to the system allocator, which typically satisfies a request that large with its own mapping and can return it to the kernel on free. So dropping one huge buffer frequently does lower RSS immediately, while dropping a million small dicts of the same total size usually does not.
  • Does calling gc.collect() shrink the process?
    No. The cycle collector only finds objects kept alive by reference cycles so they can be deallocated; it never moves objects and never returns pages by itself. After it runs, the freed memory sits in pymalloc's pools exactly as any other freed memory does. If RSS matters, lower the peak or run the heavy phase in a short-lived child process.
  • If the process plateaus and stays flat across many batches, is that a leak?
    No, and that flatness is the evidence. A leak keeps climbing because live objects accumulate; retention climbs to the workload's high-water mark and stops, because the same freed blocks are reused. Confirm by logging `sys.getallocatedblocks()` once per batch: a stable live block count with flat RSS is retention, a rising count is a leak.

The parking space you vacate is immediately free for the next car, but the garage keeps its whole floor leased until every space on it is empty.

saying these in an interview costs you the question

  • Says del hands the memory straight back to the OS
  • Calls a flat, high RSS plateau a memory leak
  • Thinks gc.collect() shrinks the process
  • Believes CPython never returns memory to the OS
  • Confuses deallocating an object with releasing pages

context