skip to content

How does CPython's pymalloc allocator handle small objects differently from malloc?

level: middleimportance: should knowfreq 35%

answer

  1. Too many tiny objects for malloc
  2. Three levels of granularity
  3. One pool serves one size class
  4. Free-list pop, no system call
  5. 512-byte threshold, 16 KiB pools, 1 MiB arenas

basics

~20 s

pymalloc groups requests of 512 bytes or less into size-class pools carved from 1 MiB arenas, so most Python objects are served by a free-list pop instead of a system call path. Larger requests fall through to malloc.

solid answer

~50 s

Python objects are small and numerous, so calling the C library once per object would dominate object creation. CPython interposes pymalloc for every request of **512 bytes or less**: it obtains 1 MiB **arenas** from the OS, splits each into 16 KiB **pools**, and dedicates a pool to one size class — multiples of 16 bytes up to 512, 32 classes in all. Allocating is a free-list pop, freeing is a push; no system allocator call is involved. Requests above 512 bytes go straight to `malloc`. `sys._debugmallocstats()` prints this whole structure, starting with `Small block threshold = 512, in 32 size classes.`, and `sys.getallocatedblocks()` reports the live pymalloc block count. In the standard build pymalloc relies on the GIL for mutual exclusion, which is why the free-threaded build uses the thread-safe mimalloc as its object allocator instead.

code

python · 6 lines
python
import sys

print("live pymalloc blocks:", sys.getallocatedblocks())
# prints the size-class table and arena summary to stderr,
# starting with: Small block threshold = 512, in 32 size classes.
sys._debugmallocstats()

go deeper

for a junior

Know the headline: CPython has its own allocator for small objects sitting in front of the C library's, so creating many tiny objects is cheap. The exact pool and arena sizes are not expected yet.

for a middle

Explain the arena/pool/block hierarchy and the 512-byte cutoff, and why batching small allocations into size-class free lists is faster than one malloc per object. Naming sys._debugmallocstats() as where you would read it off shows real familiarity.

for a senior

Connect the layout to behaviour you have actually observed: arena-granular release, fragmentation from scattered survivors, and the different profile of bulk data held above the 512-byte line.

for a principal

Frame the tradeoff being bought: throughput on small-object churn in exchange for arena-granular fragmentation and a GIL dependency — which is exactly what the free-threaded build had to renegotiate when it adopted a thread-safe allocator.

### The problem pymalloc solves Almost every Python value is a heap object with a header, and typical objects are tiny: a small `tuple`, a `float`, a dict entry table, an `int`. A program that builds a million of them would make a million `malloc` calls and a million `free` calls. General-purpose system allocators are built for a mix of sizes and for arbitrary threads; paying their bookkeeping for a 32-byte object dominates the cost of creating it. CPython therefore puts its own **small-object allocator**, pymalloc, in front of the system allocator. On CPython 3.14's standard build it handles every request of **512 bytes or less**; anything larger falls straight through to `malloc`. ### The three-level layout: arena, pool, block * **Arena** — the unit pymalloc buys from the OS, 1 MiB on 3.14 (raised from 256 KiB in 3.10), obtained with `mmap` where available. * **Pool** — a 16 KiB slice of an arena (4 KiB before 3.10). A pool is assigned to exactly one *size class* while it holds anything. * **Block** — a fixed-size slot inside a pool. Sizes are multiples of 16 bytes: 16, 32, 48 … 512, giving 32 size classes. A 40-byte request is rounded up to the 48-byte class, and the difference is what `sys._debugmallocstats()` calls "bytes lost to quantization". Allocation is a free-list pop from the pool that currently serves that size class; deallocation is a push back onto it. Both are a handful of instructions with no system call, which is the entire point. Empty pools return to their arena's pool pool, and a fully empty arena is handed back to the OS. ### What it means in practice **Fragmentation is arena-shaped.** Because an arena is released only when it is entirely empty, a few long-lived small objects scattered across many arenas can keep megabytes resident long after the bulk of the data is gone. This is the mechanism behind "my RSS never comes back down". **The 512-byte line changes behaviour.** Bulk data held in one large object — a `bytes`, a `bytearray`, an `array.array`, a big buffer viewed through a `memoryview` — never touches pymalloc, so its memory is returned on the system allocator's terms rather than pinned by arena occupancy. Holding the same bytes as a million small objects is the opposite case. **pymalloc is not general-purpose thread-safe.** In the standard build it relies on the GIL for mutual exclusion, which is why the free-threaded build (experimental in 3.13, officially supported in 3.14 by PEP 779) uses **mimalloc** as its object allocator instead: mimalloc is thread-safe by design with per-thread heaps. ### Inspecting it Two stdlib entry points expose the allocator directly: ```python import sys print(sys.getallocatedblocks()) # live pymalloc blocks sys._debugmallocstats() # size-class table, pool and arena counts, to stderr ``` `sys._debugmallocstats()` prints the per-size-class table (pools in use, blocks in use, available blocks), then the arena summary: arenas allocated, bytes in allocated blocks, bytes in available blocks, unused pools, and bytes lost to pool headers, quantization and alignment. Comparing *bytes in allocated blocks* against *arenas × 1 MiB* is a direct read on how fragmented the heap is. `sys.getallocatedblocks()` counts pymalloc blocks specifically — under `PYTHONMALLOC=malloc` it reports 0, because pymalloc is not the one doing the allocating. ### The boundaries worth stating in an interview pymalloc sits between Python objects and the C library, not between your program and virtual memory: it does not do compaction, it never moves an object (C extensions hold raw pointers into these blocks, so addresses must be stable), and it has nothing to do with reclaiming *garbage* — deciding that an object is dead is reference counting's job, and breaking cycles is the collector's. pymalloc only answers "where do the bytes come from, and where do they go back to".

  • What happens to a request of exactly 40 bytes?
    It is rounded up to the next size class. Classes are multiples of 16 bytes, so a 40-byte request is served from a 48-byte block and the 8 bytes of difference are what `sys._debugmallocstats()` reports as bytes lost to quantization. Rounding is what lets a pool hold uniform, interchangeable blocks so allocation stays a free-list pop.
  • Does pymalloc ever move an object to compact the heap?
    No. C extensions and the interpreter itself hold raw pointers into these blocks, so object addresses must stay stable for the object's lifetime. pymalloc has no compaction and no relocation; the only defragmentation available is emptying an arena completely so it can be released, which is why long-lived small objects scattered across arenas are so costly.
  • Why does the free-threaded build not just keep pymalloc?
    pymalloc's pool and arena bookkeeping is not thread-safe on its own; in the standard build the GIL serialises access to it. Without a GIL every allocation would need locking, which would cost more than it saves, so the free-threaded build ships mimalloc, which is designed around per-thread heaps and thread-safe free lists.

saying these in an interview costs you the question

  • Says every Python object is a separate malloc call
  • Thinks pymalloc handles allocations of any size
  • Confuses pymalloc with garbage collection
  • Claims pymalloc compacts or moves objects
  • Describes a pool as holding mixed object sizes

context