skip to content

Performance and Memory

Where Python spends time and memory: the CPython object model, the profilers that find the real cost, and the idioms and native escape hatches that fix it. It separates measurers from guessers.

part ofPythonoverview, primer and where to startread it →
on this pageshow

explore

questions

91 · 5 sections

Why can reference counting alone never free two Python objects that point at each other?

level: juniorimportance: must knowfreq 55%
basics
~20 s

Each object in the cycle holds a reference to the other, so neither count drops to zero even after every outside name is gone. CPython runs a second, tracing collector — the gc module — to find and free such groups.

open as a page

In CPython, what does `del x` actually do — does it free the object?

level: juniorimportance: must knowfreq 62%
basics
~20 s

del x only unbinds the name x from its namespace and drops one reference to the object. CPython deallocates the object at the moment its reference count reaches zero, so if a list, attribute or another name still holds it, nothing is freed.

open as a page

What does weakref.ref give you that an ordinary Python reference does not?

level: juniorimportance: must knowfreq 40%
basics
~10 s

A weakref.ref points at an object without counting toward the references that keep it alive. Call the ref like a function to get the object back, or None once the object has been collected.

open as a page

How much memory does adding __slots__ to a Python class actually save?

level: middleimportance: must knowfreq 60%
basics
~20 s

__slots__ fixes the attribute names up front so instances store values at known offsets and carry no per-instance dictionary. On CPython 3.14 that is roughly a 35-45% cut for a small class - real, but far less than the old folklore.

open as a page

A Python inventory-sync worker's RSS climbs each batch and never falls: how do you tell a real leak from allocator retention?

level: seniorimportance: must knowfreq 50%
basics
~20 s

Log sys.getallocatedblocks() once per batch. A live block count that returns to baseline while RSS does not means the objects are gone and pymalloc is holding fragmented arenas; a climbing count is a real leak, localised with tracemalloc.

open as a page

Why does a Python `import` statement execute code rather than just bind a name?

level: juniorimportance: must knowfreq 50%
basics
~20 s

The first import of a module runs its whole body top to bottom, then stores the finished module object so later imports are just a cache lookup. Every top-level statement in it is startup cost.

open as a page

Why do objects appended to a module-level list never get freed in Python?

level: juniorimportance: must knowfreq 55%
basics
~20 s

A module-level list stays reachable from sys.modules for the whole process, so every object it holds keeps a live reference and can never be freed. Python reclaims an object only when nothing reachable still refers to it.

open as a page

What is the difference between wall-clock time and CPU time when timing Python code?

level: juniorimportance: must knowfreq 55%
basics
~20 s

Wall-clock time is elapsed real time, read with time.perf_counter(). CPU time is the time the processor actually spent executing the process, read with time.process_time(). Sleeping, waiting on a socket or blocking on a lock adds wall time but no CPU time.

open as a page

Why does timeit.timeit take a separate setup argument instead of one block of code?

level: juniorimportance: must knowfreq 55%
basics
~20 s

timeit runs the setup string once, before the clock starts, and only the stmt string inside the timed loop. Imports and test-data construction are therefore excluded, so you measure the operation itself rather than its fixtures.

open as a page

In a cProfile report, what is the difference between tottime and cumtime?

level: middleimportance: must knowfreq 55%
basics
~20 s

tottime is the time spent inside a function's own body, excluding the calls it makes. cumtime adds everything it called. Sort by tottime to find hot code, by cumtime to find the expensive call path.

open as a page

Why is ''.join(parts) preferred over building a string with += in a loop?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Strings are immutable, so each += builds a brand-new string and copies everything accumulated so far. str.join walks the sequence once in C, sizes the result, allocates it once, and copies each piece exactly once.

open as a page

Why don't changes a multiprocessing child process makes to a global variable reach the parent?

level: juniorimportance: must knowfreq 55%
basics
~20 s

Each multiprocessing worker is a separate OS process with its own address space and its own copy of every module-level name. What the child rebinds or mutates changes only the child's memory; results must be sent back explicitly.

open as a page

What is the difference between a generator expression and a list comprehension in Python?

level: juniorimportance: must knowfreq 78%
basics
~20 s

A list comprehension builds the whole list in memory at once. A generator expression builds nothing up front: it produces items one at a time on demand, holds only the current value, can be consumed once, and supports neither len() nor indexing.

open as a page

Why does functools.lru_cache raise TypeError when a list is passed as an argument?

level: middleimportance: must knowfreq 55%
basics
~20 s

The wrapper builds a key from the call's arguments and looks it up in a dictionary, so every argument must be hashable. A list is not, and the lookup raises TypeError: unhashable type: 'list' before the function body ever runs.

open as a page

How do you stream a 40 GB log file in Python without loading it into memory?

level: middleimportance: must knowfreq 70%
basics
~20 s

Iterate the open file object directly with a for loop: it is its own iterator and hands back one buffered line at a time. Never call read() or readlines(), and keep every downstream stage lazy so nothing collects the whole file.

open as a page

Why is reading a local variable cheaper than reading a global in CPython?

level: middleimportance: must knowfreq 55%
basics
~20 s

A local read compiles to a LOAD_FAST-family instruction, an index into the frame's array of local slots. A global read compiles to LOAD_GLOBAL, which hashes the name and searches the module dictionary, then builtins. An array index beats a hash lookup.

open as a page

Why is an element-wise Python loop far slower than one vectorized array-library call?

level: middleimportance: must knowfreq 62%
basics
~20 s

The loop pays interpreter overhead on every element: bytecode dispatch, a heap-allocated object per value, reference-count updates and runtime type checks. One vectorized call pays that once, then runs a typed machine-code loop over contiguous memory.

open as a page

How does CPython's specializing adaptive interpreter (PEP 659) speed up hot code?

level: middleimportance: must knowfreq 42%
basics
~20 s

CPython watches how each bytecode instruction is actually used and rewrites it in place into a form specialized for the types it keeps seeing, behind a cheap guard check. No machine code is produced; the interpreter simply does less work per instruction.

open as a page

Why do C extensions block PyPy adoption, and which bindings port cleanly?

level: seniorimportance: must knowfreq 34%
basics
~20 s

CPython's C API hands extensions raw pointers to objects and makes them maintain reference counts. PyPy has a moving collector and no reference counts, so it emulates that API — costly at every crossing and not always complete.

open as a page

What does Python's `dis.dis()` print when you pass it a function?

level: juniorimportance: should knowfreq 35%
basics
~20 s

dis.dis(f) prints the CPython bytecode the compiler produced for f: one line per instruction, with the source line, the byte offset, the opcode name and its resolved argument. It shows what the interpreter executes, not how long it takes.

open as a page

Why slice a memoryview of a large bytes object instead of the bytes itself?

level: juniorimportance: must knowfreq 45%
basics
~20 s

Slicing a bytes object allocates a new bytes and copies the requested range. A memoryview exposes the same memory through the buffer protocol, so slicing it returns another view over the original block in constant time, copying nothing.

open as a page

What happens to the GIL while zlib.compress runs on a large buffer?

level: juniorimportance: must knowfreq 55%
basics
~20 s

The C code inside zlib drops the global interpreter lock before it starts deflating and takes it back when the call finishes. Other Python threads run real bytecode during that window, so threaded compression genuinely overlaps.

open as a page

Why must a ctypes foreign function have argtypes and restype set?

level: middleimportance: must knowfreq 38%
basics
~20 s

Without them ctypes guesses: it decodes the result as a C int, truncating any 64-bit pointer, and passes arguments by default rules that ignore the real signature. Declaring argtypes and restype makes ctypes convert and type-check every call.

open as a page

What does ctypes.CDLL do, and how do you call a function from that library?

level: juniorimportance: should knowfreq 30%
basics
~20 s

ctypes.CDLL loads a shared library into the running process through the platform's dynamic loader. Attribute access on the object it returns resolves an exported C symbol by name and hands back a callable you invoke like a normal function.

open as a page

Why read a binary file with readinto and a preallocated bytearray instead of read?

level: middleimportance: should knowfreq 32%
basics
~20 s

read(n) allocates a fresh bytes object on every call. readinto(buf) writes into a writable buffer you already own and returns how many bytes it actually wrote, so a streaming loop reuses one allocation instead of producing garbage per chunk.

open as a page