Performance and Memory
Where Python spends time and memory: the CPython object model, the profilers that find the real cost, and the idioms and native escape hatches that fix it. It separates measurers from guessers.
part ofPythonoverview, primer and where to startread it →on this pageshowhide
explore
- Memory Model23 questions
- Reference Counting4 questions
- Cycle Collection and the gc Module4 questions
- Weak References4 questions
- Object Overhead and __slots__4 questions
- pymalloc and Arena Behaviour4 questions
- Interning and Object Reuse3 questions
- Profiling and Benchmarking23 questions
- timeit and Micro-Benchmarks4 questions
- cProfile and pstats3 questions
- tracemalloc and Memory Profiling4 questions
- Finding What Holds Objects4 questions
- Sampling and Wall-Clock Time4 questions
- Startup and Import Cost4 questions
- Optimization Idioms17 questions
- Pushing Loops into C4 questions
- Streaming vs Materializing4 questions
- Caching and Memoization5 questions
- Sharing Data Across Processes4 questions
- Execution and Acceleration16 questions
- Bytecode and dis4 questions
- Adaptive Specialization and JIT4 questions
- Leaving Pure Python4 questions
- PyPy and Alternative Runtimes4 questions
- Native Extension Boundary12 questions
- ctypes and Foreign Functions4 questions
- Buffer Protocol and memoryview4 questions
- Releasing the GIL4 questions
questions
91 · 5 sectionsWhy can reference counting alone never free two Python objects that point at each other?
basics
~20 sEach object in the cycle holds a reference to the other, so neither count drops to zero even after every outside name is gone. CPython runs a second, tracing collector — the gc module — to find and free such groups.
In CPython, what does `del x` actually do — does it free the object?
basics
~20 sdel x only unbinds the name x from its namespace and drops one reference to the object. CPython deallocates the object at the moment its reference count reaches zero, so if a list, attribute or another name still holds it, nothing is freed.
What does weakref.ref give you that an ordinary Python reference does not?
basics
~10 sA weakref.ref points at an object without counting toward the references that keep it alive. Call the ref like a function to get the object back, or None once the object has been collected.
How much memory does adding __slots__ to a Python class actually save?
basics
~20 s__slots__ fixes the attribute names up front so instances store values at known offsets and carry no per-instance dictionary. On CPython 3.14 that is roughly a 35-45% cut for a small class - real, but far less than the old folklore.
A Python inventory-sync worker's RSS climbs each batch and never falls: how do you tell a real leak from allocator retention?
basics
~20 sLog sys.getallocatedblocks() once per batch. A live block count that returns to baseline while RSS does not means the objects are gone and pymalloc is holding fragmented arenas; a climbing count is a real leak, localised with tracemalloc.
Why does a Python `import` statement execute code rather than just bind a name?
basics
~20 sThe first import of a module runs its whole body top to bottom, then stores the finished module object so later imports are just a cache lookup. Every top-level statement in it is startup cost.
Why do objects appended to a module-level list never get freed in Python?
basics
~20 sA module-level list stays reachable from sys.modules for the whole process, so every object it holds keeps a live reference and can never be freed. Python reclaims an object only when nothing reachable still refers to it.
What is the difference between wall-clock time and CPU time when timing Python code?
basics
~20 sWall-clock time is elapsed real time, read with time.perf_counter(). CPU time is the time the processor actually spent executing the process, read with time.process_time(). Sleeping, waiting on a socket or blocking on a lock adds wall time but no CPU time.
Why does timeit.timeit take a separate setup argument instead of one block of code?
basics
~20 stimeit runs the setup string once, before the clock starts, and only the stmt string inside the timed loop. Imports and test-data construction are therefore excluded, so you measure the operation itself rather than its fixtures.
In a cProfile report, what is the difference between tottime and cumtime?
basics
~20 stottime is the time spent inside a function's own body, excluding the calls it makes. cumtime adds everything it called. Sort by tottime to find hot code, by cumtime to find the expensive call path.
Why is ''.join(parts) preferred over building a string with += in a loop?
basics
~20 sStrings are immutable, so each += builds a brand-new string and copies everything accumulated so far. str.join walks the sequence once in C, sizes the result, allocates it once, and copies each piece exactly once.
Why don't changes a multiprocessing child process makes to a global variable reach the parent?
basics
~20 sEach multiprocessing worker is a separate OS process with its own address space and its own copy of every module-level name. What the child rebinds or mutates changes only the child's memory; results must be sent back explicitly.
What is the difference between a generator expression and a list comprehension in Python?
basics
~20 sA list comprehension builds the whole list in memory at once. A generator expression builds nothing up front: it produces items one at a time on demand, holds only the current value, can be consumed once, and supports neither len() nor indexing.
Why does functools.lru_cache raise TypeError when a list is passed as an argument?
basics
~20 sThe wrapper builds a key from the call's arguments and looks it up in a dictionary, so every argument must be hashable. A list is not, and the lookup raises TypeError: unhashable type: 'list' before the function body ever runs.
How do you stream a 40 GB log file in Python without loading it into memory?
basics
~20 sIterate the open file object directly with a for loop: it is its own iterator and hands back one buffered line at a time. Never call read() or readlines(), and keep every downstream stage lazy so nothing collects the whole file.
Why is reading a local variable cheaper than reading a global in CPython?
basics
~20 sA local read compiles to a LOAD_FAST-family instruction, an index into the frame's array of local slots. A global read compiles to LOAD_GLOBAL, which hashes the name and searches the module dictionary, then builtins. An array index beats a hash lookup.
Why is an element-wise Python loop far slower than one vectorized array-library call?
basics
~20 sThe loop pays interpreter overhead on every element: bytecode dispatch, a heap-allocated object per value, reference-count updates and runtime type checks. One vectorized call pays that once, then runs a typed machine-code loop over contiguous memory.
How does CPython's specializing adaptive interpreter (PEP 659) speed up hot code?
basics
~20 sCPython watches how each bytecode instruction is actually used and rewrites it in place into a form specialized for the types it keeps seeing, behind a cheap guard check. No machine code is produced; the interpreter simply does less work per instruction.
Why do C extensions block PyPy adoption, and which bindings port cleanly?
basics
~20 sCPython's C API hands extensions raw pointers to objects and makes them maintain reference counts. PyPy has a moving collector and no reference counts, so it emulates that API — costly at every crossing and not always complete.
What does Python's `dis.dis()` print when you pass it a function?
basics
~20 sdis.dis(f) prints the CPython bytecode the compiler produced for f: one line per instruction, with the source line, the byte offset, the opcode name and its resolved argument. It shows what the interpreter executes, not how long it takes.
Why slice a memoryview of a large bytes object instead of the bytes itself?
basics
~20 sSlicing a bytes object allocates a new bytes and copies the requested range. A memoryview exposes the same memory through the buffer protocol, so slicing it returns another view over the original block in constant time, copying nothing.
What happens to the GIL while zlib.compress runs on a large buffer?
basics
~20 sThe C code inside zlib drops the global interpreter lock before it starts deflating and takes it back when the call finishes. Other Python threads run real bytecode during that window, so threaded compression genuinely overlaps.
Why must a ctypes foreign function have argtypes and restype set?
basics
~20 sWithout them ctypes guesses: it decodes the result as a C int, truncating any 64-bit pointer, and passes arguments by default rules that ignore the real signature. Declaring argtypes and restype makes ctypes convert and type-check every call.
What does ctypes.CDLL do, and how do you call a function from that library?
basics
~20 sctypes.CDLL loads a shared library into the running process through the platform's dynamic loader. Attribute access on the object it returns resolves an exported C symbol by name and hands back a callable you invoke like a normal function.
Why read a binary file with readinto and a preallocated bytearray instead of read?
basics
~20 sread(n) allocates a fresh bytes object on every call. readinto(buf) writes into a writable buffer you already own and returns how many bytes it actually wrote, so a streaming loop reuses one allocation instead of producing garbage per chunk.