skip to content

Why is PyPy often much faster than CPython on long-running pure-Python code?

level: middleimportance: should knowfreq 32%

answer

  1. Remove the interpreter, not the algorithm
  2. Hot loops get recorded, then compiled
  3. Assumptions become guards on machine code
  4. The benefit arrives only after warm-up
  5. Nothing gained inside compiled libraries

basics

~20 s

PyPy has a tracing just-in-time compiler: once a loop runs hot it records the operations actually executed, specializes them on the types it saw, and compiles that to machine code. The gain arrives only after warm-up.

solid answer

~50 s

CPython executes bytecode one instruction at a time, and every integer add allocates a boxed object and dispatches through the evaluation loop. PyPy instead watches for hot loops, records a **trace** — the linear sequence of low-level operations that one pass through the loop actually performed — and compiles it to machine code, specialized on the types it observed. Type checks become cheap **guards** on that machine code, temporary objects can be optimized away, and calls get inlined across the trace, which is where the several-fold speedups on pure-Python numeric, string and algorithmic loops come from. The cost is warm-up: before a loop is hot it runs in PyPy's own interpreter, and tracing plus compiling costs CPU and memory. A script that runs for a second never amortizes that. Time spent inside compiled libraries or waiting on I/O is not accelerated either.

code

python · 14 lines
python
import time


def work(n):
    total = 0
    for i in range(n):
        total += (i * i) % 7
    return total


for phase in range(3):
    start = time.perf_counter()
    work(2_000_000)
    print(phase, round(time.perf_counter() - start, 3))

go deeper

for a junior

Be ready to say that PyPy is a different implementation of the same language and that it compiles hot code to machine code while it runs, rather than translating your program ahead of time like a C compiler.

for a middle

Explain the mechanics: hot loops are traced into a linear, type-specialized sequence, assumptions become guards, and the compiled trace removes dispatch and boxing. Then explain why that benefit needs a long-lived process to be worth anything.

for a senior

Show that you benchmark it honestly — steady state measured apart from warm-up, on realistic input — and that you know which parts of a real service the JIT cannot touch: compiled dependencies, I/O waits, and anything that recycles processes before they warm up.

for a principal

Own the framing that a tracing JIT trades interpreter overhead for warm-up cost and memory, and decide whether your fleet's process lifetime, deployment cadence and memory budget make that trade profitable at all.

**Where CPython's time goes.** Running `total += i * i` in CPython is far more work than the arithmetic. The evaluation loop fetches a bytecode instruction, dispatches on it, pops boxed operand objects off a stack, asks each object's type how to multiply, allocates a fresh integer object for the result, adjusts reference counts, and pushes it back. The multiply itself is a single machine instruction; almost everything around it is interpretation overhead. That overhead is the ceiling a faster runtime is trying to remove. **What tracing means.** PyPy is a separate implementation of Python, itself written in a restricted dialect of Python from which its authors generate the runtime and, crucially, a just-in-time compiler. It counts how often loops execute. When a loop crosses a hotness threshold, PyPy *traces* it: it records the linear sequence of low-level operations that one actual pass through the loop performed — through function calls, through the arithmetic, through attribute lookups — as a straight line with no branches. Everything the trace assumed in order to be straight becomes a **guard**: this operand really is a machine-sized integer, this attribute really is at the offset we saw, this branch really went the way it went. The trace is then optimized and compiled to machine code. That linear, type-specialized form is what makes the optimizations possible. Integers seen to be machine-sized get unboxed and kept in registers. Objects that are allocated and immediately consumed inside the trace can be removed entirely by escape analysis. Repeated type checks and dictionary lookups collapse to constants. Calls disappear into the trace by inlining. The result on a long-running pure-Python loop is routinely several times faster than CPython, and occasionally an order of magnitude. **What warm-up costs.** None of that exists when the process starts. Early iterations run in PyPy's own bytecode interpreter, which is generally *slower* per operation than CPython's, and the tracing and compiling itself costs CPU time and memory. So a runtime's benefit is an integral over the life of the process: a batch job or a server that runs for hours amortizes warm-up in the first seconds and then runs on compiled traces; a command-line tool that starts, does one pass and exits pays warm-up and collects nothing. This is why benchmarking PyPy with a stopwatch on a single cold run is meaningless — you must measure steady state separately from the first pass. **When guards fail.** A guard that fails means reality diverged from what the trace assumed — the loop that only ever saw integers just got a float, or a polymorphic call site hit a second type. Execution leaves the compiled trace and falls back, and if the alternative path is itself hot a *bridge* is compiled for it. Occasional guard failures are normal and cheap. A hot loop that is genuinely polymorphic — different types on every iteration — produces many bridges and much less benefit, because there is no stable shape to specialize on. Monomorphic, predictable inner loops are what a tracing JIT rewards. **What it does not speed up.** - **Time inside compiled libraries.** If your hot path is a call into a shared library, the JIT has nothing to compile — that code is already machine code and the interpreter was never the bottleneck. - **I/O waits.** A service blocked on sockets or disk is not interpreter-bound; a faster interpreter moves nothing. - **Short processes.** Startup and import are heavier, and warm-up never pays back. **The other differences that come with it.** PyPy does not use reference counting; it uses a generational, moving garbage collector. Objects are therefore *not* finalized the instant their last reference vanishes, so code that relies on a file closing at the end of a function — rather than using a `with` block or an explicit close — can leak descriptors for much longer. Baseline memory is higher, and compiled traces add to it, which matters if you deploy many worker processes on a memory-bound machine. And the compatibility story is the real gate on adoption: PyPy tracks a CPython language level a few releases behind the current one, so the newest syntax may simply not parse there. **How to talk about it in an interview.** The clean framing is: a tracing JIT converts *interpreter* overhead into *warm-up* cost. If your process is long-lived and interpreter-bound, that is an excellent trade. If it is short-lived, I/O-bound, or spending its time in compiled code, you are paying the cost and buying nothing.

  • Why can a script that runs for under a second be slower on PyPy than on CPython?
    Because everything it does happens during warm-up. Startup and imports are heavier, the loops run in PyPy's own interpreter before any trace exists, and the tracing and compiling themselves consume CPU. The compiled traces are produced just in time to be thrown away at exit. A tracing JIT converts interpreter overhead into a fixed up-front cost, and a short process pays the cost without living long enough to collect the return.
  • What happens when a guard in a compiled trace fails?
    A guard encodes an assumption the trace was specialized on — that an operand is a machine-sized integer, say. When reality diverges, execution leaves the compiled code and falls back; if the alternative path is itself hot, a bridge is compiled for it. Occasional failures are cheap. A hot loop that genuinely sees a different type every iteration has no stable shape to specialize on, produces many bridges, and gets far less benefit.
  • How does PyPy's garbage collector change behaviour you may have relied on?
    PyPy uses a generational, moving collector with no reference counting, so an object is not destroyed the moment its last reference disappears. A file left unclosed at the end of a function stays open until the collector runs, which can exhaust descriptors under load. The fix is the one that was always correct: `with` blocks and explicit closes rather than relying on CPython's immediate finalization.

It is the difference between reading a recipe aloud step by step every time and, after the tenth batch, writing out one flowing set of motions for exactly the ingredients you keep using — fast, until someone hands you a different ingredient and you have to stop and check.

saying these in an interview costs you the question

  • Describes PyPy as ahead-of-time compiled Python
  • Expects a speedup on short scripts and CLI tools
  • Thinks the JIT accelerates time spent inside compiled libraries
  • Assumes an I/O-bound service gets faster on PyPy
  • Believes PyPy implements the same language version as current CPython
  • Times one cold run and calls it a benchmark

context