skip to content

Execution and Acceleration

What CPython does with your source: compiles it to bytecode, dispatches one instruction at a time, specializes the hot parts, plus the faster runtimes underneath. Interviewers ask why Python is slow.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

16

Why is reading a local variable cheaper than reading a global in CPython?

level: middleimportance: must knowfreq 55%

answer

  1. Three different amounts of work, not three flavours
  2. Scope is fixed at compile time
  3. An array slot versus a hashed dictionary probe
  4. Globals fall back to the builtins dictionary
  5. Attribute access runs the whole descriptor protocol

basics

~20 s

A local read compiles to a LOAD_FAST-family instruction, an index into the frame's array of local slots. A global read compiles to LOAD_GLOBAL, which hashes the name and searches the module dictionary, then builtins. An array index beats a hash lookup.

solid answer

~50 s

Scope is decided at compile time. The compiler numbers every name assigned in a function body, records them in `co_varnames`, and emits `LOAD_FAST` (on 3.14 usually `LOAD_FAST_BORROW`) to fetch slot *n* from the frame's array — a pointer offset. A name it cannot prove local becomes `LOAD_GLOBAL`, which hashes the string and probes the module's global dictionary, then falls back to the builtins dictionary; that is why `len` costs two dictionary probes. `LOAD_ATTR` is more expensive still, because `obj.attr` runs the whole attribute protocol: the type's MRO is searched for a data descriptor before the instance `__dict__` is even consulted. Hence the classic hoist — binding `append = out.append` before a hot loop turns a per-iteration attribute lookup into one local read. The win is real but small, and narrower than it was before 3.11 cached global lookups, so hoist only where you have measured.

code

python · 12 lines
python
import dis

GAIN = 1.05

def with_global(rows):
    return [row * GAIN for row in rows]

def with_local(rows, gain=GAIN):
    return [row * gain for row in rows]

dis.dis(with_global)
dis.dis(with_local)

go deeper

for a junior

Be able to say that Python looks up locals, globals and attributes by different mechanisms, and that locals are the fastest. Knowing the ordering and that a disassembly can show it is enough at this stage.

for a middle

Explain the mechanics: numbered frame slots for locals versus a hashed dictionary probe plus a builtins fallback for globals, and the full descriptor walk for attributes. Expect to be asked why the compiler can decide this ahead of time.

for a senior

Show judgement about when hoisting is worth the readability cost, and know that global lookups have been cached per instruction since 3.11, so old benchmark folklore overstates the gap. Bring a measurement before you rewrite a loop.

for a principal

Own the guidance for the team: micro-optimisations like hoisting are local, measured exceptions, not conventions. The larger call is usually whether the hot path should stay in pure Python at all.

The three name-loading instructions you meet first in a disassembly — `LOAD_FAST`, `LOAD_GLOBAL`, `LOAD_ATTR` — are not three flavours of the same operation. They are three genuinely different amounts of work, and the difference is decided before your program runs. ## Scope is a compile-time decision When CPython compiles a function it builds a symbol table for the body. Any name **assigned** anywhere in that body — by `=`, `for`, `with ... as`, `import`, `def`, an augmented assignment — is a local for the *entire* body, including lines above the assignment. Locals are numbered and their names stored in `co_varnames`; the frame created for each call carries a flat array of that many slots. Two consequences follow. First, `LOAD_FAST n` is a pointer offset into that array plus a reference-count bump: no hashing, no dictionary, no fallback. Second, if the compiler decided a name is local and the slot is still empty when you read it, you get `UnboundLocalError`, not `NameError` — the slot exists, it just holds nothing. Names from an enclosing function are a third case: they compile to `LOAD_DEREF`, which reads a cell object and then its contents, one indirection more than a local and still far cheaper than a global. ## What a global read actually costs A name the compiler cannot prove local becomes `LOAD_GLOBAL`, whose oparg indexes `co_names`. At run time the interpreter hashes the name string and probes the module's global dictionary; on a miss it probes the builtins dictionary. So every bare `len`, `range`, `print` or `int` in a loop body is a dictionary lookup that misses once and hits once. The lookup must happen every iteration because the semantics demand it: any thread, any imported module, any line of code could rebind the module global between iterations, so the interpreter is not allowed to cache the value naively. `LOAD_ATTR` is heavier again. Evaluating `obj.attr` invokes `type(obj).__getattribute__`, which searches the type's MRO for the name; if it finds a **data descriptor** (a `property`, for example) that wins outright, otherwise the instance `__dict__` is consulted, and only then the class attribute or non-data descriptor — with `__getattr__` as a last resort on failure. A method access also builds a bound-method object. In 3.12 `LOAD_METHOD` was merged into `LOAD_ATTR`, so a method call now shows up as a single `LOAD_ATTR` with a flag bit set, followed by `CALL`. ## Hoisting, and its honest value The classic idiom follows directly: ```python def import_rows(rows, out): append = out.append # one LOAD_ATTR, once for row in rows: append(row) # one local read per iteration ``` Disassemble both spellings and the loop body shrinks from `LOAD_FAST` + `LOAD_ATTR` + `CALL` to `LOAD_FAST` + `CALL`. On the per-row loop of a payroll CSV import running against a 92nd-percentile latency budget, that can be the few percent you need. But be honest about the size of it: since 3.11 the interpreter caches global lookups per instruction, so the global-versus-local gap narrowed considerably, and the remaining win is mostly in attribute chains inside genuinely hot loops. Hoisting costs readability every time — a reader must now track an alias — so it belongs behind a measurement, never in a style guide. The same mechanism explains a favourite interview nugget: identical code runs faster inside a function than at module level, because at module scope *every* name is a global and there are no fast local slots at all. Wrapping a script's body in `def main():` is sometimes a free speedup. Writes mirror reads. Binding a local is `STORE_FAST`, a store into the same slot array, while binding a module global is `STORE_NAME` or `STORE_GLOBAL` into a dictionary. This is also why mutating `globals()` from inside a function cannot create a local: the compiler already decided which names are slots, and nothing done to a dictionary at run time can add one. It is the same fact seen from the other side — the shape of the frame is fixed when the function is compiled, not when it is called. ## Verifying rather than believing All of this is checkable in one line: disassemble the two versions and read the loop body. On 3.14 you will mostly see `LOAD_FAST_BORROW` rather than `LOAD_FAST` — a 3.14 addition that pushes a *borrowed* reference and skips the reference-count bump where the compiler can prove the value stays alive — and `LOAD_FAST_CHECK` where a local might legitimately be unbound. They are all the same family: an index into the frame's slots, and the cheapest name read Python has.

  • Where does a name from an enclosing function fit in that cost order?
    Between a local and a global. It compiles to `LOAD_DEREF`, which reads a cell object and then the value inside it — one indirection more than a local slot, but still no hashing and no dictionary. The compiler records those names in `co_freevars` on the inner code object and `co_cellvars` on the outer one.
  • Why does adding an assignment to a name break a function that previously read the module global?
    Because scope is decided for the whole body at compile time. Once a name is assigned anywhere in the function, every read of it compiles to a local slot access, including reads on earlier lines. Reading the still-empty slot raises `UnboundLocalError`. Declaring `global name` or renaming the local is the fix.
  • How much does hoisting a lookup actually buy today?
    Less than the folklore claims. Since 3.11 global lookups are cached per instruction, so the biggest remaining win is removing a per-iteration attribute lookup from a genuinely hot loop. Treat it as a measured micro-optimisation on a proven hot path, not a house style: it costs readability on every line it touches.

A local is a numbered pigeonhole you reach into directly; a global is asking the front desk to look the name up in a register, and then asking a second desk when the first has never heard of it.

saying these in an interview costs you the question

  • Says Python resolves every name the same way
  • Thinks LOAD_FAST searches a dictionary of local names
  • Blames the GIL for slow global lookups
  • Believes attribute access costs the same as a local read
  • Hoists lookups everywhere as a style rule, unmeasured
  • Claims the interpreter decides scope at run time

context

open as a page

Why is an element-wise Python loop far slower than one vectorized array-library call?

level: middleimportance: must knowfreq 62%

basics

~20 s

The loop pays interpreter overhead on every element: bytecode dispatch, a heap-allocated object per value, reference-count updates and runtime type checks. One vectorized call pays that once, then runs a typed machine-code loop over contiguous memory.

open as a page

How does CPython's specializing adaptive interpreter (PEP 659) speed up hot code?

level: middleimportance: must knowfreq 42%

basics

~20 s

CPython watches how each bytecode instruction is actually used and rewrites it in place into a form specialized for the types it keeps seeing, behind a cheap guard check. No machine code is produced; the interpreter simply does less work per instruction.

open as a page

Why do C extensions block PyPy adoption, and which bindings port cleanly?

level: seniorimportance: must knowfreq 34%

basics

~20 s

CPython's C API hands extensions raw pointers to objects and makes them maintain reference counts. PyPy has a moving collector and no reference counts, so it emulates that API — costly at every crossing and not always complete.

open as a page

What does Python's `dis.dis()` print when you pass it a function?

level: juniorimportance: should knowfreq 35%

basics

~20 s

dis.dis(f) prints the CPython bytecode the compiler produced for f: one line per instruction, with the source line, the byte offset, the opcode name and its resolved argument. It shows what the interpreter executes, not how long it takes.

open as a page

Why is PyPy often much faster than CPython on long-running pure-Python code?

level: middleimportance: should knowfreq 32%

basics

~20 s

PyPy has a tracing just-in-time compiler: once a loop runs hot it records the operations actually executed, specializes them on the types it saw, and compiles that to machine code. The gain arrives only after warm-up.

open as a page

Why does hoisting `self.rows.append` into a local before a hot loop help, when hoisting `out.append` barely does?

level: middleimportance: should knowfreq 26%

basics

~10 s

self.rows.append is two attribute lookups per iteration; out.append is one. Hoisting deletes lookups from the loop body, and on CPython 3.14 a single lookup is already cheap enough that removing it buys nothing measurable.

open as a page

When would you bind an existing shared library rather than write your own Python extension?

level: middleimportance: should knowfreq 34%

basics

~20 s

Bind when a mature, tuned implementation already exists, so the only code you own is a thin wrapper. Write your own when the hot code is your domain logic, or when a large native dependency buys very little.

open as a page

What does CPython 3.11's zero-cost exception handling change about the price of try/except?

level: middleimportance: should knowfreq 30%

basics

~20 s

Since 3.11, entering a try block runs no setup instruction at all. Handler locations live in a table compiled beside the bytecode and consulted only when something is actually raised, so the whole cost moved onto the raising path.

open as a page

How can a stale `.pyc` in `__pycache__` make a patched module keep running old code?

level: seniorimportance: should knowfreq 26%

basics

~20 s

CPython caches compiled bytecode beside the source and by default recompiles only when the source's recorded modification time or size has changed. A deploy that rewrites a file without changing either silently keeps running the old bytecode.

open as a page

What determines whether a compiled extension actually speeds up a metrics scraper's hot loop?

level: seniorimportance: should knowfreq 44%

basics

~20 s

How much of total runtime the loop owns, how often you cross the boundary, and what the data costs to hand over. A loop worth 70% of runtime caps the overall gain near 3.3x, and per-element calls or copies erase even that.

open as a page

Why does CPython 3.14's free-threaded build run single-threaded work slower than the default build?

level: seniorimportance: should knowfreq 26%

basics

~20 s

Without the GIL, every reference count update and every container access has to be safe on its own, so the interpreter pays for synchronisation the default build got for free. PEP 779 puts the single-threaded overhead in 3.14 at roughly five to ten percent.

open as a page

How would you decide if PyPy suits a webhook receiver peaking at 1,200 requests per minute?

level: principalimportance: should knowfreq 20%

basics

~20 s

Profile first: an alternative runtime pays off only when long-lived processes spend most CPU in pure Python. At twenty requests a second, confirm there is a CPU problem at all, then weigh warm-up, memory and the dependency audit against cheaper fixes.

open as a page

Which acceleration strategy do you pick when pure Python is too slow, and what does it cost the team?

level: principalimportance: should knowfreq 38%

basics

~20 s

Climb a ladder of rising cost and stop at the first rung that meets a stated target: better algorithm, compiled built-ins, bulk array operations, an ahead-of-time compiler, a compiled extension, a different runtime. Each rung trades speed for debuggability, portability and maintainers.

open as a page

What does platform.python_implementation() tell you, and why check it?

level: juniorimportance: nice to knowfreq 18%

basics

~10 s

platform.python_implementation() returns a string naming the interpreter running your code: 'CPython', 'PyPy', 'Jython' or 'IronPython'. You check it to guard behaviour that one implementation happens to guarantee and the language itself does not.

open as a page

Does CPython 3.14 ship a JIT compiler, and is it enabled by default?

level: juniorimportance: nice to knowfreq 20%

basics

~20 s

CPython carries an experimental machine-code JIT, first added in 3.13, but it is opt-in twice over: the interpreter must be built with it, and the process must be started with PYTHON_JIT=1. Most installed builds have it off or absent.

open as a page