skip to content

Why is reading a local variable cheaper than reading a global in CPython?

level: middleimportance: must knowfreq 55%

answer

  1. Three different amounts of work, not three flavours
  2. Scope is fixed at compile time
  3. An array slot versus a hashed dictionary probe
  4. Globals fall back to the builtins dictionary
  5. Attribute access runs the whole descriptor protocol

basics

~20 s

A local read compiles to a LOAD_FAST-family instruction, an index into the frame's array of local slots. A global read compiles to LOAD_GLOBAL, which hashes the name and searches the module dictionary, then builtins. An array index beats a hash lookup.

solid answer

~50 s

Scope is decided at compile time. The compiler numbers every name assigned in a function body, records them in `co_varnames`, and emits `LOAD_FAST` (on 3.14 usually `LOAD_FAST_BORROW`) to fetch slot *n* from the frame's array — a pointer offset. A name it cannot prove local becomes `LOAD_GLOBAL`, which hashes the string and probes the module's global dictionary, then falls back to the builtins dictionary; that is why `len` costs two dictionary probes. `LOAD_ATTR` is more expensive still, because `obj.attr` runs the whole attribute protocol: the type's MRO is searched for a data descriptor before the instance `__dict__` is even consulted. Hence the classic hoist — binding `append = out.append` before a hot loop turns a per-iteration attribute lookup into one local read. The win is real but small, and narrower than it was before 3.11 cached global lookups, so hoist only where you have measured.

code

python · 12 lines
python
import dis

GAIN = 1.05

def with_global(rows):
    return [row * GAIN for row in rows]

def with_local(rows, gain=GAIN):
    return [row * gain for row in rows]

dis.dis(with_global)
dis.dis(with_local)

go deeper

for a junior

Be able to say that Python looks up locals, globals and attributes by different mechanisms, and that locals are the fastest. Knowing the ordering and that a disassembly can show it is enough at this stage.

for a middle

Explain the mechanics: numbered frame slots for locals versus a hashed dictionary probe plus a builtins fallback for globals, and the full descriptor walk for attributes. Expect to be asked why the compiler can decide this ahead of time.

for a senior

Show judgement about when hoisting is worth the readability cost, and know that global lookups have been cached per instruction since 3.11, so old benchmark folklore overstates the gap. Bring a measurement before you rewrite a loop.

for a principal

Own the guidance for the team: micro-optimisations like hoisting are local, measured exceptions, not conventions. The larger call is usually whether the hot path should stay in pure Python at all.

The three name-loading instructions you meet first in a disassembly — `LOAD_FAST`, `LOAD_GLOBAL`, `LOAD_ATTR` — are not three flavours of the same operation. They are three genuinely different amounts of work, and the difference is decided before your program runs. ## Scope is a compile-time decision When CPython compiles a function it builds a symbol table for the body. Any name **assigned** anywhere in that body — by `=`, `for`, `with ... as`, `import`, `def`, an augmented assignment — is a local for the *entire* body, including lines above the assignment. Locals are numbered and their names stored in `co_varnames`; the frame created for each call carries a flat array of that many slots. Two consequences follow. First, `LOAD_FAST n` is a pointer offset into that array plus a reference-count bump: no hashing, no dictionary, no fallback. Second, if the compiler decided a name is local and the slot is still empty when you read it, you get `UnboundLocalError`, not `NameError` — the slot exists, it just holds nothing. Names from an enclosing function are a third case: they compile to `LOAD_DEREF`, which reads a cell object and then its contents, one indirection more than a local and still far cheaper than a global. ## What a global read actually costs A name the compiler cannot prove local becomes `LOAD_GLOBAL`, whose oparg indexes `co_names`. At run time the interpreter hashes the name string and probes the module's global dictionary; on a miss it probes the builtins dictionary. So every bare `len`, `range`, `print` or `int` in a loop body is a dictionary lookup that misses once and hits once. The lookup must happen every iteration because the semantics demand it: any thread, any imported module, any line of code could rebind the module global between iterations, so the interpreter is not allowed to cache the value naively. `LOAD_ATTR` is heavier again. Evaluating `obj.attr` invokes `type(obj).__getattribute__`, which searches the type's MRO for the name; if it finds a **data descriptor** (a `property`, for example) that wins outright, otherwise the instance `__dict__` is consulted, and only then the class attribute or non-data descriptor — with `__getattr__` as a last resort on failure. A method access also builds a bound-method object. In 3.12 `LOAD_METHOD` was merged into `LOAD_ATTR`, so a method call now shows up as a single `LOAD_ATTR` with a flag bit set, followed by `CALL`. ## Hoisting, and its honest value The classic idiom follows directly: ```python def import_rows(rows, out): append = out.append # one LOAD_ATTR, once for row in rows: append(row) # one local read per iteration ``` Disassemble both spellings and the loop body shrinks from `LOAD_FAST` + `LOAD_ATTR` + `CALL` to `LOAD_FAST` + `CALL`. On the per-row loop of a payroll CSV import running against a 92nd-percentile latency budget, that can be the few percent you need. But be honest about the size of it: since 3.11 the interpreter caches global lookups per instruction, so the global-versus-local gap narrowed considerably, and the remaining win is mostly in attribute chains inside genuinely hot loops. Hoisting costs readability every time — a reader must now track an alias — so it belongs behind a measurement, never in a style guide. The same mechanism explains a favourite interview nugget: identical code runs faster inside a function than at module level, because at module scope *every* name is a global and there are no fast local slots at all. Wrapping a script's body in `def main():` is sometimes a free speedup. Writes mirror reads. Binding a local is `STORE_FAST`, a store into the same slot array, while binding a module global is `STORE_NAME` or `STORE_GLOBAL` into a dictionary. This is also why mutating `globals()` from inside a function cannot create a local: the compiler already decided which names are slots, and nothing done to a dictionary at run time can add one. It is the same fact seen from the other side — the shape of the frame is fixed when the function is compiled, not when it is called. ## Verifying rather than believing All of this is checkable in one line: disassemble the two versions and read the loop body. On 3.14 you will mostly see `LOAD_FAST_BORROW` rather than `LOAD_FAST` — a 3.14 addition that pushes a *borrowed* reference and skips the reference-count bump where the compiler can prove the value stays alive — and `LOAD_FAST_CHECK` where a local might legitimately be unbound. They are all the same family: an index into the frame's slots, and the cheapest name read Python has.

  • Where does a name from an enclosing function fit in that cost order?
    Between a local and a global. It compiles to `LOAD_DEREF`, which reads a cell object and then the value inside it — one indirection more than a local slot, but still no hashing and no dictionary. The compiler records those names in `co_freevars` on the inner code object and `co_cellvars` on the outer one.
  • Why does adding an assignment to a name break a function that previously read the module global?
    Because scope is decided for the whole body at compile time. Once a name is assigned anywhere in the function, every read of it compiles to a local slot access, including reads on earlier lines. Reading the still-empty slot raises `UnboundLocalError`. Declaring `global name` or renaming the local is the fix.
  • How much does hoisting a lookup actually buy today?
    Less than the folklore claims. Since 3.11 global lookups are cached per instruction, so the biggest remaining win is removing a per-iteration attribute lookup from a genuinely hot loop. Treat it as a measured micro-optimisation on a proven hot path, not a house style: it costs readability on every line it touches.

A local is a numbered pigeonhole you reach into directly; a global is asking the front desk to look the name up in a register, and then asking a second desk when the first has never heard of it.

saying these in an interview costs you the question

  • Says Python resolves every name the same way
  • Thinks LOAD_FAST searches a dictionary of local names
  • Blames the GIL for slow global lookups
  • Believes attribute access costs the same as a local read
  • Hoists lookups everywhere as a style rule, unmeasured
  • Claims the interpreter decides scope at run time

context