skip to content

Why does hoisting `self.rows.append` into a local before a hot loop help, when hoisting `out.append` barely does?

level: middleimportance: should knowfreq 26%

answer

  1. Count the work each iteration repeats
  2. Two dots cost twice one dot
  3. The loop body's LOAD_ATTR count
  4. One indexed local read replaces the chain
  5. Specialization already made one dot cheap

basics

~10 s

self.rows.append is two attribute lookups per iteration; out.append is one. Hoisting deletes lookups from the loop body, and on CPython 3.14 a single lookup is already cheap enough that removing it buys nothing measurable.

solid answer

~40 s

Each dot compiles to its own `LOAD_ATTR`, which runs the attribute protocol every iteration: an MRO walk on the type, a data-descriptor check, then the instance `__dict__`. `self.rows.append(line)` therefore performs two of those per row; binding `add = self.rows.append` before the loop runs the whole chain once and leaves an indexed local read, printed as `LOAD_FAST_BORROW` on 3.14. That measures a single-digit-percent win over a large loop. Hoisting a single dot is different: since 3.11 CPython specializes a `LOAD_ATTR` that keeps seeing the same type, and the inline call form (`dis` shows `append + NULL|self`) skips building a bound method entirely, so `push = out.append` is a wash and can be slightly slower. Read the loop body in `dis.dis()`, count the `LOAD_ATTR`s you actually remove, then confirm with `timeit`.

code

python · 20 lines
python
import dis


class PayrollImport:
    def __init__(self):
        self.rows = []

    def load(self, lines):
        for line in lines:
            self.rows.append(line)

    def load_hoisted(self, lines):
        add = self.rows.append
        for line in lines:
            add(line)


dis.dis(PayrollImport.load)
print("=" * 30)
dis.dis(PayrollImport.load_hoisted)

go deeper

for a junior

Be ready to read a disassembly and say what each dot in a loop body costs. Know that a local name is reached by slot index while an attribute is searched for at runtime, every iteration.

for a middle

You are expected to explain the attribute protocol that LOAD_ATTR runs - MRO walk, data-descriptor check, instance __dict__ - and then show, in dis output, exactly which instructions the hoist removes from the loop body.

for a senior

Demonstrate that you measure before believing: know that specialization since 3.11 killed the single-dot version of this advice, keep the rewrite to loops profiling has flagged, and name the staleness risk in review.

for a principal

Own where micro-optimisation stops. Decide when a codebase should spend readability on single-digit percentages versus restructuring the hot path or moving it out of pure Python, and make that the written norm rather than a per-review argument.

## Every dot is a runtime operation A dotted name in Python source is not a compile-time address; it is work the interpreter performs each time control reaches it. In the row loop of a payroll CSV import, `self.rows.append(line)` is three lookups per iteration: read the local `self`, look `rows` up on it, then look `append` up on the result. CPython compiles the second and third of those dots into their own `LOAD_ATTR` instruction, and each one runs the full attribute protocol on every pass. ## What LOAD_ATTR has to do `LOAD_ATTR` executes what `object.__getattribute__` describes: walk `type(obj).__mro__` looking for the name; if the class attribute found there is a data descriptor (it defines `__set__` or `__delete__`) it wins immediately; otherwise the instance `__dict__` is consulted; otherwise the class attribute is used, and if it is a non-data descriptor its `__get__` runs. Plain functions are non-data descriptors, which is exactly where a bound method comes from. None of that is free, and none of it is memoised by the source text. ## What an indexed local read has to do Locals are not stored in a dictionary. The compiler assigns each local name a slot index, recorded in the code object's `co_varnames`, and the instruction carries that index as its argument, so reading one is an array index into the frame's local slots: no hashing, no MRO walk, no descriptor check. On 3.14 `dis.dis()` prints most of these reads as `LOAD_FAST_BORROW`, which is new in 3.14; on earlier releases the same reads print as `LOAD_FAST`. Bytecode is version-specific, so a disassembly you memorised from an older release will not match line for line. ## Reading it off the disassembly Disassemble the two spellings and compare loop bodies rather than whole functions. The inline version's body carries two `LOAD_ATTR`s per iteration. The hoisted version's body carries none: the whole chain runs once before the loop, and each iteration reads one local and calls it. The `FOR_ITER` jump distance shrinks correspondingly, which is a quick visual check that the body really did get shorter. This is the concrete evidence to point at in an interview, instead of asserting that "attribute access is slow". ## The half that candidates get wrong Fewer instructions is not automatically faster, and on 3.14 the classic advice is half dead. Since 3.11 CPython specializes hot instructions, so a `LOAD_ATTR` at a call site that keeps seeing the same type is quickened into a much cheaper cached form. The consequence, which is the part that matters here, is that a single hoisted dot no longer buys anything: rewriting `out.append(x)` as `push = out.append` then `push(x)` measures as a wash on 3.14, and can come out marginally slower. There is a second reason for that. The inline call form compiles to `LOAD_ATTR` with the low bit of its argument set, which `dis.dis()` annotates as `append + NULL|self`; that form pushes the callable and the instance separately and never materialises a bound method object at all. The hoisted local, by contrast, holds a real bound method built by the one plain `LOAD_ATTR` before the loop. Where the hoist still earns its keep is when it removes a *chain*. `self.rows.append` is two attribute lookups per iteration, and lifting both out of a loop over ten thousand payroll rows measured a single-digit-percent improvement on 3.14. The rule that survives is not "hoist lookups" but "count the attribute lookups the loop body performs, and delete the ones that cannot change". ## What the hoisted name actually holds `xs.append is xs.append` is `False`, because each access constructs a fresh bound method object; `xs.append == xs.append` is `True`, because bound methods compare by their `__func__` and `__self__`. Hoisting simply performs that construction once instead of never. Two consequences follow. The binding is a snapshot: if `self.rows` is later rebound to a different list, or the method is patched, the loop keeps using the object captured before it started. And the local holds a strong reference to the owning instance for as long as the frame lives, which matters when the loop is long and the object is large. ## The judgement to show Use `dis.dis()` to explain the difference and `timeit` to decide whether it is worth having. Apply it only inside a loop that profiling has already identified as hot, because in ordinary code it trades a clear expression for an opaque one and buys nothing measurable. And treat the technique as the small end of the scale: deleting attribute lookups from a loop body is worth single-digit percentages, whereas restructuring the work so the loop runs fewer times, or does not run in pure Python at all, is worth multiples.

  • Does `out.append(x)` allocate a new bound method object on every iteration in CPython 3.14?
    No. The compiler emits the method-call form of `LOAD_ATTR` - `dis` annotates it `append + NULL|self` - which pushes the underlying callable and the instance separately, so nothing is materialised. A bound method is only built when you access the attribute without immediately calling it, which is precisely what the hoist does, once, before the loop. That is why hoisting a single dot removes a lookup but adds an object, and nets out to roughly nothing.
  • How would you prove the win rather than assert it?
    Disassemble both spellings with `dis.dis()` and compare the loop bodies, counting the `LOAD_ATTR`s that actually disappear - that explains the mechanism. Then measure with `timeit.repeat()` on realistic input, taking the minimum across repeats to suppress scheduler noise. Do both: the disassembly without the measurement is a plausible story that specialization may have already invalidated, and the measurement without the disassembly gives you no explanation to reason from.
  • What can break when you hoist a method into a local before a long loop?
    The binding is a snapshot taken before the first iteration. If the attribute is rebound during the loop - the object swaps its list, or a test patches the method - the loop keeps calling the captured original, which is a genuinely hard bug to see. The local also holds a strong reference to the owning instance for the life of the frame. Both are acceptable in a tight numeric loop and unacceptable in a long-lived one over mutable state.

Looking up one extension in the company directory is quick enough that copying it onto a sticky note saves nothing; looking up the department, then the person, then the extension, on every single call, is the part worth writing down once.

saying these in an interview costs you the question

  • Claims attribute access is slow without saying what LOAD_ATTR does
  • Asserts hoisting any lookup is always faster, without measuring
  • Thinks the inline call allocates a bound method every iteration
  • Says LOAD_ATTR and an indexed local read cost the same
  • Confuses attribute lookup with global-namespace lookup
  • Applies the trick everywhere instead of one profiled hot loop

context