Why does timeit.timeit turn off garbage collection during the timed loop?
answer
- One subsystem is switched off around the loop
- It is restored in a finally block
- Allocation-heavy code looks better than it is
- Reference counting is untouched by it
- Re-enable it from inside setup
basics
~20 stimeit disables the cyclic garbage collector around the timed loop and restores it afterwards, so a collection triggered by unrelated earlier allocations cannot land inside your measurement. It buys comparability at the cost of flattering allocation-heavy code.
solid answer
~40 sBefore running the timed loop, `timeit` records whether the collector was enabled, calls `gc.disable()`, and re-enables it in a `finally` block. The motivation is independence: whether a generational collection happens to fire during your ten milliseconds should not decide which of two implementations looks faster. The important caveat is that this makes the measurement unrepresentative for code that allocates heavily or creates reference cycles, because in production that code pays for the collections it provokes. Note also what is *not* disabled — reference counting still frees objects immediately, so `gc.disable()` only stops cycle detection, not deallocation. If you want the collector in the picture, re-enable it inside setup: `timeit.timeit(stmt, setup='import gc; gc.enable()')`, since setup runs before the clock starts but inside the same call.
code
python · 7 linesimport timeit
stmt = "d = {}; d['self'] = d"
off = timeit.timeit(stmt, number=200000)
on = timeit.timeit(stmt, setup="import gc; gc.enable()", number=200000)
print(f"collector disabled: {off:.3f}s")
print(f"collector enabled: {on:.3f}s")go deeper
It is enough to remember that timeit deliberately controls the environment around the code it times, and that a benchmark number therefore describes conditions you chose rather than the ones your program runs in.
Be able to state that the cyclic collector is disabled for the timed loop and restored afterwards, that reference counting still frees objects, and how to put collection back into the measurement via the setup string.
Use this to reason about which comparisons a micro-benchmark can settle. Show that you notice when two implementations differ mainly in allocation pressure, and that you validate such a comparison with the collector live or at the service level.
Frame the general point for the team: every control a benchmark applies is a difference from production. Decide which controls your performance claims are allowed to assume, and require the conditions to be stated with the number.
## What timeit actually does The timed loop is wrapped, in effect, like this: ```python import gc was_enabled = gc.isenabled() gc.disable() try: result = inner(iterations, timer) finally: if was_enabled: gc.enable() ``` So the collector is off for the duration of the measurement and its previous state is restored afterwards — including the case where it was already disabled by the caller. This is documented behaviour, not an implementation accident. ## Why it is defensible CPython's cyclic collector is generational and threshold-driven: it runs when the difference between allocations and deallocations crosses a threshold, sweeping a generation of tracked container objects. Whether that threshold is crossed inside your timed window depends on what the process did *before* the benchmark — how much data setup built, what the previous repetition left behind, what the import machinery allocated. That is a source of variance entirely unrelated to the statement under test, and it can easily attribute one implementation's collection pause to its rival. Disabling the collector removes that coupling, which is the same instinct behind the best-of-several-repetitions rule: strip out disturbance so two measurements are comparable. ## Why it can mislead The cost is realism. If the statement allocates aggressively — building containers, creating short-lived objects with references to each other, constructing exception tracebacks — then in production it *causes* collections, and the time those collections take is genuinely part of its cost. A micro-benchmark run with the collector off assigns that cost to nobody. Two implementations that differ mainly in allocation pressure can therefore look equivalent under timeit and diverge noticeably in a real service. There is a second-order effect too. With the collector off, cyclic garbage created inside the loop is never reclaimed for the whole run. A loop that builds a self-referencing dictionary a million times leaves a million of them alive, and the growing heap changes allocator behaviour and cache locality. Long timeit runs over cycle-creating code can drift slower for that reason alone. ## What is not disabled A frequent misconception is that `gc.disable()` stops memory from being freed. It does not. CPython frees an object the moment its reference count drops to zero, and that path is untouched by the collector's state. What stops is the periodic scan that finds groups of objects which reference each other and are unreachable from anywhere else. So a loop creating and dropping ordinary lists still returns memory promptly under timeit; only genuine cycles accumulate. ## Turning it back on The documented way to include collection in the measurement is to re-enable it in setup, which runs inside the same call but before the clock starts: ```python timeit.timeit(stmt, setup='import gc; gc.enable()', number=100000) ``` The explicit `import gc` makes this work regardless of which namespace the generated function runs in, including when you pass your own `globals`. Do this when the question you are asking is 'what will this cost in a running process', and leave the default when the question is 'which of these two expressions is intrinsically cheaper'. ## The interview point behind the trivia The fact itself is small. What it stands for is not: a micro-benchmark is a controlled experiment, and every control it applies is a way in which it stops resembling production. Garbage collection is the clearest example because it is documented and easy to toggle, but the same argument covers a pre-warmed cache in setup, a fixture that fits in L2, a single-threaded loop where the real workload has contention, and a process with no other work competing for the allocator. A candidate who knows the collector is off and can say why that matters is really demonstrating that they read benchmark numbers as conditional on their setup. Related tooling exists for the opposite direction: `gc.freeze()` moves everything currently tracked into a permanent generation so later collections do not rescan it, which is a production tuning knob rather than a benchmarking one. And `gc.collect()` called explicitly between repetitions is a reasonable way to start each measurement from a known-clean heap. ## Two practical consequences First, the collector's state is process-wide, so while a timed loop runs, every other thread in the process also runs without cycle collection. That rarely matters in a benchmark script that does nothing else, but it is worth knowing before you embed a timing call inside a live application. Second, if you want each repetition to start from a comparable heap, call `gc.collect()` yourself between repetitions — the collector being disabled inside the measurement says nothing about how much uncollected garbage the previous repetition left behind. Running `timeit.repeat` and collecting explicitly between the calls gives measurements that are both undisturbed and starting from the same state, which is usually what people mean when they say a benchmark is clean.
- Does disabling the collector mean nothing is freed during a timeit run?No. CPython frees an object as soon as its reference count reaches zero, and that is independent of the collector's state. `gc.disable()` only stops the periodic generational scan that reclaims groups of objects referencing each other. So ordinary short-lived lists and strings are still released promptly inside a timed loop; only genuine reference cycles accumulate until the collector is re-enabled.
- When would you deliberately re-enable collection inside a benchmark?When the statement's allocation pressure is part of what you are comparing — building large containers, creating cyclic structures, or code you suspect provokes collections in production. Pass `setup='import gc; gc.enable()'` so the loop runs with the collector live. Expect noisier results, and lean on more repetitions and the spread rather than a single best-of figure.
saying these in an interview costs you the question
- Believes timeit leaves the garbage collector untouched
- Thinks disabling the collector stops all memory from being freed
- Claims timeit leaves collection off after the call returns
- Treats an allocation-heavy result as production-representative
- Calls gc.enable() before timeit and assumes it sticks