skip to content

How do you use timeit to measure the cost of a try/except block?

level: juniorimportance: should knowfreq 45%

answer

  1. Do not hand-roll a stopwatch
  2. One stdlib module owns micro-benchmarks
  3. Setup runs once and is not timed
  4. Repeat several times, take the minimum
  5. Beware constant-folded statements

basics

~20 s

Write both variants as statement strings, put shared state in a setup string that timeit runs once untimed, call timeit.repeat with a large number of loops, and compare the minimum of the repeats rather than the average.

solid answer

~40 s

Use `timeit`, not a hand-rolled stopwatch. `timeit.timeit(stmt, setup, number=N)` runs `setup` once without timing it, then executes `stmt` N times inside a tight loop and returns the total seconds; `timeit.repeat(..., repeat=5)` gives you five such totals and you report the **minimum**, because noise can only make a run slower. Pass the same `number` to both variants so the totals are comparable, and make N large enough — a million loops for something measured in tens of nanoseconds. `timeit` disables the cyclic garbage collector for the duration and uses `time.perf_counter` as its clock, both of which remove noise you would otherwise fight by hand. The classic trap is measuring nothing: `1 + 1` is folded at compile time, so time a real dictionary lookup or conversion instead.

code

python · 23 lines
python
import timeit

SETUP = "budgets = {'svc-1': 12.5}"

HIT = """
try:
    budgets['svc-1']
except KeyError:
    pass
"""

MISS = """
try:
    budgets['svc-2']
except KeyError:
    pass
"""

n = 200_000
hit = min(timeit.repeat(HIT, SETUP, repeat=5, number=n))
miss = min(timeit.repeat(MISS, SETUP, repeat=5, number=n))
print(f"no raise: {hit / n * 1e9:.1f} ns/op")
print(f"raises:   {miss / n * 1e9:.1f} ns/op")

go deeper

for a junior

Know the shape of a measurement: statement string, setup string that is not timed, a large loop count, and the minimum of several repeats. Being able to produce a defensible number beats any opinion about which form is faster.

for a middle

Explain why the minimum is the right statistic and why timeit disables the cyclic collector during the loop. Be ready to spot the constant-folding trap that makes two snippets look identically fast.

for a senior

Demonstrate experimental hygiene: same loop count on both sides, per-operation nanoseconds rather than raw totals, repeats reversed to rule out warm-up, and a stated threshold below which you call a difference noise.

for a principal

Set the bar for what counts as evidence in a performance argument on your team, and say where micro-benchmarks stop: they choose between two known lines, and never substitute for finding out where the program's time actually goes.

## Why `timeit` rather than a stopwatch The instinct is to bracket a call with `time.time()` and subtract. For anything at the scale of exception handling that is useless: a single guarded lookup takes tens of nanoseconds, which is at or below the resolution of the clock, and one sample is dominated by whatever else the machine was doing. `timeit` exists to fix exactly that. It runs the statement many times in a tight loop, times the whole loop with a high-resolution clock, and lets you repeat the whole experiment. ## The three entry points **`timeit.timeit(stmt, setup, number=N)`** compiles `setup` and runs it once — untimed — then runs `stmt` N times and returns the total elapsed seconds as a float. Divide by N yourself if you want per-operation time. **`timeit.repeat(stmt, setup, repeat=R, number=N)`** does that whole thing R times and returns a list of R totals. This is what you should normally use. **`timeit.Timer`** is the object behind both, and it carries `autorange`, which picks a loop count for you by growing N until the run takes at least 0.2 seconds. There is also a command-line form, `python -m timeit`, where `-s` supplies setup lines and the remaining arguments are joined with newlines into the statement — handy for a multi-line `try/except`. ## Minimum, not mean Report `min(timeit.repeat(...))`. Timing noise is one-sided: another process, a page fault or a scheduler preemption can only make a run *slower*, never faster than the code actually is. The mean therefore drifts with whatever else the machine did, while the minimum is the run least disturbed. Report the mean only when you are deliberately characterizing variance rather than comparing two implementations. ## Setting the experiment up correctly Everything the statement needs but should not be timed goes in `setup`: imports, building the dictionary, creating the object. `setup` runs once per repeat and is never included in the returned time. If you want to reuse objects that already exist in your module rather than rebuilding them in a string, pass `globals=globals()` and the statement is executed against your namespace. Use the **same `number`** on both sides of a comparison, or the totals are not comparable. Keep the machine otherwise quiet, and repeat the comparison in the reverse order once — if the second-measured variant is always the faster one, you are measuring warm-up, not code. ## The trap that ruins exception benchmarks The single most common mistake is timing something the compiler already removed. `1 + 1` is constant-folded, so both a guarded and an unguarded version of it measure an empty loop and you "prove" the guard is free for the wrong reason. Put a real operation inside: a dictionary subscript, a `float()` conversion, an attribute access. The same applies in the other direction — if you want to measure the cost of a *raise*, make sure the raise actually happens on every loop, and remember that you are then also measuring the exception's construction and the traceback allocated for each frame it unwinds through. ## Reading the numbers honestly A `timeit` result is a total for N loops on one machine, one build and one moment. Two rules keep it meaningful. First, always divide by N and state the per-operation figure in nanoseconds so the number can be compared with a mental model — a dictionary hit is around ten nanoseconds, a raise caught in the same frame costs a few times that, and one thrown from several frames down runs into the hundreds of nanoseconds or beyond. Second, treat differences smaller than the spread of your repeats as noise; if `min` and the next-smallest repeat differ by more than the gap you are claiming, you have not measured anything. ## What `timeit` does for you quietly Two behaviours are worth knowing because interviewers probe them. `timeit` disables the cyclic garbage collector for the duration of the timing loop and restores it afterwards, so a collection triggered by unrelated allocation does not land inside your measurement. And its clock is `timeit.default_timer`, which is `time.perf_counter` — a monotonic, highest-resolution clock rather than wall-clock time. If you specifically want to measure with collection enabled, you can re-enable it in the setup string. ## Where `timeit` stops `timeit` answers "which of these two small snippets is faster". It does not tell you where the time in a real program goes; that is a profiler's job and a different question entirely. Reach for `timeit` when you already know which two lines you are choosing between.

  • Why do you report the minimum of the repeats instead of the mean?
    Because timing noise is one-sided. Interference from other processes, page faults or scheduling can only add time, never remove it, so the fastest repeat is the one least disturbed and the best estimate of what the code actually costs. The mean mostly measures how busy the machine was; use it only when you are deliberately studying variance.
  • What does timeit do with the cyclic garbage collector while it runs?
    It disables it for the duration of the timing loop and restores the previous state afterwards, so a collection pass triggered by unrelated allocation cannot land inside a measurement. If your snippet's real-world cost depends on collection, re-enable it in the setup string and say so when you report the number.
  • Your guarded and unguarded snippets time identically at 8 ns each. What do you check first?
    Whether you measured anything at all. Expressions like `1 + 1` are constant-folded at compile time, so both variants become an empty loop. Replace the body with a real operation — a dictionary subscript, a conversion, an attribute access — and confirm the per-operation figure moves into a plausible range before drawing any conclusion.

saying these in an interview costs you the question

  • Wraps one call in time.time() and subtracts
  • Averages the repeats instead of taking the minimum
  • Thinks the setup string is included in the timing
  • Times a constant-folded expression and proves nothing
  • Uses different number values for the two variants
  • Reports a total without dividing by the loop count

context