Why does timeit report the best of several repeats instead of the average?
answer
- The noise only goes one way
- Nothing external makes code run faster
- One knob amortizes the clock, one samples noise
- Best-of is reproducible; the mean drifts
- Report the spread when the code itself varies
basics
~20 sTiming noise on a real machine is one-sided: other processes, interrupts and cache evictions can only make a run slower. The minimum repetition is therefore the least disturbed sample and the most reproducible estimate of the code's cost.
solid answer
~40 s`timeit.repeat(stmt, setup, repeat=5, number=1000000)` returns one total per repetition, and the conventional summary is `min(results) / number`. The reasoning is that the distribution has a hard floor and a long right tail — nothing external makes your code faster, but a scheduler preemption or a cache eviction makes it slower — so the mean is dragged around by contamination while the minimum estimates the underlying cost. `python -m timeit` follows the same rule and prints a best-of figure, warning you when the slowest repetition is several times the best, which signals a machine too noisy to trust. The minimum is not always the right summary: when the code's own runtime genuinely varies with input or state, report the median or the whole spread, because then the variability is the signal rather than the noise.
code
python · 12 linesimport statistics
import timeit
runs = timeit.repeat(
stmt="sorted(data)",
setup="data = list(range(2000, 0, -1))",
repeat=7,
number=200,
)
print(f"min {min(runs) / 200 * 1e6:.1f} us")
print(f"median {statistics.median(runs) / 200 * 1e6:.1f} us")
print(f"max {max(runs) / 200 * 1e6:.1f} us")go deeper
Know that a single timing is not a result: run the snippet several times and compare the best figures. Recall that the command-line form reports a best-of number rather than an average.
Explain why noise is one-sided and what -n and -r each control. Be able to say that one repetition is already a mean over the inner loop, and that the minimum is chosen across repetitions.
Demonstrate judgement about when the minimum lies — data-dependent work, warm-up since the 3.11 specializing interpreter, tail-latency questions — and show that you report the spread and the conditions alongside any number you publish.
Decide what the organisation treats as evidence: which summary statistic goes in a performance claim, on what hardware, and what spread makes a result inadmissible. Without that convention, benchmark numbers become rhetoric in review threads.
## Two knobs that do different jobs `number` and `repeat` are often confused, and the difference explains the whole min-versus-mean argument. `number` is how many times the statement runs inside **one** pair of clock reads. Its job is to amortize the cost of reading the clock and to lift the measured interval well above the timer's resolution. The figure you divide out of one repetition is already an arithmetic mean over those `number` iterations — you cannot avoid averaging at that level. `repeat` is how many times that whole timed loop is run. Its job is to sample the *environment* several times so you can see, and discard, disturbance. On the command line these are `-n` and `-r`; `python -m timeit -n 1000 -r 7 'stmt'` runs the statement a thousand times per measurement and takes seven measurements. Since Python 3.7 the default is `-r 5`, and when `-n` is omitted timeit autoranges it — trying 1, 2, 5, 10 and so on until one timed loop takes at least about 0.2 seconds. ## Why the minimum Benchmark noise is not symmetric. Consider what can perturb a timed loop: the OS scheduler runs another process on your core, an interrupt fires, the CPU drops frequency for thermal reasons, another tenant evicts your data from a shared cache, a page fault takes a trip to the kernel. Every one of these adds time. There is no corresponding mechanism that makes the same instructions execute faster than the machine can execute them. So the distribution of repetition totals looks like a wall on the left with a tail stretching right. Under that model the sample minimum is the best estimator of the wall's position — the run during which the least interference happened — while the mean and standard deviation describe your neighbours' behaviour as much as your code's. That is why `min()` is the documented recommendation and why the command line prints a best-of figure. The practical payoff is reproducibility. A best-of number taken twice on the same idle laptop tends to agree to a few percent; two means taken on a busy laptop can differ by tens of percent, which makes before-and-after comparison useless. ## When the minimum is the wrong summary The argument above assumes your code has one true cost that noise only inflates. That assumption fails more often than people expect: - The statement's own work varies with data — a lookup that sometimes misses, a branch that sometimes takes the slow path, a container that occasionally resizes. Here the spread *is* the result, and reporting only the floor describes the luckiest case. - The timed loop has warm-up effects. Since Python 3.11 the specializing adaptive interpreter rewrites hot bytecode into specialized forms after a code object has been executed enough times, so early repetitions can be measurably slower than later ones. A minimum silently picks the fully warmed state, which may or may not be the state your production code reaches. - You care about tail latency. If the question is 'what does the p99 request pay', the minimum answers the opposite question. When any of those apply, keep the raw list from `repeat` and report the median plus the range, or plot the distribution. `statistics.median` and `statistics.stdev` are one import away, and the honest report is 'median 41.2 microseconds, min 39.8, max 44.1' rather than a single number with no error bar. ## Reading the spread as a health check The spread is diagnostic even when you intend to quote the minimum. If the worst repetition is several times the best, something outside your code is interfering — the command line says so explicitly — and the right response is to fix the machine, not to publish the number. Close the browser, stop the build, pin the process, run more repetitions, and check that the total per repetition is large enough that clock resolution is irrelevant. ## Putting it together A defensible micro-benchmark reports the per-call minimum for comparability, keeps the full list so the spread is visible, and states the conditions. `min(timeit.repeat(...)) / number` is the number to compare across a change; the distribution around it is what tells you whether the comparison means anything.
- What is the difference between the -n and -r options of python -m timeit?`-n` is how many times the statement runs inside one timed loop; it amortizes the clock read and keeps the measured interval far above timer resolution. `-r` is how many times that whole loop is repeated and separately timed; it samples the environment so disturbance can be seen and discarded. Omit `-n` and timeit autoranges it to a loop lasting roughly 0.2 seconds.
- When would you report the median or the full distribution instead of the minimum?Whenever the variation comes from the code rather than the machine: data-dependent branches, occasional container resizes, cache hits and misses, or interpreter warm-up across repetitions. Also whenever the decision is about tail behaviour — a minimum is the best case and says nothing about p99. Keep the list `timeit.repeat` returns and quote median plus range.
- How large should one repetition be for the timer's resolution to stop mattering?Large enough that the total dwarfs the cost of a `time.perf_counter` call and its resolution — a few tenths of a second is the target timeit's own autoranging uses. If a whole timed loop finishes in microseconds, raise the inner loop count rather than the repetition count; more repetitions of an unreliably short measurement do not fix it.
saying these in an interview costs you the question
- Averages the repetitions because averaging sounds more rigorous
- Confuses the inner loop count with the number of repetitions
- Quotes a minimum for code whose runtime genuinely varies
- Ignores a worst repetition several times the best
- Compares numbers taken on a machine running other work