timeit says a rewritten bid-scoring function is 3x faster, but the ad-auction service's p99 did not move. What do you check?
answer
- A small share caps any speedup
- The fixture was small, warm and uncontended
- Ask whether the snippet consumed anything
- Setup built a steady state production never has
- Confirm at the service, not in the loop
basics
~20 sCheck, in order: the function's share of a request, whether the benchmark's inputs and warm state resemble production, and whether the timed snippet did the work at all. Then confirm end to end, not in a loop.
solid answer
~50 sStart with share of total time — a function that costs 2% of a request cannot move p99 by more than 2%, however fast it gets. Then interrogate the benchmark's conditions: the fixture was probably small and warm, built once in setup, while production sees larger and colder inputs; a memo cache pre-populated in setup makes every timed call a hit, whereas eleven concurrent bidder workers mutating that shared cache invalidate entries and rarely hit at all. Then check the snippet itself — timing a generator expression that is never consumed, or a call whose result is discarded lazily, measures almost nothing. Also remember the loop ran with the collector disabled, on warmed bytecode and warmed CPU caches. The resolution is always the same: reproduce the win against production-shaped inputs, and confirm it with a before-and-after latency histogram from the service.
code
python · 7 linesimport timeit
data = list(range(1000))
lazy = timeit.timeit("(x * x for x in data)", globals=globals(), number=100000)
real = timeit.timeit("[x * x for x in data]", globals=globals(), number=100000)
print(f"generator expression: {lazy:.4f}s")
print(f"list comprehension: {real:.4f}s")go deeper
Remember that a snippet timing describes the snippet under the conditions you set up, not the program. Before claiming a speedup, check that the code you timed actually ran and that the input resembled a real one.
Explain the concrete ways a benchmark diverges from production: input size, pre-warmed fixtures built in setup, lazily consumed statements that compute nothing, and a timed loop that runs with collection disabled.
Walk the diagnosis in order — share of request time, input shape, state warmth and contention, snippet validity, harness controls — then insist on end-to-end confirmation from the service's latency distribution before the win is claimed.
Set the bar for what counts as a performance result: which measurement a change must carry, who reproduces it, and when a micro-benchmark is merely a hypothesis. Without that, optimisation effort follows whichever number is easiest to produce.
## Step 1: does the function matter at all The first question is not about benchmarking; it is about arithmetic. If bid scoring is 2 milliseconds of a 90 millisecond request, making it three times faster removes at most 1.3 milliseconds. A p99 dominated by a downstream call, a lock wait, or serialization will not move, and no amount of micro-optimisation will change that. Establish the function's share of a request before doing anything else — with a whole-program measurement rather than a snippet timing — and if the share is small, the correct outcome is to stop and say so. ## Step 2: were the inputs production-shaped Micro-benchmarks are almost always run on the input the author could construct quickly. A scoring loop timed over 50 candidate bids can behave completely differently at 5,000: different container growth, different cache residency, a different balance between per-item work and fixed overhead. Two implementations frequently cross over — the one that wins on small inputs loses on large ones, because it trades a lower constant factor for worse growth. So pull real input sizes and real value distributions from the service and re-run the comparison at those sizes. If the crossover sits inside your production range, the honest answer is that neither implementation is faster; it depends. ## Step 3: was the state warm in a way production never is This is where the most convincing benchmarks lie. Suppose the scoring function consults a shared dictionary of advertiser quality scores. In the benchmark, setup populates that dictionary once, and every one of the timed iterations is a hit on a small, cache-resident structure. In the running bidder, eleven worker threads are reading and writing that same structure concurrently; entries are invalidated and rebuilt under contention, so the hot path the benchmark measured is the path production rarely takes. What you measured was a cache hit; what the service pays for is the miss plus whatever synchronisation the shared state requires. The general form: setup establishes a steady state, and the loop then measures only that steady state. Cold start, contention, eviction, growth and interleaving with other work all live outside the measurement by construction. ## Step 4: did the snippet actually do the work Python makes it easy to time nothing. A statement like `(score(b) for b in bids)` builds a generator object and stops; nothing is scored. The same trap catches `map`, `filter`, `zip`, `enumerate` and any lazily consumed iterator, an unawaited coroutine object, and a call whose expensive part happens only when the result is used. Compare against a version that forces consumption — a list comprehension, or wrapping in `list()` — and if the numbers differ by orders of magnitude you have found the bug in the benchmark rather than a speedup in the code. ## Step 5: what the harness controlled away The timed loop ran with cyclic garbage collection disabled, so the rewrite's allocation pressure was invisible. It ran the same code object hundreds of thousands of times, so since Python 3.11 the specializing adaptive interpreter had fully specialized it — a degree of warmth that a function called a few times per request may never reach. On Python 3.14 the JIT is still experimental and off unless `PYTHON_JIT=1` is set, so it is unlikely to explain a gap, but if the benchmark and the service run different build or flag configurations, that difference is worth eliminating before theorizing. The CPU was also warm: branch predictors trained, data resident in cache, clock boosted for a short burst in a way it will not be under sustained load. ## Step 6: measure where the answer lives The repair is not a better micro-benchmark. Ship the change behind a flag, compare latency histograms from the service before and after, and look at the percentile you actually care about. If you cannot ship it, build a load test whose inputs come from recorded traffic and whose concurrency matches production. A micro-benchmark is a hypothesis generator: it is excellent at telling you which of two expressions is intrinsically cheaper, and it has no authority at all over what a service's p99 will do. ## How to report this The answer an interviewer wants is a sequence, not a single cause: quantify the share, match the inputs, match the state, verify the snippet did the work, note the controls the harness applied, then confirm end to end. Saying '3x in a loop is not 3x in a service' is the conclusion; the credibility comes from the order in which you would rule things out.
- How would you turn this micro-benchmark into evidence an interviewer would accept?Re-run it at production input sizes taken from real traffic, with the shared state in the condition production leaves it in, and with the collector enabled if allocation pressure is part of the difference. Report the per-call median and spread rather than a lone best-of number, then confirm with a before-and-after latency histogram from the service behind a flag. The micro-benchmark proposes; the service decides.
- The rewrite is genuinely faster per call but the service gets slower. What could explain that?Allocation pressure that the timed loop hid because collection was disabled; extra memory that pushes a working set out of cache; a data structure that is faster to read but must be rebuilt under concurrent writes; or contention introduced by the synchronisation the new version needs. All of these are invisible to a single-threaded loop over a warm fixture and visible in a load test.
- Why can a timed loop be more optimized than the same function in production?Since Python 3.11 the adaptive interpreter specializes bytecode after a code object has run enough times, so a loop of a million iterations reaches a fully warmed state that a function called a handful of times per request never reaches. Add trained branch predictors and cache-resident data and the benchmark measures the best case the code can ever have.
saying these in an interview costs you the question
- Treats a micro-benchmark ratio as an end-to-end speedup
- Benchmarks a ten-element fixture and extrapolates to production
- Times a generator expression that is never consumed
- Ignores that the loop ran warm with collection disabled
- Assumes the optimized function is the request's bottleneck
- Compares runs taken on different machines or under other load