After a code change, your second performance trace of the same page load is 400 ms faster than the first. What could make that comparison wrong, and how would you establish that your change actually caused the improvement?
answer
- one pair is an anecdote
- noise moves hundreds of milliseconds
- interleave the runs, compare medians
- look for the cost you removed
- did the user-visible moment move?
basics
~20 sA single before/after pair proves little: run-to-run variance, cache state, background load and different code paths all move the number. Re-run both builds several times under identical interleaved conditions, compare medians, and check that the specific cost you targeted actually shrank.
solid answer
~50 sOne pair of traces is an anecdote. Page loads vary run to run — connection setup, server response time, cache and connection warmth, other processes on the machine, even which code path the data happened to take — and hundreds of milliseconds of that variance is ordinary. So I re-run both builds several times under identical conditions and compare medians rather than single runs, and I **interleave** the runs (A, B, A, B) instead of doing all of A then all of B, so drift in machine or server state cannot masquerade as my change. Then I do the check that matters more than the aggregate: open the after-trace and confirm the *specific* cost I targeted is smaller or gone. If total time dropped but the block I set out to remove is still there, something else moved and my conclusion is wrong. Finally I sanity-check that the improvement lands where a user notices it, not just in a total.
go deeper
Know that page-load timings vary between runs, so one faster recording is not proof; the same scenario has to be measured several times under the same conditions.
Explain concretely what makes runs vary — cache and connection warmth, server variability, background load — and why medians of repeated runs beat a single comparison.
Demonstrate the discipline: interleaved runs, identical conditions, and above all the mechanistic check that the specific cost you targeted actually disappeared from the trace, plus a re-check after deploying.
Own the standard of evidence the team applies before an optimization is allowed to add complexity, and be willing to say that a change whose effect sits inside the noise should not ship at all.
## Why one pair of traces is not evidence A performance recording is a sample from a noisy distribution, not a reading from a ruler. Load two identical builds of the same page twice and the totals will differ — sometimes by tens of milliseconds, sometimes by hundreds. If your change is worth 400 ms and typical noise is worth 300 ms, a single before/after comparison genuinely cannot distinguish "the fix worked" from "the second run got lucky". Senior engineers get caught by this constantly, because the result confirms what they hoped and confirmation feels like proof. ## Where the noise comes from - **Connection and server state.** The first request of a session pays for connection setup; a later one may not. Backend response time varies with its own load, warm caches and neighbours. - **Cache warmth on the client.** Anything already cached — resources, compiled script, DNS answers — makes a later run cheaper for reasons that have nothing to do with your diff. - **Machine state.** Other applications, background tabs, indexing, a build running in another terminal, and thermal throttling on a laptop all steal CPU from the recording. - **The page's own variability.** Different data, a different experiment bucket, a different ad, or a conditional code path can make two runs of "the same page" do genuinely different work. - **The profiler itself.** Recording adds overhead, and that overhead is not perfectly constant. ## Designing a comparison that survives scrutiny Hold everything constant except the change: - The **same build type** (production against production), served the same way. - The **same throttling settings** and the **same cache state** for every run. - The **same data and the same scenario**, performed identically each time. - A **quiet machine**, ideally a clean browser profile without extensions. Then run each build several times and compare medians rather than single values — a median resists the one anomalous run that a mean would absorb. Crucially, **interleave** the runs. Measuring all of build A first and all of build B afterwards confounds your change with everything that drifted in between: server warmth, machine temperature, a colleague's deploy. Alternating A, B, A, B distributes that drift across both arms. And take the magnitude seriously. If the spread within a single build overlaps the difference between builds, the honest conclusion is "no measurable effect", not "probably a small win". ## The mechanistic check, which is stronger than the aggregate Statistics tell you *whether* something changed; the trace tells you *what*. This is the step that distinguishes a real verification from a hopeful one: open the after-trace and look for the specific thing you set out to remove. - You deferred a script — is its evaluation block gone from the critical span, or merely moved somewhere it still blocks? - You removed a repeated computation — has that function's aggregated self time collapsed? - You made a resource load earlier — does its request bar now start earlier in the trace? If the total improved but the targeted cost is unchanged, you have learned that something else moved, and shipping the change on that evidence means shipping complexity that buys nothing. Conversely, if the targeted cost is clearly gone but the total barely moved, you have learned something equally valuable: that cost was not on the critical path, and the fix — however correct — was aimed at the wrong thing. ## When the number improves and the experience does not A change can remove real work and leave the user-visible moment exactly where it was, because the work was running in parallel with something slower, or after the moment that mattered. Always tie the verification back to the number you named at the start of the investigation, and if possible to the moment a user would perceive: when did the content actually appear, when did the click actually respond. Total scripting time is a diagnostic quantity, not a user experience. ## Confirm it survives the trip to production A local verification is necessary and not sufficient. The deployed build is compiled, minified, compressed and served from different infrastructure, and any of those can change the picture — a lazily-loaded chunk that behaved locally may be fetched over a slow connection in production. Re-check after deployment, with the same discipline: same scenario, same conditions, and look for the mechanism, not just the total. ## Knowing when to stop Verification has diminishing returns too. For a change with a large, clearly visible mechanistic effect, a handful of interleaved runs is plenty. For a marginal change where the effect sits inside the noise, the right answer is usually to drop it: an optimization you cannot demonstrate is complexity you cannot justify to the next engineer who reads the code.
- Why interleave the runs instead of measuring all of the old build and then all of the new one?Because anything that drifts over time — server warmth, machine temperature, a background process, a colleague's deploy — becomes perfectly correlated with the build order and is then indistinguishable from your change. Alternating the two builds spreads that drift across both arms, so a consistent difference is more plausibly caused by the diff rather than by when each block happened to run.
- The total time improved but the block of work you tried to remove is still in the trace. What do you conclude?That the improvement is not yours until proven otherwise. Either the change did something different from what you intended, or unrelated variance produced the gap. Re-run with the mechanism in mind, and if the targeted cost is genuinely unchanged, revert the change and re-open the investigation. Shipping on an unexplained improvement means carrying complexity with no demonstrated benefit.
- Your change removes 300 ms of scripting, yet content still appears at the same moment. How do you interpret that?The work you removed was not on the critical path to that paint — it ran in parallel with something slower, or after the moment you care about. That is genuinely useful information: it narrows where the real blocker is, and it warns you against reporting the scripting reduction as a user-visible win. Re-aim the investigation at whatever the paint is actually waiting for.
saying these in an interview costs you the question
- Treats a single faster run as proof the fix worked
- Compares a warm-cache run against a cold-cache one
- Runs all of build A, then all of build B
- Reports a total improvement without checking the mechanism
- Ships an optimization whose effect sits inside the noise